Annotation Quality6 min read

How Tier 2 cities annotate differently, and why it matters

Give two annotators the same task, the same guidelines, and the same training, and you would expect roughly the same output. That assumption is baked into how most annotation work gets hired and managed. Set the rubric once, hire enough people, and consistency follows.

It does not hold up the way most teams expect, and the reason has nothing to do with skill or effort. It comes down to where the annotator is actually from.

What we saw in our own pilot

Before taking this to market, we ran a structured pilot to test our process end to end. 105 annotations across 7 distinct task types, completed by 14 annotators spread across 5 Indian states, all validated through our Intent Preservation Engine at approximately 93 percent intent preservation.

The point of the pilot was not just to test whether our validation process worked. It was also to see how much annotator background actually shapes the output, even when the guidelines are identical. What we saw lined up with something the research has been pointing to for a while, that location, dialect, and lived context change how a person interprets meaning, tone, and intent, even on a task that looks completely objective on paper.

Why this happens

Most annotation guidelines are written once, usually by a small team, usually in one city, sometimes in one country entirely removed from where the annotation work is actually done. That guideline document becomes the single source of truth for what "correct" looks like, and every annotator is expected to interpret it the same way, regardless of their own linguistic and cultural context.

That is a reasonable assumption for something like verifying a date format or checking a bounding box. It falls apart quickly for anything involving tone, sentiment, intent, or cultural nuance. A phrase that is read as neutral in one region can be read as sarcastic, dismissive, or even offensive in another. A response that sounds appropriately formal in one context can sound cold or evasive in a different one. None of that is a mistake on the annotator's part. It is the guideline failing to account for the fact that meaning is not the same everywhere.

Metro hub annotation pools tend to share more in common with each other than they do with the broader population a model is actually meant to serve. That is not a knock on metro annotators; it is simply a fact about who tends to get hired when annotation work is sourced from a small number of easy-to-access hubs. The result is a kind of hidden consistency that looks like quality but is really just narrowness.

Why this actually matters for model performance

This is not just an interesting quirk of human behavior; it shows up directly in how models perform once they are deployed. A model trained on annotation from a narrow geographic and cultural pool tends to inherit that pool's blind spots. It gets very good at understanding the kind of language, tone, and intent common to that specific group, and noticeably worse at understanding everyone else.

Sarvam AI's results are a strong real-world example of the opposite happening. Its model, built with a deliberately diverse annotator base rather than a metro-concentrated one, outperformed a model four times its size on Indic language benchmarks. That is not a marginal improvement. That is what happens when the annotator pool actually resembles the population using the model, instead of a narrow proxy standing in for it.

What we changed because of this

Running annotation across five states instead of one hub was not just about spreading out labor. It meant building guidelines that could hold up across different regional interpretations of tone and intent, not just one. It meant treating annotator disagreement, when two annotators from different states read the same text differently, as a signal worth investigating rather than noise to average out. And it meant validating every batch through our Intent Preservation Engine specifically to check whether meaning survived the process, not just whether the label matched a format.

That last part matters more than it sounds like it should. Diversity in the annotator pool only helps if the validation layer is actually built to catch and preserve that diversity, rather than smoothing every answer down to whatever reads as the most standard or safe interpretation.

The takeaway

If your model needs to work for more than one region, one dialect, or one cultural context, the annotation guidelines cannot be the only thing doing the work. Who is actually doing the annotating matters just as much, and in a lot of cases, matters more. A single rubric written in one place was never going to capture how the rest of the world actually communicates.

We are continuing to scale this approach, working toward 300 plus annotators by the end of the year, expanding the range of states, task types, and languages covered along the way. The pilot was never meant to be the finished product. It was proof that the approach holds up, and a foundation to keep building on.

Talk to us

Curious what this looks like on your own data?

Book a demo and we'll walk you through the process, the validation, and the real numbers behind it.