Data annotation
The broad term — and a job description that has changed more in three years than in the decade before.
Definition
Data annotation is the practice of adding labels, judgments or structure to raw material so that a model can learn from it or be measured against it. It is the broad term, and it spans an enormous range of difficulty: drawing a box around an object, marking a sentence as positive or negative, transcribing speech, writing a rationale for why one answer is better than another, adjudicating whether a clinical recommendation is safe. What unites them is that a human decision is being recorded in a form a training or evaluation process can consume.
Every annotation programme is really two artefacts: the guidelines and the annotations. The guidelines define what a label means, what to do at the boundaries, and how to handle the cases nobody anticipated. Most annotation failures are guideline failures rather than annotator failures, and they surface as disagreement between annotators who are each behaving reasonably. Mature programmes iterate the guidelines against real disagreements, keep reference items with a known correct answer, route hard cases to arbitration, and record the guideline version each judgment was made under.
The character of the work has shifted with the models. Tasks that were once bulk and mechanical are increasingly pre-filled by a model, and the human contribution moves to verification, correction and the cases the model cannot settle — which raises the expertise required per decision even as the volume per decision falls. On frontier material the annotator is no longer applying a label from a fixed set; they are supplying a judgment only someone trained in the discipline can make. That is why quality here is now argued in terms of who did the work and how it was measured, rather than how much of it there was.
How DeepenSkill approaches it
DeepenSkill engages the specialists who can make those judgments, and reports the measurement rather than a testimonial. Everyone clears the same bar before starting: credentials verified, a standardized domain test passed, then a practical task scored against gold reference data. On live work the team measuring the experts is separate from the team running them — the platform rejects a validation report authored by the party that delivered the work — agreement statistics are reported with the method named, the item and rater counts stated and a 95% interval attached, and every judgment traces to a verified contributor, a task spec and a timestamp. See expert data annotation work, the narrower data labeling case, and a sample verification bundle you can recompute yourself.
FAQ
What is data annotation?
Adding labels, judgments or structure to raw material so a model can learn from it or be measured against it — from a simple tag through to a written rationale only a specialist can produce.
What is the difference between data annotation and data labeling?
Annotation is the broad term for any added judgment or structure. Labeling usually means assigning a value from a defined set. In practice the two words are used interchangeably.
How is annotation quality measured?
With reference items whose correct answer is established independently, agreement statistics between annotators who worked independently, arbitration of disagreements, and a record of which guideline version each judgment was made under.
Where this shows up in the work