Skip to main content
AI & Model Data

How to Evaluate an AI Training Data Vendor: 12 Questions That Reveal Everything

Ananya Ploesu · · 6 min read

Checklist-style diagram of twelve evaluation questions for choosing an AI training data vendor1. Label2. Review3. EvaluateAI & MODEL DATAHow to Evaluate an AITraining Data Vendor:12 Questions ThatReveal EverythingDataplexLabs InsightsData · AI · Decisions

The short answer

Evaluate an AI training data vendor by asking direct questions about annotator qualification, disagreement adjudication, published agreement metrics, edge-case handling, data ownership, and independent eval-set construction. A vendor's evasiveness on any one of these, more than their headline scale, tells you whether the data will hold up under real use.

Why do the right questions matter more than the pitch

Most AI training data vendor pitches sound similar. Scale, speed, quality assurance, global annotator networks. The differences that actually determine whether a dataset works for fine-tuning or evaluation live underneath that pitch, in operational detail most vendors will only give you if you ask directly.

Model-ready data: Model-ready data is data that meets defined, measurable acceptance criteria for accuracy, agreement and coverage, not simply data that has been labelled.

We cover that distinction in full in what model-ready data actually means. This post assumes you already know you need labelled or annotated data and are trying to choose who builds it. Twelve questions, asked in order, will tell you almost everything you need to know before a contract is signed.

What are the 12 questions to ask an annotation vendor

  1. How are annotators recruited and qualified for this specific domain? Good answer: a described screening process with a domain test relevant to your task. Bad answer: a vague reference to a large global pool with no domain filter mentioned.
  2. How are disagreements between annotators adjudicated? Good answer: a named adjudication step, often a senior reviewer or domain expert, with disagreement documented. Bad answer: majority vote with no visibility into what happens on a tie or a genuinely ambiguous case.
  3. Are agreement metrics measured and will you share them for our project? Good answer: yes, reported per batch, with a defined threshold below which data is reworked. Bad answer: agreement is tracked internally but never shared with the client.
  4. How do you surface edge cases rather than force them into existing categories? Good answer: an escalation path where annotators flag unclear items for review rather than guessing. Bad answer: any item must be forced into one of the given labels, no matter how ambiguous.
  5. Who owns the resulting data and its IP? Good answer: a clear contractual statement that you own the delivered data outright. Bad answer: ambiguous language, or a right for the vendor to reuse your data for other clients.
  6. How is sensitive or regulated data handled? Good answer: specific controls named, such as access restriction, anonymisation options and retention limits. Bad answer: a generic compliance statement with no specifics offered when pressed.
  7. Is the evaluation set built independently of the training set? Good answer: yes, drawn from a separate pool with no overlap, explained clearly. Bad answer: the eval set is a random split of the same pool, which tells you little about generalisation.
  8. What happens when the specification changes mid-project? Good answer: a defined change process with impact on cost and timeline discussed upfront. Bad answer: no process exists, or every change is treated as a dispute.
  9. How is rework priced when quality falls short of the agreed bar? Good answer: rework against agreed acceptance criteria is included or clearly priced. Bad answer: rework is billed as new work with no reference to the original criteria.
  10. What are the acceptance criteria, in writing, before work starts? Good answer: measurable criteria such as agreement thresholds and error rates, agreed before delivery. Bad answer: quality is described only as “high” with nothing measurable attached.
  11. Is the claimed subject-matter expertise real or generic? Good answer: named qualifications or experience relevant to your domain, checkable on request. Bad answer: “our annotators are trained on your domain” with no detail on what that training involved.
  12. What will the vendor refuse to do? Good answer: a vendor names real limits, such as declining domains outside their expertise or data they cannot legally handle. Bad answer: a vendor claims no limitations at all, which is rarely true and rarely reassuring.

None of these questions are exotic. That is exactly why a vendor's fluency, or lack of it, on all twelve is a fast, reliable filter.

Which annotation quality metrics should you actually check

MetricWhat it measuresWhy it mattersWarning sign
Inter-annotator agreementHow consistently independent annotators label the same itemLow agreement means the labelling task itself is ambiguous, or annotators are undertrainedVendor cannot produce an agreement figure, or reports one figure for an entire project regardless of task difficulty
Adjudication rateProportion of items requiring a tie-break or senior reviewA very low rate on a genuinely hard task suggests disagreements are being suppressed, not resolvedAdjudication is described as “rare” with no actual rate offered
Rework rateProportion of delivered data rejected against acceptance criteria and redoneSome rework is normal; a vendor unwilling to discuss it usually has no process for measuring itVendor claims a zero rework rate, which most experienced buyers should treat as implausible
Eval-set independenceWhether the evaluation set is drawn from a separate pool to the training setAn eval set contaminated by the training pool overstates how well a model will generaliseThe eval set is simply a random holdout from the same labelling run
Edge-case escalation rateProportion of items flagged as ambiguous rather than forced into a labelA near-zero escalation rate on a genuinely ambiguous domain suggests annotators are guessingVendor cannot describe what happens when an annotator is unsure
Annotation quality metrics worth checking before you commit

Why does 'we have 50,000 annotators' answer a question you did not ask

Scale is a real advantage for some tasks: high-volume, low-ambiguity labelling where speed and cost genuinely matter more than nuance. It is a poor proxy for quality on anything that requires domain judgement.

A vendor answering “how do you handle disagreement on ambiguous medical text” with “we have 50,000 annotators across 40 countries” has told you their scale, not their quality process. The two are not the same question, and a vendor who consistently answers scale questions with scale answers is usually avoiding the process questions because the process is thin.

This is the real fault line in the market. The top of it, including Scale AI, Surge, Appen and Labelbox, is largely enterprise-locked and built for volume. The gap underneath that, in our view, is not more commodity labelling capacity. It is domain-expert review for tasks where a non-specialist genuinely cannot adjudicate a disagreement, which is a different service to buy and a different question to ask a vendor.

If your task needs domain judgement rather than volume, AI training data covers how we structure that work.

Should you build an annotation team in-house instead

Sometimes, yes. If your labelling task requires deep, ongoing domain expertise that only your own staff hold, and volume is modest, an in-house reviewer working directly with your ML team can outperform any outsourced process, because the feedback loop is immediate.

  • In-house makes sense when your domain expertise is rare, internal, and hard to specify well enough for an outside vendor to apply consistently
  • In-house makes sense when volume is low enough that a small number of trained staff can keep up without becoming a bottleneck
  • A vendor makes more sense when volume is variable or high, when you need annotators qualified across multiple domains, or when building and retaining an internal team is not a good use of scarce headcount

Whichever path you choose, apply the same twelve questions to your own process. An in-house team with no documented agreement metric and no independent eval set has the same weaknesses as a vendor who cannot answer those questions.

A readiness check before committing either way is a reasonable low-cost step, particularly if you are unsure whether your current dataset would pass acceptance criteria you have not yet written down.

What does a genuinely strong vendor answer actually sound like

Strong answers share a pattern regardless of the specific question. They are specific rather than aspirational. They name a process rather than a value. They admit limits rather than claiming none. And they are willing to show you a sample of real output against a real specification before you sign anything.

Our own answers to these twelve questions, along with how we structure delivery and acceptance criteria, are set out on how we work with AI teams and trust. We would rather you ask a competitor the same twelve questions and compare the answers directly than take either page on faith.

For teams testing an assistant built on data like this before it reaches customers, testing an AI assistant before it talks to customers covers the evaluation step that follows vendor selection.

Key takeaways

  • Twelve direct operational questions reveal more about a training data vendor than any pitch deck, and the good/bad answer pattern is consistent across all of them
  • Inter-annotator agreement, adjudication rate, rework rate, eval-set independence and edge-case escalation are the five metrics worth asking to see
  • A vendor's claimed scale answers a different question to their quality process, and leaning on scale when asked about process is a warning sign
  • The enterprise top of the market is largely locked to large accounts; the real mid-market gap is domain-expert review, not more commodity labelling
  • Apply the same twelve questions to an in-house team before assuming it is automatically cheaper or better than outsourcing

Questions buyers ask

How much does AI training data annotation typically cost to engage a vendor for?

Cost depends on task complexity, required domain expertise and volume, and reputable vendors price against an agreed specification rather than a flat per-label rate that ignores difficulty. Rather than comparing headline rates, ask what is included in that rate, particularly rework, adjudication and eval-set construction.

Can we build our own annotation team in-house instead of hiring a vendor?

Yes, and it is often the right choice when your domain expertise is specialised, internal and hard to specify externally, with volume low enough for a small trained team to manage. It becomes a poor fit when volume grows or domains multiply faster than you can hire and retain qualified reviewers.

What is inter-annotator agreement and why does it matter so much?

Inter-annotator agreement measures how consistently independent annotators assign the same label to the same item. Low agreement usually means the task definition is ambiguous or annotators are undertrained, and a dataset built on low agreement produces a model that has learned inconsistency rather than the intended distinction.

Is a vendor with more annotators always a safer choice?

No. A large annotator pool suits high-volume, low-ambiguity tasks well, but scale says nothing about whether disagreements are properly adjudicated or whether domain expertise is genuine. Ask process questions directly rather than treating headline scale as a proxy for quality.

How is the evaluation set kept independent of the training set?

A properly independent eval set is drawn from a separate sampling pool, ideally collected or reserved before training data is selected, rather than split randomly from the same labelling run afterwards. Ask a vendor to describe exactly how their eval set was sourced, since a random split from the same pool tells you very little about how the model will generalise.

What happens if we need to change the labelling specification partway through a project?

A vendor with a mature process will have a defined change procedure that accounts for cost and timeline impact, applied to affected items rather than the whole dataset. Vendors without such a process tend to either refuse changes outright or treat every change as a dispute, both of which slow the project down.

Should we ask for a sample delivery before signing a full contract?

Yes, this is one of the most useful steps available and most credible vendors will agree to it. A sample built against your real specification, checked against your own acceptance criteria, tells you far more than any reference call or case study.

One-page checklist

AI Data Vendor Evaluation Scorecard

Score a prospective annotation vendor against these groups before committing budget or a specification to them.

Three fields, delivered immediately. No newsletter spam.

Ananya Ploesu

Data & AI Lead, DataplexLabs

Works with operations, finance and machine learning teams on data collection, margin analysis and model-ready datasets.

Related reading

Not sure which model fits your team?

Ten questions, no sales call attached. The check scores your data, workflows and controls, and tells you which delivery model suits where you are.

One business-day response · NDA on request · No newsletter spam.

Next step

Discuss your use case

Bring one pain point, a data source, a workflow, a margin question. We'll come back with a focused assessment and a clear ROI hypothesis.

Get a focused reply within one business day

One business-day response · NDA on request · No newsletter spam.

Book a meeting

Talk to a data and AI lead, not a sales rep

Pick a 30-minute slot. Bring one problem. You leave with a scoped approach and a rough ROI range.

  • 30 minutes
  • Video call
  • Reply within 1 business day