Most preference and evaluation data is collected in English by non-native annotators or, more and more, by a model judging another model. That works until you need to know whether a Swahili or Hindi response is actually fluent and culturally right, which is exactly where a non-native annotator and a synthetic judge both fall down. Boltwork is built for that gap: native speakers judging fluency, preference, and cultural correctness in long-tail languages, with the quality controls that make small-crowd data trustworthy.
This post is the honest version of how it works, including the numbers, a mistake our own QA layer caught, and the limits.
The lane we picked, and the one we did not
We are not the cheap expert-data vendor. The premium shops own the vetted-English-expert lane and vetting is their product. We deliberately do not compete there. Our lane is the native-speaker long tail: the underserved languages the English-first shops cannot source and the generic crowd platforms cannot retain workers for. Caliber in our lane is set by a QA layer, not a hiring bar, which is what lets us operate without worker KYC. The flip side of that choice is a hard boundary: we do not take regulated, PII, or safety-critical work.
Quality is the product
Every item is judged independently by multiple native speakers across distinct workers and distinct IPs, so a single person cannot carry a label. Three things gate the data: hidden gold honeypots with a known answer, seeded invisibly, where misses cost trust; consensus across distinct judges, so a record releases only when independent native speakers agree; and per-worker trust scores built from gold performance and agreement history, which weight whose judgment counts.
The numbers, both eras of them. On our earlier pilot set of easier pairs, the crowd ran about 98% gold-task accuracy and about 95% inter-annotator agreement. We then retired that pool on purpose, because easy pairs teach a model nothing, and rebuilt the task set around genuinely hard items. On the current hard set, trusted judges run about 89% on gold checks, and per-item agreement sits around 66%. That lower agreement number is not a defect. It is what informative data looks like: if every judge always agreed, the pairs would be too easy to be worth labeling.
Hard pairs, not obvious ones
Easy preference data, a clearly strong response versus a clearly weak one, is close to worthless because everyone agrees and the model learns nothing. The current set is hard pairs: near-ties and planted-flaw items, judge-filtered so that telling them apart actually requires native fluency. That is where native-speaker judgment earns its keep.
Our QA layer caught our own generator
Here is the mistake, because a methodology post that only reports wins is marketing. Part of our gold pool is built by planting a single subtle flaw into a strong answer. An audit of consensus behavior found that our generator had been planting flaws without verifying them: in many pairs the rewrite model returned the answer essentially unchanged, which made the honeypot unanswerable, and on some of them every single judge unanimously picked against the official gold answer. Unanimous disagreement from independent native speakers is not a bad crowd. It is a broken answer key. We retired 62% of the honeypot pool, recomputed every worker's trust score against the remaining valid gold, and changed the generator so a planted flaw must be independently verifiable before it is ever allowed to gate a worker. The consensus layer surfaced the miscalibration exactly the way it is supposed to surface bad data. The 89% figure above is measured after that correction, against valid gold only.
Rationale and provenance on every record
A bare label is hard to trust and impossible to audit. So each judgment ships with the worker's written rationale, a confidence value, an error-category tag such as fluency, cultural fit, or factuality, and full provenance. You can filter on error type, inspect why a call was made, and trace where each item came from, rather than taking a label on faith.
Why we pay in Bitcoin over Lightning
This is an operational choice, not a slogan. At the long tail, the workers are in regions where PayPal and KYC friction quietly break the unit economics of paying people for small tasks. Instant Bitcoin payouts over Lightning are how we reach and retain native speakers there at all. To date we have paid workers across the network with zero fraud incidents, which the consensus and gold layer, not identity checks, is responsible for.
The honest limits
We are at pilot scale. The crowd is large enough for a real design-partner pilot of a few hundred records, not yet for a large contract, and coverage per language is uneven: our deepest queues today are English and Spanish, while the languages we most want to serve are exactly where we are recruiting native speakers hardest. We say so here and in the dataset stats. We would rather under-promise and deliver a clean small pilot than oversell capacity we do not have yet.
Try it on your data, free
If you run evals or RLHF in a language the English-first shops cannot cover, send us 100 to 300 of your own items and we will return labeled JSONL in about a week with full QC stats. No cost, no commitment. You can see exactly what a delivered record looks like, with rationales, error tags, and provenance, at lightningfaucet.com/boltwork/data/