Native-speaker human data for the languages your evals keep missing.
Boltwork is a live human-data network of native speakers judging fluency, preference, and cultural correctness in long-tail languages the English-first shops cannot source and the crowd platforms cannot keep workers for. Every record carries a written rationale, an error-category tag, confidence, and full provenance.
The premium labeling shops are English-first and priced for US experts. The crowd platforms are losing the workers who made them work. Native-speaker preference data in the long tail stays scarce, expensive, and slow. Boltwork is built for exactly this gap.
Native speakers in the long tail
Swahili, Hindi, Tagalog, Vietnamese, Indonesian, Arabic, and more, judging the fluency and cultural correctness a non-native annotator or a synthetic judge simply cannot. This is the lane the premium expert shops do not staff.
Quality is the product, not a hiring bar
Hidden gold honeypots, multi-judge consensus across distinct workers and IPs, and per-worker trust scoring. We set caliber with a QA layer, so we do not need worker KYC, and we do not take regulated, PII, or safety-critical jobs.
Rationale and provenance on every record
Each judgment ships with the worker's written reason, a confidence value, an error-category tag, and provenance. You can audit and filter the data, not just trust a label.
A workforce that actually stays
Built on LightningFaucet's Bitcoin-paid global userbase. Instant Lightning payouts are how we retain a native-speaker panel in regions where conventional payout rails do not reach, instead of churning it.
How it works
You bring the prompts and candidate outputs. We return judged, quality-scored, provenance-tagged data.
Send us the task. Pairwise preference, quality rating, ranking, or eval, in the language(s) you need.
Native speakers judge it. Multiple distinct judges per item, each giving a reasoned choice plus an error tag on the weaker answer.
Consensus and gold QA resolve it. Distinct-IP consensus, gold catch-rate, and trust scores are computed per batch.
You get a clean export plus a quality report. JSONL with the choice, rationale, confidence, error tags, and provenance. Drop-in for RLHF and eval pipelines.
What a record looks like
Each resolved item carries the consensus choice, agreement, and the judges' rationales, confidence, and error tags. Illustrative of the export format:
[sw] "Andika barua pepe ya kuomba kazi" → preference: A (3 judges, 100% agreement)
judge 1: A | high | "A ni rasmi zaidi na ina muundo kamili wa barua"
judge 2: A | medium | "B inakosa anwani na hitimisho" tags: too_short
[hi] "मुझे योग के बारे में बताइए" → preference: B (3 judges, 67% agreement)
judge 1: B | high | "B स्पष्ट और सही है" tags (on A): factual_error
EnglishEspañolPortuguêsFrançaisहिन्दीBahasaTiếng ViệtالعربيةTagalogKiswahili+ more on request
What we do not do
Trust comes from being clear about the boundary.
No regulated, PII, or safety-critical data.
No worker KYC. Our QA layer, not identity, sets quality.
We are not the cheap expert-data vendor. If you need vetted English domain experts, the premium shops own that lane. We own the native-speaker long tail.
No over-promising on scale. We are at pilot scale today, a few hundred records per batch, which is exactly the size of a first design-partner pilot.
See the data before you commit
Send us 100 to 300 of your own items and target languages. We run a free pilot, native-speaker judged, and return labeled JSONL with the quality report so you can evaluate it directly. No cost, no commitment.
Boltwork is built on LightningFaucet, an established Lightning platform with real earners across many countries. That existing global, instantly-payable workforce is why our native-speaker panel exists on day one.
Boltwork · native-speaker multilingual human data for AI teams · [email protected]