Quick context for procurement reviewers + ML engineers who haven't sourced African-language data before.
- What is RLHF data and why does it cost more than ASR?
- RLHF (Reinforcement Learning from Human Feedback) data is pairs or rankings of model outputs labeled by humans for preference, factuality, or safety. It costs more than speech (ASR) collection because the annotator has to read model outputs, judge them, and write rationale — skilled bilingual work, not just recording a clip. Our rate of $120/annotator-hour sits at the Mercor/Surge benchmark for expert preference annotation, well under the US specialist tier of $250-1,000/hour.
- ASR vs TTS data — what's the difference?
- ASR (Automatic Speech Recognition) data is many speakers saying many things — used to train a model to convert speech to text. TTS (Text-to-Speech) data is one identity-locked speaker recording 100+ utterances, used to train a model to generate speech in that voice. ASR values speaker diversity; TTS values speaker consistency. We price ASR at $45/audio-hour (volume game) and TTS at $100/audio-hour (identity + studio quality).
- How much does it cost to label an audio-hour of training data?
- Industry-wide 2026 rates vary by language coverage and annotation depth. Bulk-collection benchmarks (Karya, Sama, Scale AI, Surge) cluster around $7-15/hour for high-resource speech, while quality vendors quote $60-95/hour for African-language collection with QA. Our delivered rates sit between those bands on purpose: low-resource Nigerian languages price above bulk English because recruitment, dialect tagging and two-pass QA are real costs — and below the big-vendor quotes because the panel is ours. We disclose buyer-billed rates openly above; tester pay is published on /pricing so the spread is transparent.
- Who pays for Nigerian-language training data today?
- Public 2025-2026 buyers: Google (WAXAL · 11K hours · Feb 2026), Meta (funded NaijaVoices · 1,867 hours via Lacuna Fund), Cohere Labs (Aya-Earth · Feb 2026), Microsoft (via Karya + Awarri partnership), the Nigerian government ($3.5M for the N-ATLAS LLM). Conspicuously absent from public African-language buying: Anthropic, OpenAI, xAI, Mistral. African languages are not in Claude 4.7 / Llama-3 / Gemma2 official support — that gap is the prospect list we're building for.
- Choosing an AI model and need proof it handles Nigerian languages first?
- Our benchmark arm runs paid private evaluations: your candidate systems, tested blind on your own scenarios, judged by this same rater network. Fixed prices at 9jaBench Eval.