Skip to main content
For AI labs & data teams·Nigerian languages

The Nigerian panel your Pidgin model deserves

Real Nigerian voice contributors for ASR, TTS, RLHF, and translation pairs. Ten Nigerian languages, narrated by the region you're shipping to. Consented, naira-paid, with AI-summarised QA on every clip.

Curated panel · early-access pricing · pilot decks under NDA on request.

First-party proof

We use this data to build our own model

9jatesters isn't just a data vendor. We're building NGPT, our own Nigerian-language model, on this exact panel — real Nigerian speakers, consented and naira-paid, across ten languages. Every collection and QA pipeline is proven on a first-party model before it's ever sold to a lab.

The factory eats its own cooking.

Free sample

See the data before you commit

A Nigerian Pidgin pilot preview: real speakers across 6 NG zones, 16-kHz audio + a Common Voice-compatible JSONL manifest. Read the actual transcripts — no email required.

Explore the free Pidgin sample
Language coverage

Ten languages, 480M+ speakers total

Nigerian Pidgin, Yoruba, Igbo, Hausa and Nigerian English, plus Fulani, Kanuri, Tiv, Edo and Efik — narrated by the region you're shipping to, not a generic English panel.

Pidgin
Most-spoken West African creole · 75M+ speakers
Yoruba
South-West Nigeria · 45M+ speakers
Igbo
South-East Nigeria · 30M+ speakers
Hausa
Northern Nigeria + Niger · 70M+ speakers
Nigerian English
Accented English · 200M+ speakers
Fulani (Fulfulde)
Northern Nigeria + Sahel · 40M+ speakers
Kanuri
North-East Nigeria + Chad · 10M+ speakers
Tiv
North-Central Nigeria · 5M+ speakers
Edo
Edo State · 5M+ speakers
Efik
Cross River · 4M+ speakers

What we collect for you

Voice, text, preference signals — all native to the languages we speak.

ASR / speech-to-text

Short clean reads + spontaneous narration. We tag dialect + age band + state for stratified training splits.

TTS voice modelling

Same speaker, 100+ utterances, identity-locked so the whole dataset stays one consistent human voice.

RLHF / preference data

Bilingual annotators rate AI outputs across Pidgin/English code-switching. Useful when your foundation model is English-trained.

Translation pairs

Sentence-level English ↔ Pidgin/Yoruba/Igbo/Hausa, source-tagged with regional variants.

Why this panel is different

Real Nigerians. AI-summarised QA.

Real people on real Nigerian phones and networks — no emulators, no VPNs, no synthetic personas. Every clip runs two-pass QA: Gemini transcribes and summarises it, then a human admin reviews, so you get decisions and friction points, not just raw audio. Contributors are consented and paid in naira to verified Nigerian bank accounts — withdrawals settle within 24 hours. Real people. Real accents. Honest provenance for your dataset.

Real Nigerians, real phones

Real people on real Nigerian phones and networks — no emulators, no VPNs, no synthetic personas.

AI-summarised QA

Two-pass QA: Gemini transcribes and summarises each clip, then a human admin reviews. You get decisions and friction points, not just raw clips. Failed clips never reach the buyer.

Fast naira withdrawals

Contributors are paid in naira to verified Nigerian bank accounts — withdrawals settle within 24 hours. Curated panel across every region of Nigeria; pilot jobs route within hours of brief acceptance and volume scales as the panel grows.

NDPA-aligned, self-hosted on hardware we own

Self-hosted on our own hardware — not a third-party cloud — with an honest privacy posture. We invoice in USD or NGN — your call. Standard data-licence agreement included.

Pricing

Priced in audio-hours, like the rest of the industry

Four tiers — ASR, TTS, RLHF, translation. Headline rates are early-access pilot pricing in USD. Volume + multi-language discounts on contract. NDAs welcome. Failed clips never billed.

Most popular

Most popular · entry tier

ASR / speech collection

$45/ audio-hour delivered

Equivalent to ~$3 / approved clip at typical prompt cadence.

Short prompts + spontaneous narration in your chosen language. Tagged by dialect, age band, state, and gender for stratified training splits.

  • 16-kHz mono .wav + JSON manifest with prompt text, duration, anonymised contributor id
  • Two-pass QA: Gemini transcript-back + human admin review
  • Failed clips never billed — you pay for approved-only
  • Stratified demographics across all 6 Nigerian geopolitical zones
  • Most pilots come back within a few days; scale by batch from there
Start an ASR pilot

Single-speaker, identity-locked

TTS voice modelling

$100/ audio-hour delivered

Same speaker, 100+ utterances, identity-locked so the whole dataset stays one consistent human voice. Studio-grade re-takes included.

  • Single speaker per contract, identity-locked for voice consistency
  • Scripted prompts (you supply) + emotional range coverage on request
  • 44.1-kHz mono .wav with per-utterance metadata
  • Up to 3 re-take rounds per session at no extra cost
  • Talent rights signed via standard data-licence agreement
Brief us a TTS speaker

Where the dollars are · frontier-lab tier

RLHF / preference data

$120/ annotator-hour

Bilingual Nigerian annotators rate LLM outputs across Pidgin/English code-switching, cultural appropriateness, and factual grounding. Useful when your foundation model is English-trained but your users aren't. Priced at the Mercor/Surge benchmark for expert preference annotation — well under US $250-1,000/hr specialist tier.

  • Side-by-side preference rating (Bradley-Terry compatible exports)
  • Free-text rationale per ranking for SFT/DPO training data
  • Annotators screened on bilingual fluency + cultural reasoning
  • Adversarial red-team prompts available on contract
  • Calibration set + IAA (inter-annotator agreement) report included
Scope an RLHF pilot

Sentence- or document-level

Translation pairs

$0.20/ word translated

English ↔ Pidgin/Yoruba/Igbo/Hausa. Native translators (not back-translation from English MT) with regional-variant tagging. Useful for MT bootstrapping and benchmark sets.

  • Sentence-aligned pairs in TMX, CSV, or JSONL
  • Tagged with regional variant (e.g. SW-Yoruba vs Ekiti, Naija vs Sapele Pidgin)
  • Per-pair quality score from two-translator overlap
  • Source-text de-duplication + length-balance audit
  • Available with reference back-translations for adversarial sets
Brief a translation set

Safety panel · enterprise

Adversarial red-team

$120/ red-team-hour

Trained Nigerian red-teamers stress-test your model in Pidgin, Yoruba, Igbo, and Hausa. Low-resource languages are a known jailbreak vector, and most published eval cards do not cover them at all. We deliver structured adversarial prompts, severity ratings, and reproduction steps — at a fraction of US safety-contractor rates ($90-200/hr).

  • Adversarial prompt library scoped to your model's safety policy
  • Severity rubric + reproduction steps per finding
  • Cross-language jailbreak coverage (Pidgin/Yoruba/Igbo/Hausa)
  • Real Nigerian red-teamers, consented and naira-paid — no emulators, no VPNs, no synthetic personas
  • Aggregated report under NDA + raw finding logs on request
Scope a red-team engagement
Standard delivery

16-kHz mono .wav (ASR) or 44.1-kHz (TTS), JSONL manifest, S3-presigned or direct GCS bucket.

Billing

Invoiced in USD or NGN. Net-30 on contract. Standard data-licence agreement included.

Pilot timing

Most pilots come back within a few days; bigger jobs ship in batched deliveries. We'll confirm timing when you book.

Frequently asked

AI-data buyer FAQ

Quick context for procurement reviewers + ML engineers who haven't sourced African-language data before.

What is RLHF data and why does it cost more than ASR?
RLHF (Reinforcement Learning from Human Feedback) data is pairs or rankings of model outputs labeled by humans for preference, factuality, or safety. It costs more than speech (ASR) collection because the annotator has to read model outputs, judge them, and write rationale — skilled bilingual work, not just recording a clip. Our rate of $120/annotator-hour sits at the Mercor/Surge benchmark for expert preference annotation, well under the US specialist tier of $250-1,000/hour.
ASR vs TTS data — what's the difference?
ASR (Automatic Speech Recognition) data is many speakers saying many things — used to train a model to convert speech to text. TTS (Text-to-Speech) data is one identity-locked speaker recording 100+ utterances, used to train a model to generate speech in that voice. ASR values speaker diversity; TTS values speaker consistency. We price ASR at $45/audio-hour (volume game) and TTS at $100/audio-hour (identity + studio quality).
How much does it cost to label an audio-hour of training data?
Industry-wide 2026 rates vary by language coverage and annotation depth. Bulk-collection benchmarks (Karya, Sama, Scale AI, Surge) cluster around $7-15/hour for high-resource speech, while quality vendors quote $60-95/hour for African-language collection with QA. Our delivered rates sit between those bands on purpose: low-resource Nigerian languages price above bulk English because recruitment, dialect tagging and two-pass QA are real costs — and below the big-vendor quotes because the panel is ours. We disclose buyer-billed rates openly above; tester pay is published on /pricing so the spread is transparent.
Who pays for Nigerian-language training data today?
Public 2025-2026 buyers: Google (WAXAL · 11K hours · Feb 2026), Meta (funded NaijaVoices · 1,867 hours via Lacuna Fund), Cohere Labs (Aya-Earth · Feb 2026), Microsoft (via Karya + Awarri partnership), the Nigerian government ($3.5M for the N-ATLAS LLM). Conspicuously absent from public African-language buying: Anthropic, OpenAI, xAI, Mistral. African languages are not in Claude 4.7 / Llama-3 / Gemma2 official support — that gap is the prospect list we're building for.
Choosing an AI model and need proof it handles Nigerian languages first?
Our benchmark arm runs paid private evaluations: your candidate systems, tested blind on your own scenarios, judged by this same rater network. Fixed prices at 9jaBench Eval.