WER vs. CER vs. FER: Why African Language ASR Teams Need Better Evaluation Metrics
Word Error Rate lies about African language ASR performance. Feature Error Rate reveals what WER hides—and changes how you evaluate models.

Word Error Rate (WER) is the standard metric for evaluating automatic speech recognition systems. It counts how many words the model got wrong. For English ASR, that works reasonably well. For African language ASR, new research from AfricaNLP 2026 demonstrates that WER actively mischaracterizes model performance—collapsing phonological, tonal, and lexical errors into a single number that hides what the model actually learned. If you're evaluating ASR vendors for Yoruba IVR, fine-tuning Whisper on Wolof call-center data, or deciding whether a Hausa transcription model is production-ready, the metrics you track determine whether you ship the right system.
The AfricaNLP paper, "Linguistically Informed Evaluation of Multilingual ASR for African Languages," tested Whisper and MMS (Meta's Massively Multilingual Speech) models on Yoruba and found something striking: a model with 91% WER—complete failure by conventional standards—had only 26.7% Feature Error Rate (FER). The model was getting most of the phonological features correct (consonant place and manner, vowel height and backness, tone) but struggling with lexical identity. WER said the system was useless. FER revealed it was closer to correct than the topline metric suggested.
Why Word Error Rate Breaks Down for African Languages
WER treats every incorrect word equally. If the reference transcription is "ọmọ" (child in Yoruba) and the hypothesis is "ọmọ́" (with high tone instead of mid), WER counts that as one full error—the same penalty as transcribing "ọmọ" as "dog." For languages where tone, vowel length, and nasalization carry lexical meaning, this flattening erases the distinction between near-misses and catastrophic failures.
Character Error Rate (CER) helps—it counts individual character substitutions, insertions, and deletions—but it still can't distinguish a tonal diacritic error (which might change meaning entirely) from a random typo. The AfricaNLP paper argues that WER and CER are "coarse-grained" and "do not reflect the linguistic nuances" of tone languages, agglutinative morphology, or underdocumented phoneme inventories.
The practical consequence: you might reject a model that's 70% of the way to production-ready because WER makes it look like a 10% solution. Or you might ship a model with acceptable WER that systematically confuses tone, rendering the transcriptions semantically incorrect even when the words look right.
What Feature Error Rate (FER) Reveals That WER Hides: The AfricaNLP 2026 Evidence
Feature Error Rate decomposes every phoneme in the hypothesis and reference into a vector of binary phonological features—things like [+voice], [+nasal], [+high tone]—and calculates error at the feature level. The AfricaNLP 2026 study applied FER to Whisper and MMS outputs on Yoruba CommonVoice data.
Key findings:
- Whisper medium on Yoruba: 91.0% WER, but only 26.7% FER. The model captured most consonant and vowel features correctly; it just couldn't map those features to the correct lexical items.
- MMS 1B on Yoruba: 31.8% WER, 15.7% FER. Better overall, but still a 2:1 ratio between word-level and feature-level error.
- Tone Error Rate (TER): Whisper achieved 38.9% TER on Yoruba—meaning it got tone wrong nearly 40% of the time, even when the segmental phonemes were correct.
The authors conclude that WER "mischaracterizes" ASR quality for African languages because it conflates multiple error types. A high WER might mean the model has bad acoustic training, or it might mean the model learned the phonetics but lacks lexical coverage. FER and TER separate those failure modes.
For engineering teams, this matters at procurement time and during fine-tuning. If you're evaluating a vendor's Hausa ASR, ask for FER and TER alongside WER. If you're fine-tuning on 50 hours of Wolof call-center audio, track feature-level error to diagnose whether you need more acoustic data (high FER) or more lexical coverage (high WER, low FER).
When to Use WER, CER, FER, and TER in Production
Each metric answers a different question. Here's when to use each:
| Metric | What It Measures | When to Use It | African Language Consideration |
|---|---|---|---|
| WER (Word Error Rate) | Lexical accuracy: did the model pick the right words? | Benchmarking vendor claims; tracking end-to-end quality for non-technical stakeholders. | Can hide near-miss tonal errors or morphological confusions that break meaning. |
| CER (Character Error Rate) | Surface-level accuracy; helpful for scripts with complex orthography (e.g., Amharic, N'Ko). | Evaluating models where word boundaries are ambiguous or where you're testing transcription fidelity. | More granular than WER but still treats all character errors equally. |
| FER (Feature Error Rate) | Phonological accuracy: did the model get the consonants, vowels, tone, and nasalization right? | Diagnosing training issues during fine-tuning; deciding whether to prioritize acoustic or lexical data. | Reveals whether a high-WER model is actually close to correct. |
| TER (Tone Error Rate) | Tonal accuracy (for Yoruba, Hausa, Igbo, etc.). | Any production system where tone changes meaning; essential for tonal languages. | Must be calculated separately; not included in standard ASR evaluation libraries. |
For CPaaS platforms shipping IVR in Yoruba or Twi, we recommend tracking WER for stakeholder reporting, FER for internal model tuning, and TER as a gate check before production deployment. A model with 40% TER will produce transcriptions that look plausible to a non-speaker but convey the wrong meaning—a dangerous failure mode for voice bots handling customer service or financial transactions.
The Procurement Implication: How to Read Vendor Benchmarks for Yoruba, Wolof, and Hausa ASR
When evaluating ASR vendors or cloud APIs for African languages, the benchmark they publish probably reports WER on a held-out test set. That number is not useless, but it's incomplete. Here's what to ask:
- What's the test set? If it's Common Voice (which the AfricaNLP paper used), that's read speech, not conversational. Read-speech WER will be 10-20 points lower than real-world call-center or IVR audio.
- Do they report TER? If the language is tonal and they don't mention tone error, they haven't evaluated the failure mode that matters most.
- What's the confusion matrix for phonologically similar phonemes? A vendor with low WER might be systematically confusing /p/ and /b/ (which matters for Wolof) or collapsing nasal vowels (which matters for Fulfulde). Ask for phoneme-level diagnostics.
- Is the model licensed for commercial use? Many "state-of-the-art" African language models are fine-tuned on Common Voice, which is CC0 for the audio but restrictive for derivatives if you're building a SaaS product. This is orthogonal to evaluation metrics, but it's the other half of the procurement decision.
Clear Global's research on African language ASR scaling found that "data quality and linguistic diversity" matter more than raw data volume. A vendor with 500 hours of clean, phonetically balanced Hausa audio and FER-informed evaluation will outperform a vendor with 2,000 hours of scraped, unannotated audio evaluated only on WER.
What Afriklang's Wolof Benchmark Measures (and Why It Matters)
Afriklang's published Wolof benchmark fine-tuned Whisper medium on 30 hours of image-prompted Wolof speech and evaluated on sentiment classification (F1 macro 90.0% vs. 45% for the best frontier LLM). We didn't report WER for the benchmark task because sentiment classification is a downstream task, not transcription—but the training data was evaluated with inter-annotator agreement >80% at the phonetic feature level before fine-tuning.
Why that matters: phonetic feature annotation (marking tone, vowel length, nasalization) at capture time means the training data encodes the distinctions that WER erases. When you fine-tune on Afriklang data, you're training on speech where the model can learn tone and nasalization from clean labels, not noisy heuristics. That's the difference between a model that hits 30% WER with 15% TER (useful) and 30% WER with 40% TER (not deployable).
Our catalogue includes Luganda, Ewe, Fulfulde, Hausa, Wolof, Twi, and Yoruba. All datasets are captured with image-prompted elicitation (natural, conversational speech) and annotated by native speakers with AI-assisted quality filtering. Every dataset ships with a compliance trail and a commercial license.
Evaluation Checklist: What to Track Beyond WER When Fine-Tuning African Language Models
If you're fine-tuning Whisper, MMS, or a custom model on African language data, track these metrics alongside WER:
- Feature Error Rate (FER): Decompose errors into consonant, vowel, and suprasegmental features. The AfricaNLP paper provides a Python implementation using PanPhon for feature extraction.
- Tone Error Rate (TER): Calculate separately for any tonal language. If TER is >30%, your model is guessing.
- Phoneme Confusion Matrix: Which phonemes does the model conflate? For Yoruba, watch /ɔ/ vs. /o/; for Wolof, watch prenasalized stops.
- Semantic Error Rate: For downstream tasks (IVR intent classification, sentiment analysis), measure how many WER-counted errors actually change meaning. A high WER with low semantic error is often shippable.
- Per-speaker WER: If variance is high, you need more speaker diversity in your training set.
Kili Technology's ASR guide notes that "benchmarks are only as good as the test data distribution." For African languages, that means test sets must include tonal minimal pairs, dialectal variation, and code-switching if your production use case involves any of those.
Word Error Rate is not wrong—it's just blind to the linguistic structure of African languages. Feature Error Rate and Tone Error Rate reveal what WER hides: whether the model learned the phonology, whether tone is under control, and whether a high-WER system is actually 70% correct instead of 10% correct. For CPaaS platforms and AI teams shipping voice products in West Africa, the right evaluation metrics are the difference between rejecting a near-ready model and shipping a plausible-looking system that gets the meaning wrong.
If you're building or procuring ASR for African languages and need speech data with clean feature-level annotation, browse Afriklang's catalogue or book a discovery call to discuss your evaluation requirements.