Measured on 8 and 9 October 2026
Results and method
All our numbers, with their limits. They come from test sets the models did not see in training; where a test is less clean than it looks, we say so.
Speech recognition
Transcriber v4
Base: the w2v-BERT 2.0 family (MIT licence), trained by us for Baoulé. Measure: character error rate (CER) and word error rate (WER), after lower-casing and removing punctuation; tone marks are kept. Lower is better.
553 sentences read over the phone by 4 speakers who are not in the training data. WER 28.3%.
Same test, through our server, spaces not counted. 95% CI [10.5–11.8]. WER 28.3% [26.9–29.9].
| Test set | Clips | CER | WER |
|---|---|---|---|
| Phone, speakers never heard | 553 | 10.3% | 28.3% |
| Field recordings, public test set | 479 | 8.4% | 25.7% |
| Sentences read by volunteers, public test set ¹ | 562 | 8.0% | 20.1% |
| Read texts, public test set ² | 192 | 4.1% | 11.6% |
| Bible reading, held-out books | 420 | 4.1% | 9.8% |
In this table, CER counts spaces as characters. ¹ Sentences differ from training, but all 15 speakers in this test also appear in our training data: an optimistic figure. ² The same kind of texts as the phone test; this set does not let us check whether its speakers also appear in its training part.
| Test set | CER, no spaces | CER, spaces counted | WER |
|---|---|---|---|
| Phone, speakers never heard (553) | 11.2% [10.5–11.8] | 10.3% [9.7–10.9] | 28.3% [26.9–29.9] |
| Sentences read by volunteers (562) ¹ | 8.7% [7.7–10.0] | 8.0% [7.0–9.2] | 19.9% [18.3–21.7] |
Text overlaps
An automatic check compares the phone test with everything used in training. No shared recordings or speakers. But 16 test sentences appear word for word in our training texts (13 are one- to three-word expressions, 3 are full sentences found in earlier training data), and 3 more are near-duplicates.
Same kind of text
The phone test and one of the training sets added for v4 are readings of the same kind of text. Identical and near-identical sentences were removed from training, but part of the gain on this test may come from that similarity.
Two definitions of CER
Our internal evaluation counts spaces; the API benchmark does not. Measured the same way, the API and the internal evaluation agree (10.3%).
Translation
Baoulé ↔ French
“Everyday” test set: 316 sentence pairs never seen in training. Measure: chrF++ (how closely the characters and words match a human reference translation, 0 to 100). Higher is better. Baoulé is scored without tone marks.
| Model | Base and licence | Baoulé → French | French → Baoulé |
|---|---|---|---|
| Commercial (round 5, 9 October) | MADLAD-400 family, Apache 2.0 licence: commercially usable | 34.7 | 34.1 |
| Research (round 3, 8 October) | NLLB-200 family, non-commercial licence (CC BY-NC 4.0) | 32.3 | 38.0 |
- Confidence intervals. Measured through the API, the Research model scores 32.2 [30.3–34.3] and 37.5 [35.5–39.6] (the server decodes with slightly different settings). For the Commercial model, the gain over its previous round is +1.6 [0.6–2.6] into French and +2.2 [1.2–3.2] into Baoulé (bootstrap, 1,000 resamples).
- Which one to use? The Research model is best into Baoulé; the Commercial model is best into French and the only one you can use in a paid product.
| Part | Sentences | Commercial bci → fra | Commercial fra → bci | Research bci → fra | Research fra → bci |
|---|---|---|---|---|---|
| Free conversation | 166 | 31.6 | 29.7 | 28.2 | 32.2 |
| Dictionary examples | 112 | 37.4 | 38.8 | 37.3 | 47.0 |
| Health questions and answers | 38 | 43.8 | 46.5 | 43.4 | 53.1 |
| Register | Sentences | Commercial bci → fra | Commercial fra → bci | Research bci → fra | Research fra → bci |
|---|---|---|---|---|---|
| Tales | 54 | 23.1 | 25.7 | 20.6 | 28.4 |
| Bible, held-out passages | 227 | 37.8 | 35.0 | 47.4 | 39.5 |
A test close to training
The test sentences come from the same sources (same documents, same speakers) as part of the training data. The sentences themselves differ: any training pair identical or very close (90% similar or more) to a test sentence was removed. The numbers are therefore on the optimistic side.
Our weak spot
Our weakest part is free conversation, and tales are weaker still. That is exactly what the Voix Baoulé programme is meant to improve: real conversations, in real conditions.
Small tests, for now
316 sentences is not many: a gap of one or two points between two models can be chance. There is no standard public test set for Baoulé yet; we are building one.
Our method
Four steps, the same for speech and for text.
1. Gather
Bilingual books scanned and read by our OCR, public recordings with their text, then our own recordings, consented and paid.
2. Unify the spelling
Older texts are converted to the modern spelling (ɛ, ɔ…). Every rule is measured: on a 14,499-word trial corpus, the final conversion recognises 88.9% of words, against 70.8% with simple rules.
3. Verify the alignment
Every audio clip is aligned with its text and scored, every sentence with its translation; doubtful pairs are removed before training, and samples are checked by hand.
4. Measure on held-out data
Test sets are split off before training; an automatic check looks for shared recordings, speakers and sentences; every score comes with its confidence interval (1,000 bootstrap resamples).
What we don't publish
We do not publish our model weights, the detailed list of our data sources, or the details of our processing pipelines. A question about these measurements? Write to contact@abyshire.ci.
Last updated: 10 October 2026.