Note
How the machine learned to hear: the companion to the second article
The exam before the training: 34.6 % of words wrong, four audiobooks that sat in the material twice, an hour and a half of fine-tuning on a home GPU — and 17.0 % at the end. With the transcripts, lined up side by side.
The second article of the series is out on ana-yurt.com, and it is about the ears: speech recognition for Crimean Tatar — a computer listening and writing down what it hears. This page is its companion: the same transcripts as in the article, lined up row by row, so the difference can be seen rather than taken on trust.
It started with a humbling number. The project already had a large recogniser fine-tuned on our language, and it was already in use — but it had never sat a fair exam. The first thing I did in this stage was not to improve anything: I set the exam. Two whole books, two narrators the model had never heard, 893 clips, 1 hour 52 minutes. The result was 34.6 % of words wrong — every third word. And before running anything at all, I compared the training material against the test material not by file name but by text — and found that four audiobooks were sitting in my material twice, cut by two different tools under completely unrelated names. For the very book I was about to test on, 96.9 % of clips had a twin in the training set (651 of 672). By file name there was not a single overlap.
Then came the fastest and dullest part. The main model is left alone and a small add-on is trained beside it — 2 % of all parameters. An hour and a half on one home GPU; the add-on weighs 126 MB and fits in a chat message. On the exam: 20.1 % instead of 34.6 %. And then the errors went down again with no training at all, purely by tuning how the model picks words as it answers: 17.0 %, without a single weight changing.
Half of that last gain came from three clips out of 893. Not because the model misheard them, but
because it looped on them: once greedy decoding turns down the wrong path it cannot correct itself
and starts going in circles — – dedi Akimoviç, – dedi Akimoviç, … until it hits the length cap.
Three clips, 0.34 % of the material, gave 39 % of the whole gain; the other 890 improved steadily and
without tricks. Both are visible below. There is also a story about a setting that was supposed to
cure the looping and instead broke a perfectly ordinary Crimean Tatar repetition. And finally: the
17 % is a studio number. When background speech is as loud as the speech that matters, the error rate
becomes 69 % — the model stops working. That, and not «more data», is the real job of the next
stage.
Examples
Three clips and three different voices. None of these clips was in the training material — that was checked against the training manifests clip by clip, not by file name. All transcripts are copied verbatim from the actual run outputs; nothing is retyped by ear. «In the book» is the human transcript the error rate is measured against.
Example 1
- In the book
- – dep dudaqlarını qaltıratıp qıçırdı. – Tanımayım, – dedi Vedut. – Tanıyım, – dedi tekrar Oksana.
- Base model
- Dev dudaqlarını qaltır atıp qışırdın. Tanımayım, dedi Vedut. Tanıyım, dedi tekrar aqşıdışqan
- Our model, greedy decoding
- – dedi Vedut,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Ak… «– dedi Akimov» repeated 32 times, then cut off by the length cap
- Our model, beam search
- – dep, dudaqlarını qaltıratıp qıçırdı. – Tanımayım... – dedi Vedut. – Tanıyım... – dedi tekrar. One error left: the name Oksana is missing.
Example 2
- In the recording
- - Angisi aceba? - mısqılnen tikilip baqtı Velide. - Mına baq, Eska ketti?
- Base model
- Añgi saace ba, mızqılnen tikilip baqtı Velide Mına baq, Eska ketti, 40 % word error rate
- Our model
- – Angisi aceba? – mısqılnen tikilip baqtı Velide. – Mına baq! Eska ketti... 0 %, not one word wrong
Example 3
- In the recording
- Men, yalanayaq, bir uzun, divarları suvuq qoyum avutiske boyalangan koridornıñ ucunda turam.
- Base model
- Ben yağanayaq bir uzun, divarları sıvuq qoyun avütüske boylanğan koridornıñ ucunda tura 58 % word error rate
- Our model
- Men, yalanayaq, bir uzun, divarları suvuq qoyum avutüske boyalangan koridornıñ ucunda turam. 8 %, one word of 12
The ban that made things worse
There is a blunt setting against looping: forbid the model to repeat itself. On the selection set it broke 17 clips and fixed one. Here is why. A child comes running and calls «Qartbaba, qartbaba!» The model heard that perfectly well — and the ban would not let it write down what it heard, so it corrupted the second word just to keep it different from the first.
«Qartbaba, qartbaba!»
- In the book
- – Qartbaba, qartbaba, – tez-tez çapıp kelgen kiçkene Aziz
- No ban
- – Qartbaba! Qartbaba! – tez-tez çapıp kelgen kiçkene Aziz Heard correctly.
- With the ban
- – Qartbaba! Qartbabaa! – tez-tez çapıp kelgen kiçkene Aziz
- With a stricter ban
- – Qartbaba! Qartbava! – tez-tez çapıp kelgen kiçkene Aziz The second «qartbaba» has turned into a word that does not exist.
This is my favourite story of the whole stage, because it is not about engineering. A language cannot be fixed by prohibitions: any rule of the form «that does not happen» will sooner or later hit the way people actually speak.
Read the full story
- The full article on ana-yurt.com (in Russian)
- Хабр (RU)
- Medium (EN) — link to follow
- dev.to (EN) — link to follow
- Hugging Face blog (EN) — link to follow
- LinkedIn (EN) — link to follow