Note

How the machine learned to hear: the companion to the second article

The exam before the training: 34.6 % of words wrong, four audiobooks that sat in the material twice, an hour and a half of fine-tuning on a home GPU — and 17.0 % at the end. With the transcripts, lined up side by side.

◆ ASR◆ audio

The second article of the series is out on ana-yurt.com, and it is about the ears: speech recognition for Crimean Tatar — a computer listening and writing down what it hears. This page is its companion: the same transcripts as in the article, lined up row by row, so the difference can be seen rather than taken on trust.

It started with a humbling number. The project already had a large recogniser fine-tuned on our language, and it was already in use — but it had never sat a fair exam. The first thing I did in this stage was not to improve anything: I set the exam. Two whole books, two narrators the model had never heard, 893 clips, 1 hour 52 minutes. The result was 34.6 % of words wrong — every third word. And before running anything at all, I compared the training material against the test material not by file name but by text — and found that four audiobooks were sitting in my material twice, cut by two different tools under completely unrelated names. For the very book I was about to test on, 96.9 % of clips had a twin in the training set (651 of 672). By file name there was not a single overlap.

Then came the fastest and dullest part. The main model is left alone and a small add-on is trained beside it — 2 % of all parameters. An hour and a half on one home GPU; the add-on weighs 126 MB and fits in a chat message. On the exam: 20.1 % instead of 34.6 %. And then the errors went down again with no training at all, purely by tuning how the model picks words as it answers: 17.0 %, without a single weight changing.

Half of that last gain came from three clips out of 893. Not because the model misheard them, but because it looped on them: once greedy decoding turns down the wrong path it cannot correct itself and starts going in circles — – dedi Akimoviç, – dedi Akimoviç, … until it hits the length cap. Three clips, 0.34 % of the material, gave 39 % of the whole gain; the other 890 improved steadily and without tricks. Both are visible below. There is also a story about a setting that was supposed to cure the looping and instead broke a perfectly ordinary Crimean Tatar repetition. And finally: the 17 % is a studio number. When background speech is as loud as the speech that matters, the error rate becomes 69 % — the model stops working. That, and not «more data», is the real job of the next stage.

Examples

Three clips and three different voices. None of these clips was in the training material — that was checked against the training manifests clip by clip, not by file name. All transcripts are copied verbatim from the actual run outputs; nothing is retyped by ear. «In the book» is the human transcript the error rate is measured against.

Example 1

In the book
– dep dudaqlarını qaltıratıp qıçırdı. – Tanımayım, – dedi Vedut. – Tanıyım, – dedi tekrar Oksana.
Base model
Dev dudaqlarını qaltır atıp qışırdın. Tanımayım, dedi Vedut. Tanıyım, dedi tekrar aqşıdışqan
Our model, greedy decoding
– dedi Vedut,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Akimov,– dedi Ak… «– dedi Akimov» repeated 32 times, then cut off by the length cap
Our model, beam search
– dep, dudaqlarını qaltıratıp qıçırdı. – Tanımayım... – dedi Vedut. – Tanıyım... – dedi tekrar. One error left: the name Oksana is missing.
The looping clip. Greedy decoding repeated «– dedi Akimov» 32 times in a row and ran into the length cap; beam search brought the sentence back.

Example 2

In the recording
- Angisi aceba? - mısqılnen tikilip baqtı Velide. - Mına baq, Eska ketti?
Base model
Añgi saace ba, mızqılnen tikilip baqtı Velide Mına baq, Eska ketti, 40 % word error rate
Our model
– Angisi aceba? – mısqılnen tikilip baqtı Velide. – Mına baq! Eska ketti... 0 %, not one word wrong
A female voice, a different reader: 9.1 s of live dialogue. The base model crumbled on the opening words and lost the dialogue punctuation entirely; today's model wrote the line out in full and put the dashes back.

Example 3

In the recording
Men, yalanayaq, bir uzun, divarları suvuq qoyum avutiske boyalangan koridornıñ ucunda turam.
Base model
Ben yağanayaq bir uzun, divarları sıvuq qoyun avütüske boylanğan koridornıñ ucunda tura 58 % word error rate
Our model
Men, yalanayaq, bir uzun, divarları suvuq qoyum avutüske boyalangan koridornıñ ucunda turam. 8 %, one word of 12
A third voice, 7.6 s. The base model began the sentence in Turkish — «Ben» instead of «Men» — and then mangled half the words: 7 errors out of 12. Today's model got one word wrong.

The ban that made things worse

There is a blunt setting against looping: forbid the model to repeat itself. On the selection set it broke 17 clips and fixed one. Here is why. A child comes running and calls «Qartbaba, qartbaba!» The model heard that perfectly well — and the ban would not let it write down what it heard, so it corrupted the second word just to keep it different from the first.

«Qartbaba, qartbaba!»

In the book
– Qartbaba, qartbaba, – tez-tez çapıp kelgen kiçkene Aziz
No ban
– Qartbaba! Qartbaba! – tez-tez çapıp kelgen kiçkene Aziz Heard correctly.
With the ban
– Qartbaba! Qartbabaa! – tez-tez çapıp kelgen kiçkene Aziz
With a stricter ban
– Qartbaba! Qartbava! – tez-tez çapıp kelgen kiçkene Aziz The second «qartbaba» has turned into a word that does not exist.
A doubled vocative is ordinary Crimean Tatar. That is how people speak. The ban cannot tell a living repetition from a machine loop.

This is my favourite story of the whole stage, because it is not about engineering. A language cannot be fixed by prohibitions: any rule of the form «that does not happen» will sooner or later hit the way people actually speak.

Read the full story