Notes
What it sounds like
Companions to the articles of the series, audio samples and short notes on how the work is going.
-
How the machine learned to hear: the companion to the second article
The exam before the training: 34.6 % of words wrong, four audiobooks that sat in the material twice, an hour and a half of fine-tuning on a home GPU — and 17.0 % at the end. With the transcripts, lined up side by side.
- ASR
- audio
-
From a dictionary to a voice assistant: the audio companion to the first article
Where the first stage of the road ended up: speech recognition from 34.6 % down to 17.0 % word errors, a whole novel — 16 chapters, 15,444 words, 1 hour 58 minutes — read by a synthetic voice, all of it on one consumer GPU. With samples you can listen to.
- TTS
- ASR
- audio