What does proof-listening find?
Places where the recording and the manuscript differ: a word read differently, a word or line not heard, and a word added. Each comes with its time in the recording and a button to hear it.
Proofing a recorded chapter
Give this page a recorded chapter and the text it was read from. A speech recogniser transcribes the recording, the transcript is lined up with the manuscript word by word, and you get every place where they differ — a skipped line, an added word, a misread — with its time in the recording and a button to hear it. The model runs in your browser; nothing is uploaded.
Ready.
Read this before trusting the list
The transcript comes from Whisper, a speech-recognition model, and it mishears: an unusual name, a quiet word, a sentence over music. When it does, the difference it produces looks the same as a narrator’s slip. So nothing here says the narrator was wrong. Each line is a timecode and a Hear it button; your ears decide.
When the written word and the heard word sound alike — their and there, a name and its nearest dictionary word — the line says probably the recogniser. That is a sound-alike test, not a confidence score: the model this page runs does not report one per word.
Capitals, punctuation and apostrophes are ignored, and numbers are compared as words, so a manuscript that says seven and a transcript that writes 7 agree. Timecodes are placed inside the phrase the recogniser returned, so treat them as within a second or two.
Whisper tiny.en is about 44 MB and its engine about 14 MB, both from this site, the first time you press Compare. English only. A chapter of twenty minutes takes a few minutes on a recent laptop, and the page shows how far it has got.
The rest of the check
The chapter checker measures level, peaks, noise and format and fixes what it can.
A pronunciation guide prevents the misreads this page finds.
Straight answers
Places where the recording and the manuscript differ: a word read differently, a word or line not heard, and a word added. Each comes with its time in the recording and a button to hear it.
No, and the page says so. The transcript comes from a speech-recognition model, and a word it mishears looks exactly like a narrator's misread. Every line is a place to listen, not a verdict. Differences where the two words sound alike are marked as probably the recogniser.
No. The recording and the manuscript are read in this browser tab, and the speech recogniser is downloaded from this site and run on your own device.
WAV, AIFF, MP3 and M4A up to an hour, in English: the recogniser, Whisper tiny.en, reads English only.