What it found on a real book.
A fourteen-hour non-fiction title, proofed by BookMyProof and, separately, by a professional audiobook proofer working the traditional way. Then we compared the two.
The title, author, narrator and publisher are not named, and neither are the amendments themselves. The findings quote proper nouns from the manuscript, and printing those would identify the book in one search. A case study that breaks the confidentiality it advertises is worth less than none. A full worked report is published on an invented title instead.
The book
| 13.8 hours | of finished audio |
| 34 tracks | as delivered, unsplit |
| 824,415 | characters of manuscript |
| minutes | from upload to report |
From 1,814 differences to 15 lines
This is the whole product in one table. Finding differences is easy and nearly useless. Deciding which ones deserve a person is the work.
| Count | What it is | Why it matters |
|---|---|---|
| 1,814 | raw differences between the audio and the book | What a tool that reports everything would have handed over. |
| 451 | of those were text, the rest audio events | Mouth noises, long pauses and clipping, rolled into per-track summaries. |
| 15 | amendments in the finished report | One page. What a person should actually spend a minute on. |
| 89 | pronunciations, listed once each | A name read the same way ninety times is one line, not ninety. |
Against the human
The proofer read the whole book against the audio and returned four faults worth fixing. That is what a fourteen-hour production actually contains: not four hundred problems, four.
BookMyProof found three of those four, at her exact timestamps, in a report she had not seen. It missed one.
We are not going to dress that up. Three of four is not everything, and a tool that told you otherwise would be lying. What it did do was hear all 13.8 hours in minutes, which is the part that costs a person a working day and their concentration by hour nine.
Where it is still wrong
Of the fifteen lines in the report, about thirteen are genuine. Two are artefacts of reading the PDF rather than faults in the narration: a piece of front matter the narrator correctly never read, and a place where the manuscript extractor dropped a space between two words.
A proofer ticks past those in a few seconds, but each one costs a little trust, so they are being fixed rather than explained away. Three similar classes were removed while this study was being written: a spoken height written by the transcriber as a numeral, an initialism spelled out in the book and said as letters, and a chapter heading the PDF glued onto the end of the previous paragraph.
What this is not
One book. That is a promising result and a single data point, and anyone presenting it as proof of a 99% detection rate would be overselling it. It is why the offer is a free pilot on your own title rather than a subscription: you compare it against your own proofer’s sheet, on a book you know, and decide from that.