Automatic speech recognition has become genuinely impressive. Feed a clean recording of a single speaker into a modern engine and the transcript comes back in seconds, mostly correct, formatted, timestamped. Which is exactly why so many organisations now assume closed captioning is a solved problem and quietly stop paying anyone to check the output.

Then a deaf viewer watches a company town hall where the chief executive appears to announce redundancies for the entire quality team, because the engine heard quality where the speaker said qualified. Nobody notices for a week. This is the state of the art, and it is worth understanding precisely where it holds and where it does not.

Closed Captioning Was Never Just a Transcript

The most common misunderstanding is treating captions as text under a video. Proper closed captioning encodes information a hearing viewer receives from the audio track without noticing. Who is speaking when two voices overlap. That a door slammed offscreen. That the music turned ominous. That someone is speaking in a different language and the tone is sarcastic.

An automatic engine produces none of this. It produces words. Everything that distinguishes a caption from a transcript, the speaker identification, the non speech audio, the reading speed management, comes from a human decision about what a viewer needs in order to follow the story.

The Accuracy Number Is Misleading

Vendors quote word error rates around five percent and it sounds excellent. On a two thousand word presentation that is one hundred wrong words, and the errors are not evenly distributed. Engines fail hardest on proper nouns, technical vocabulary, numbers, and accents underrepresented in training data.

Which means the words most likely to be wrong are the ones carrying the most information. A drug name, a dosage, a company name, a legal term. A caption that renders ordinary connective tissue perfectly while mangling every specialist term is worse than useless, because it reads fluently enough that nobody suspects it.

Accents, Overlap and the Real World

Test recordings are clean. Actual meetings are not. Two people talking at once, someone dialling in from a car, a Nigerian English speaker and a Glaswegian on the same call, a whiteboard being tapped. Recognition accuracy on genuinely messy audio remains far below the marketing figures, and it degrades in exactly the situations where accurate captions matter most.

Live captioning adds latency to the problem. An engine producing text two seconds behind the speaker is usable. One that revises its own output as it goes, rewriting words a viewer has already read, produces a disorientating experience that many deaf and hard of hearing users describe as worse than nothing.

Regulators Are Not Impressed by Automation

Compliance obligations do not soften because the captions were generated by software. American broadcast rules assessed by the Federal Communications Commission require captions to be accurate, synchronous, complete and properly placed, and automatic output routinely fails at least two of those four. European accessibility legislation and public sector web rules set comparable expectations.

The practical consequence is that an organisation running unreviewed automatic captions on regulated content is carrying a compliance risk it has usually not measured. The cost of a human review pass is small. The cost of a complaint that leads to an audit of every video you have published is not.

Where the Machines Genuinely Earn Their Place

None of this is an argument for going back to typing from scratch. The sensible workflow uses recognition as a first draft and a human as an editor, which cuts caption production time substantially without surrendering quality. Feed the engine a speaker glossary in advance, because most systems accept custom vocabulary and product names stop being mangled the moment you supply them.

The same hybrid logic runs through the rest of the field. Teams that get good results from subtitling services tend to use automation for timing and alignment while keeping humans on translation, condensation and cultural judgement. Anyone comparing the underlying options will find that audio and video transcription follows the same pattern, since the transcript is the foundation everything downstream inherits.

Captions and Subtitles Are Not the Same Thing

Worth settling, because procurement teams order the wrong one constantly. The closed captioning vs subtitles distinction comes down to assumed audience. Subtitles assume you can hear the audio but do not understand the language, so they carry dialogue only. Captions assume you cannot hear at all, so they carry dialogue plus every other meaningful sound.

Ordering subtitles when accessibility law requires captions is a common and expensive mistake, and automatic tools blur the line further by labelling everything they produce as captions regardless of what it actually contains.

What a Sensible Policy Looks Like

Sort your video library by consequence. Marketing clips and internal recordings can run on automatic captions with occasional spot checks. Training material, regulated communications, anything customer facing and anything that will be watched more than a few hundred times deserves a human pass.

Build a glossary once and reuse it, because most organisations say the same fifty specialist words in every recording. Then ask the people who actually rely on captions whether yours are usable, which is a question remarkably few teams think to ask. Automatic closed captioning has closed most of the gap, and the remaining fraction is where all the trust lives.