What Title II and Title III ask for, where WCAG comes in, and why unedited automatic captions rarely carry a compliance claim. Not legal advice.
Start here
The Act requires that people with disabilities can use what you offer. On video that means captions: a viewer who cannot hear gets nothing from an audio track. How specific it gets depends on which part covers you.
Title II (state and local government) has a standard in regulation: the web rule published 2024-04-24 requires web content and mobile apps to meet WCAG 2.1 Level AA, which covers captions on prerecorded video and live captioning at AA. Under Title III (businesses open to the public), the Department of Justice says it has no regulation setting out detailed standards, while holding that the ADA reaches what is offered on the web; its guidance describes accessible video as carrying synchronized captions that are accurate and identify any speakers. Read on ada.gov, 2026-09-15.
The dates and the standard
| Who you are | What applies |
|---|---|
| Government, population 50,000 or more | Title II: WCAG 2.1 Level AA for web and mobile apps, by 2027-04-26 |
| Government, population under 50,000 | Title II: the same standard, by 2028-04-26 |
| Special district governments | Title II: the same standard, by 2028-04-26 |
| Business open to the public | Title III: no detailed standards in regulation, but the ADA still reaches what you offer online |
| Everyone, in practice | WCAG 2.1 Level AA is what policies and procurement tend to name |
The Title II dates were extended by an Interim Final Rule published 2026-04-20. The criteria sit on our WCAG caption requirements page.
The distinction that trips people up
Time-coded text appearing with the picture, in step with the speech, including non-speech audio that carries meaning. This is what an SRT or VTT file holds.
The full text, read apart from the media. The right answer for audio-only material — an addition to captions on video, never a replacement.
Captions serve the viewer watching, the transcript the reader skimming.
The uncomfortable part
Machine captioning is good now, and still wrong where it matters: proper names, acronyms, and any moment two people talk over each other.
We make the files; you make the compliance decision. Upload a video or paste a link for an editable transcript plus SRT and VTT, at 10 hours or 5 GB per file, packs from $4 in full at any length — see the caption generator or YouTube to SRT. One precision: speaker labels appear in the transcript view, and of the downloads only JSON carries them, so SRT, VTT, TXT and DOCX arrive as plain text and markers are added during review. Details on the speaker labels page.
Questions people ask first
It depends which title covers you. The Department of Justice web rule of 2024-04-24 requires state and local government web content and mobile apps to meet WCAG 2.1 Level AA, and captions sit inside that standard. For businesses, the Department says it has no regulation setting out detailed standards, while holding that the ADA reaches what is offered online. Not legal advice.
Two dates, set by population. An Interim Final Rule published 2026-04-20 in the Federal Register moved them: entities with a population of 50,000 or more comply by 2027-04-26, those under 50,000 and special district governments by 2028-04-26. They shifted once, so check ada.gov.
For video, a transcript is not a substitute. Captions are synchronised with the picture, so someone who cannot hear follows as it plays. A transcript is read separately, which is why it answers a different need: the text alternative for audio-only material such as a podcast. Publish both.
Rarely, without review. Machine captions are a draft: names, jargon and crosstalk are where they fail, and they seldom mark who is speaking or the non-speech sounds that carry meaning. Department of Justice guidance describes accessible video as carrying synchronized captions that are accurate and identify any speakers — an unchecked file often meets neither half.
Four things. Accuracy: the words match what was said, names included, punctuated so the meaning survives. Synchronisation: text appears with the speech, not ahead or behind. Completeness: the whole recording, including non-speech audio that matters. Speaker identification: a viewer can tell who is talking when the picture does not show it.
No, and treat any vendor who claims to with suspicion. We produce the caption files, SRT and VTT, plus a transcript you can edit. Whether a video meets a particular obligation is the publisher's decision, usually with qualified advice. We make the artefacts, not the determination.
SRT and VTT from any video, with an editable transcript. packs from $4 in full, at any length.
Try Whipscribe →