Captions live in Guideline 1.2, Time-based Media. Recorded video is Level A, live is Level AA. The whole family in plain English.
Where captions sit
Guideline 1.2 covers everything with a timeline. Captions are two of its criteria; the rest deal with transcripts, audio description and sign language.
The level surprises people. Captions on recorded video are 1.2.2 Captions (Prerecorded) at Level A — the minimum rung of the standard, not a later refinement. Live captioning is 1.2.4 Captions (Live) at Level AA, the level most policies name. Criteria below are as published in WCAG 2.1 by the W3C, checked 2026-09-15; how they map onto United States law is covered on our captions and ADA compliance page.
Guideline 1.2, Time-based Media
| Success criterion | Level | In plain English |
|---|---|---|
| 1.2.1 Audio-only and Video-only (Prerecorded) | A | A podcast needs a transcript; a silent video needs a text alternative or an audio track |
| 1.2.2 Captions (Prerecorded) | A | Recorded video with sound needs captions — the one most people mean |
| 1.2.3 Audio Description or Media Alternative (Prerecorded) | A | What is shown but not said must reach someone who cannot see it |
| 1.2.4 Captions (Live) | AA | Live streams and broadcasts need captions as they happen |
| 1.2.5 Audio Description (Prerecorded) | AA | At AA, recorded video needs actual audio description |
| 1.2.6 Sign Language (Prerecorded) | AAA | Sign language interpretation of recorded audio |
| 1.2.8 Media Alternative (Prerecorded) | AAA | A full text alternative beyond captions |
| 1.2.9 Audio-only (Live) | AAA | An equivalent text alternative for live audio |
AA, the usual target, means 1.2.1 through 1.2.5: captions everywhere, plus audio description on recorded video.
What counts as a caption
A caption track stands in for the whole audio experience. A viewer relying on it should learn what a hearing viewer learns — that someone laughed, that a phone rang, which of three people just started talking.
Words people use interchangeably
Subtitles assume you can hear. Captions assume you cannot, so they add sound cues and speaker changes. WCAG asks for captions.
Read apart from the media. It answers 1.2.1 for audio-only material, but does not replace captions on a video.
A spoken account of what is visible but unspoken — a different criterion that captions never satisfy.
Whipscribe produces SRT and VTT files and an editable transcript from a video or pasted link, at 10 hours or 5 GB per file, packs from $4 in full at any length. It does not certify conformance. One precision matters here: speaker labels appear in the transcript view on the site, and among the downloads only JSON carries them — SRT, VTT, TXT and DOCX come out as plain continuous text, so speaker markers in captions are added during review. See speaker labels, the caption generator or YouTube to SRT.
Questions people ask first
Two of them. 1.2.2 Captions (Prerecorded) sits at Level A and covers recorded video with sound. 1.2.4 Captions (Live) sits at Level AA and covers live streams. Because 1.2.2 is Level A, captions on recorded video belong to the lowest conformance level there is.
They are conformance levels. A is the floor; AA adds criteria and is what most policies ask for; AAA is highest and is not expected across whole sites. For media: recorded captions at A, live captions and audio description at AA, sign language at AAA.
No, it needs a transcript. Audio with no video is covered by 1.2.1 at Level A, which asks for an alternative presenting equivalent information. Captions are for synchronized media, where text keeps pace with a picture.
Everything a hearing viewer gets from the audio: the dialogue, accurately and in time with the speech, the non-speech sound that carries meaning, and an indication of who is speaking where the picture does not make it obvious. Words alone are an incomplete answer.
Not reliably, and not on their own. The criterion asks for captions of the audio content, which means captions that convey it correctly. Machine output fails on names, acronyms and crosstalk, and marks neither speakers nor meaningful sounds. Treat it as a draft and review it.
It gives you the raw material. No, we do not certify conformance and would not claim to. You get SRT and VTT files and an editable transcript; reviewing them, adding speaker markers and sound cues, and deciding whether the result meets a standard are yours. Anyone promising guaranteed conformance from an automated step is overselling.
SRT and VTT from any video or pasted link, with an editable transcript. Packs from $4, in full, any length.
Try Whipscribe →