Tiếng Việt · phụ đề SRT và VTT

Vietnamese phụ đề with every tone mark in place.

Upload the video or paste a link and download Vietnamese SRT and VTT files, cued from word-level timings. Vietnamese speech comes back as Vietnamese cue text. For the plain transcript, see the Vietnamese transcript page. Packs from $4, credits never expire;

Tone marks kept whole Cue timings per word SRT and VTT both 10 hours per file

Three steps

Getting a Vietnamese caption file.

All in the browser, with timings handled for you.

1

Send the recording

Upload video or audio, or paste a link. Files too big for the box carry on at the upload page, up to 10 hours or 5 GB.

2

Check the text

Each word has its own timing, so navigating is instant. Place names such as Đà Nẵng, Cần Thơ and Quảng Ngãi deserve a look.

3

Export the file

SRT for editors and platform uploads, VTT for browser players, with TXT, DOCX, JSON and MD from the same run.

Captioning Vietnamese

Stacked marks, and words made of separate syllables.

Vietnamese stacks two kinds of mark on one vowel: the letter-shaping marks that make ă, â, ê, ô, ơ and ư, and the tone marks that sit above or below. A single vowel can carry both, as in ưở or ệ. The file has to be UTF-8 in composed form for these to render as one character; a chain that normalises the text differently, or an editor that strips marks it does not recognise, produces vowels with the tone visibly floating off to one side.

The spacing convention is the trap for line breaks. Vietnamese writes each syllable separately, so a single word is often two or three space-separated pieces — sinh viên, Hà Nội, phát triển. A renderer wrapping at any space can split those, leaving half a word at the end of a line. Break between words rather than between the syllables of one.

SRT and VTT downloads are cue text and timings, so no line arrives labelled with a name. Voices are separated in the transcript view, and the JSON export from the same video job keeps that as data.

Delivery

Where the Vietnamese file is used.

One export covers platform tracks and burned-in text.

YouTube

Add the SRT as the Vietnamese track in Studio. A reviewed file beats automatic captions, especially on names carrying tone marks.

TikTok and Facebook

Both burn text into the picture. Import the SRT into your editor and pick a font that renders stacked diacritics without clipping them.

Premiere and Resolve

SRT lands on a caption track with the timings intact; confirm the title font has full Vietnamese coverage before the master render.

Web players

VTT loads through an HTML5 track element. For video already online, work from its link instead of re-uploading.

Pricing

What Vietnamese captioning costs

A summary is free and fully readable at any length, caption files included. After that: $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes, bought once, credits without an expiry.

Vietnamese video$8 pack$12 pack$24 pack
12-minute hướng dẫn$0.10$0.07$0.06
50-minute podcast episode$0.40$0.30$0.24
120-minute livestream replay$0.96$0.72$0.58

Each figure is the portion of one pack the recording uses. Every format is included; the pricing page has the detail.

FAQ

Vietnamese subtitles, answered.

Which language comes back in the cue file?

Vietnamese. Cue text is written in the language that was spoken, so Vietnamese audio produces Vietnamese captions with no conversion into another language.

Why do tone marks sometimes appear displaced in a player?

The text was normalised into a decomposed form, or the font lacks proper mark placement. Keep the downloaded UTF-8 file as it is and test the font in the real player.

Can a line break between Hà and Nội?

It should not. Vietnamese writes one word as separate syllables, so a break there splits the word; wrap between words instead.

Is Đ treated as a separate letter?

Yes. Đ is its own letter, not a decorated D, and it is written that way in the cue text.

Are northern and southern accents handled differently?

There is no setting to choose. All regional speech goes through the same pass and is written as it was said.

Do caption downloads carry speaker names?

No. SRT and VTT contain cue text and timings; speaker separation is shown in the transcript view and stored in the JSON export.

Upload the video. Take the phụ đề.

Instant preview, no account needed. Your first recording, up to 3 hours, is free.