Send the video or a link and download Norwegian SRT and VTT files whose cues follow word-level timings. Norwegian speech returns Norwegian cue text. For the plain transcript, that page is here. Your first recording up to 3 hours is free; $8 then covers 1,000 minutes.
Three steps
In the browser, with the timing handled.
Video or audio in. Files beyond the box continue on the upload page, which takes up to 10 hours or 5 GB.
Each word carries a timing, so navigating is fast. Place names like Trondheim, Ålesund and Tromsø are the ones to confirm.
SRT for editing suites and platform uploads, VTT for browsers, plus TXT, DOCX, JSON and MD.
Captioning Norwegian
Norway has two official written standards, Bokmål and Nynorsk, and no single spoken standard at all. People speak their own dialect in public life, on television and in interviews, and nobody is expected to switch. That is unusual, and it matters for captioning: a cue file has to represent speech that may not match either written standard exactly, and the honest approach is to write what the speaker said rather than to convert it to a house norm.
Practically, that means a video from Bergen, one from Stavanger and one from Oslo will produce cue text that differs in more than accent — pronouns, verb endings and everyday words all vary. If your organisation publishes in one standard, treat the export as the accurate starting point and make the editorial decision yourself in the caption editor rather than expecting the file to choose.
SRT and VTT downloads carry cue text and timings only, so no line begins with a name. Voices are separated in the transcript view, and the JSON export from the same job holds that structure.
Delivery
From one export to a platform, an editor or a player.
Attach the SRT as the Norwegian track in Studio. Uploaded cue files are indexed and replace the automatic captions.
Text is burned into the frame on these platforms, so bring the SRT into your editor and style it over the picture.
SRT lands on a caption track with the timings intact, ready for restyling inside the project.
A browser video element accepts VTT on its track element. For published video, start from the link rather than the master.
Pricing
Packs from $4, credits never expire and fully readable at any length, and the caption exports come with it. Then, one-time packs: $8 for 1,000 minutes, $12 for 2,000 minutes, $24 for 5,000 minutes, credits that do not expire.
| Norwegian video | $8 pack | $12 pack | $24 pack |
|---|---|---|---|
| 15-minute presentasjon | $0.12 | $0.09 | $0.07 |
| 45-minute podkast | $0.36 | $0.27 | $0.22 |
| 120-minute konferanse | $0.96 | $0.72 | $0.58 |
Each number is the share of a single pack the recording uses. No format costs extra; see the pricing page.
FAQ
Norwegian. Cue text is written in the language that was spoken, so Norwegian audio gives Norwegian captions and nothing is converted into another language.
The cue text follows what the speaker said rather than converting to one written standard. If you publish in a specific standard, make that editorial change in your caption editor.
It is written as spoken. Norwegians use their own dialect in public life, so pronouns, verb endings and vocabulary legitimately vary between recordings.
No. The two look similar on the page but differ in everyday vocabulary, so each needs its own cue file.
Yes, they are separate letters written as transcribed, and they survive if the file stays in UTF-8 through your chain.
No. Caption downloads are cue text and timings; the transcript view separates the voices and JSON keeps that record.
Instant preview with no account. The first recording of up to 3 hours costs nothing.