How to transcribe ESP32-RTSPServer recordings
ESP32-RTSPServer makes the recording — Whipscribe makes it text. Three steps, and the first preview is free with no signup.
The three steps
Get the audio out. Take the source media you're captioning in ESP32-RTSPServer and upload it here — you get SRT/VTT with word-level timing to load back into your tool.
Drop it below. Upload the file right here — no signup for the instant preview. Long, multi-hour files are fine, and video uploads work too (we transcribe the audio track).
Take the text with you. Accurate, speaker-labeled text with timestamps — export TXT, SRT, VTT or DOCX, in 100+ languages, processed on Whipscribe's own private cloud.
About ESP32-RTSPServer
“Esp32 Multiple Client RTSP Server with Video, Audio & Subtitles.”
ESP32-RTSPServer facts & alternatives → · GitHub ↗
Frequently asked
Can Whipscribe transcribe audio you're captioning with ESP32-RTSPServer?
Yes — any audio or video file ESP32-RTSPServer produces can be uploaded directly. Take the source media you're captioning in ESP32-RTSPServer and upload it here — you get SRT/VTT with word-level timing to load back into your tool.
Do I need to convert the file first?
No. Common audio and video formats upload as-is; video's audio track is transcribed automatically.
How accurate is it?
Clear speech in major languages comes back near-publishable; noisy or heavily accented audio deserves a review pass. The instant preview shows real output before you commit.
What does it cost?
About $2 per audio hour as pay-as-you-go credits that never expire — the first preview is free with no signup.