How to transcribe › MocapNET

How to transcribe MocapNET recordings

MocapNET makes the recording — Whipscribe makes it text. Three steps, and the first preview is free with no signup.

The three steps

1

Get the audio out. Copy the clip off MocapNET (any common video format) and upload it as-is.

2

Drop it below. Upload the file right here — no signup for the instant preview. Long, multi-hour files are fine, and video uploads work too (we transcribe the audio track).

3

Take the text with you. Accurate, speaker-labeled text with timestamps — export TXT, SRT, VTT or DOCX, in 100+ languages, processed on Whipscribe's own private cloud.

About MocapNET

“A real-time method that estimates the 3D human pose directly in the popular Bio Vision Hierarchy (BVH) format, given estimations of the 2D body joints originating from monocular color images. Our contributions include: (a) A novel and compact 2D pose NSRM representation. (b) A human body orientation.”

GitHub stars
948
Built in
C++
Category
Cameras Capture

MocapNET facts & alternatives →  ·  GitHub ↗  ·  Website ↗

Frequently asked

Can Whipscribe transcribe footage captured with MocapNET?

Yes — any audio or video file MocapNET produces can be uploaded directly. Copy the clip off MocapNET (any common video format) and upload it as-is.

Do I need to convert the file first?

No. Common audio and video formats upload as-is; video's audio track is transcribed automatically.

How accurate is it?

Clear speech in major languages comes back near-publishable; noisy or heavily accented audio deserves a review pass. The instant preview shows real output before you commit.

What does it cost?

About $2 per audio hour as pay-as-you-go credits that never expire — the first preview is free with no signup.