Upload
MP3, WAV, M4A, MP4, WebM, OGG, or FLAC. Up to 120 minutes per file.
Upload audio or video. Get a timestamped transcript ready for captions, editing, or narration.
A production workflow
This example shows the output shape: a transcript with timestamps that can become subtitles or a new narration script.
The same pipeline behind NarrateHQ's studio, applied in reverse.
MP3, WAV, M4A, MP4, WebM, OGG, or FLAC. Up to 120 minutes per file.
A timestamped transcript comes back, viewable as plain text or with timing.
Download as SRT, VTT, or TXT, or send the transcript into the studio as a script.
Use the transcript as a starting point, then keep the rest of the production workflow in one place.
Auto-detect works well for the common case, but selecting your language before uploading gives the most accurate result. Auto-detect can occasionally mix up similar-sounding languages.
The uploaded recording and the transcript follow different retention rules.
MP3, WAV, M4A, MP4, WebM, OGG, and FLAC, up to 120 minutes per file.
English, Thai, Spanish, Portuguese, French, German, Japanese, Korean, Mandarin Chinese, Vietnamese, Indonesian, Hindi, and Arabic.
The transcript includes timestamps from the completed transcription job. You can download the timed SRT or send the text into narration.
When a recording has more than one speaker, segments are labeled automatically (Speaker 1, Speaker 2, and so on). Speaker labels are automatically detected and may not always be accurate.
Only upload audio or video that you have the right and consent to process. The supported formats and per-file duration limit are shown before upload.
Transcription uses a different provider with its own cost per minute, so it's metered separately rather than drawn from the same pool as narration.
It's deleted as soon as the transcript is ready. The transcript itself stays on your account until you delete it. See the Privacy Policy.
Yes. Send it to the studio as a script with one click, then generate audio in any voice.