How to Transcribe a Long Recording (2+ Hours)
Author: Hushscript Published: Last reviewed:
A long recording makes every limit visible: maximum duration, browser preparation, upload reliability, minute balance, and queue safeguards. If you have a 3-hour interview, a half-day workshop, or a board meeting to turn into text, check those boundaries before the upload rather than after a failed job.
Hushscript takes uploads up to 10 hours each, with no file-size limit – large files are prepared right in your browser. There is no plan-based daily allowance; a high 40-hour rolling-24-hour safety ceiling and active-job safeguards protect the service. A long file isn’t a special case here; it’s the same workflow as a two-minute clip, just left to finish on its own.
What you need before you start
You need three things, and you probably already have all three.
- The recording itself, in any common audio or video format: MP3, M4A, WAV, FLAC, AAC, MP4, MOV, MKV, and the rest. There’s no need to convert it first; just drop the file in.
- A modern browser. The first 30 seconds are prepared as a compact preview clip in your browser and sent for preview transcription; the full file is not uploaded until you sign in. For video files, the browser also prepares the audio without sending the source video.
- Enough minutes in your balance to cover the duration. This is the one number worth checking up front. A 2.5-hour interview is 150 minutes, so a 5 h (300 min) pack covers it with room to spare. New accounts get 30 free minutes to try the full flow; see the pricing page for what each pack costs.
A quick note on the free minutes, because long recordings burn through them fast: you get 30 free minutes once. The quickest way to unlock them is to add a card (a $1 hold validates it and is released right away, never charged), and the minutes land immediately. If you’d rather use another payment method available in your country, the 30 minutes arrive with your first purchase instead. A card isn’t required; it’s just the fastest route. Thirty minutes is enough to transcribe a short segment and judge the accuracy before you commit a full pack to a three-hour file.
What to verify before a long upload
Four checks prevent most surprises:
- Per-file duration. Hushscript accepts a recording up to 10 hours. A longer source needs to be split at a natural break.
- Browser and device capacity. There is no product file-size limit, but the browser still has to read and prepare the local source. Very large files need enough memory, storage headroom, and time on the device doing the upload.
- Minute balance. Billing follows audio duration. A 150-minute recording spends 150 base minutes before any option that changes billed duration.
- Service safeguards. There is no plan-based daily allowance, but a 40-hour rolling-24-hour safety ceiling and active-job safeguards still apply. A large batch can queue rather than start every file at once.
Confirm those boundaries before you upload. Hushscript states them here so a long job can be planned from the source duration rather than discovered by trial and error.
Transcribe a long recording, step by step
There’s no long-recording mode to switch on. The process is identical to a short clip; it just runs longer in the background.
- Open the converter and drop your file. Go to audio to text and drag the recording onto the upload area, or click to browse. Any common audio or video format is accepted. If it’s a video, the audio is extracted in your browser at this point, so the heavy video file never leaves your machine.
- Check the 30-second preview. Before you sign up, Hushscript prepares a short clip in your browser, sends that clip for transcription, and shows you the speaker-labeled output. The full recording stays local at this stage. Read the preview, confirm the speakers are split sensibly and the words are right, and you will know whether the audio is clean enough for the full run.
- Sign in to transcribe the rest. If the preview looks right, sign up and the full file uploads. This is where your minute balance matters: a 150-minute file spends 150 minutes. Make sure the pack you have covers the duration.
- Let the job finish on its own. Processing time varies with file length and service load, but the job runs unattended. You do not have to keep the tab in focus; come back when it is done.
- Relabel the speakers. The transcript arrives with generic labels like
Speaker AandSpeaker B. In the editor, renameSpeaker Ato a real name once and it updates everywhere in the document, so a three-hour interview reads with the right names throughout. - Export in the format you need. Download in any of 21 formats (TXT, SRT, DOCX, PDF, and subtitle and editor-timeline formats among them) plus a password-protected .husharchive, with no watermark on any of them. SRT is the one to reach for on long files, because its timestamps let you jump straight to any moment instead of scrolling.
The result sits in your dashboard with speaker labels, timestamps, and the full export menu, and it stays there for you to come back to.
A worked example: a 2.5-hour recorded interview
Here’s how that looks with a real shape of file. Say you’ve recorded a 2-hour-30-minute interview as a single MP3: two people, one quiet, one loud, recorded over a video call.
You drop the MP3 on the upload area. The 30-second preview shows two speakers already separated in the opening sample. You sign in; the 150-minute file uploads and the job starts. When it completes, the result is one continuous document with timestamps running unbroken from 00:00:00 to 02:30:00.
The raw output reads like this:
[00:00:04] Speaker A: Thanks for making the time. Can we start with how the project began?
[00:00:11] Speaker B: Of course. So it really started back in early 2023, when...
You rename Speaker A to the interviewer and Speaker B to the subject once, and every line updates. Then you export twice: a TXT to paste into a summarizer for a first-pass digest, and an SRT so that when you write up a direct quote you can click the timestamp and hear exactly how it was said. Total cost: 150 minutes off your balance, one upload, no splitting, no truncation, no second attempt.
Keep speakers labeled across a long session
Diarization on a 3-hour recording is genuinely harder than on a 10-minute clip: voices drift, pauses stretch out, and background noise comes and goes. Hushscript runs the labeling over the entire transcript at once, so the labels stay consistent end to end, but the recording itself can help or hurt that.
At recording time:
- If your tool can capture a separate track per participant, use it. Zoom’s local recording, for example, can produce either a mixed track or per-channel audio depending on settings; the mixed stereo track usually works best for upload.
- Recording a physical room is the hard case. A directional mic, or simply seating speakers further apart, cuts the crosstalk that makes two voices read as one.
Audio quality:
- 44.1 kHz / 16-bit WAV, or a high-bitrate MP3 at 128 kbps or above, is plenty for the engine. There’s no accuracy gain from 96 kHz; it just makes a long file larger.
- If the recording is quiet or uneven, normalizing the volume in an audio editor before upload often lifts accuracy more than any setting change does.
Once the transcript is back, relabeling Speaker A with a real name is a one-time edit that propagates through the whole document. For a deeper look at how the labeling itself works, see speaker identification.
Split the file, or upload it whole?
The short answer for almost everyone: upload it whole. Splitting a long recording into parts is the thing to avoid, not the thing to do, and there are only two situations where it’s necessary.
Upload as one file when the recording is 10 hours or under, which covers the overwhelming majority of interviews, lectures, meetings, and panels – there’s no file-size ceiling to work around. One file means one continuous set of timestamps and one consistent set of speaker labels. Splitting breaks both: each part restarts its clock at zero, and the same person can get a different label in part two than in part one, leaving you to stitch and re-map by hand.
Speech systems may still divide long audio internally. Recent IWSLT 2026 research on long-form speech processing compared several segmentation strategies and found that the choice affected robustness. That is an engine-level concern, not a reason to cut your source by hand: keeping one upload preserves a continuous user-facing transcript while the processing pipeline handles its own internal boundaries.
Split only when the recording is genuinely over 10 hours, like a multi-day conference recording. In that case, cut on a natural silence (a session break) rather than mid-sentence, and keep a note of the offset so you can renumber timestamps afterward. If you’re cutting an MP3 before upload, prefer constant bitrate (CBR) over variable bitrate (VBR), since VBR can introduce sync drift in some editors.
A multi-gigabyte video file is not a problem. Because Hushscript extracts the audio in your browser before uploading, the source video’s size is not the upload size, and there’s no cap on how large that source file can be – only the much smaller audio stream travels, as long as your browser and device can handle preparing it.
Best settings for accuracy on long audio
The single biggest factor in accuracy is speech clarity, not sample rate or file format. For long sessions specifically:
- Use a compressor on the recording side (a hardware or software audio compressor, not file compression) to even out the gap between a loud interviewer and a quiet subject. This matters far more across three hours than across ten minutes, where one drift in level can lose whole exchanges.
- FLAC is lossless and accepted, but it won’t beat a good MP3 on accuracy. Its only real use is removing doubt: if you want to be certain audio quality isn’t the bottleneck, FLAC settles the question at the cost of a larger file.
- Avoid heavy noise reduction before upload. Aggressive denoising can chew up consonants and actually lower accuracy. Light normalization helps; scrubbing does not.
Troubleshooting common long-file problems
Accuracy dips in the back half of the recording. This is almost always falling audio levels or rising room noise late in a long session, not a length limit. Check the original recording around the timestamps where errors cluster; if the level dropped, normalize and re-run, and use the 30-second preview on a later segment to confirm before spending the minutes again.
Two people keep getting merged into one speaker. Overlapping speech and crosstalk are the usual cause. There’s no perfect fix after the fact, but a recording with more separation between voices (a per-channel track, or mics further apart) diarizes far more cleanly next time. In the transcript you have now, you can correct the occasional mislabel manually in the editor.
A large file is slow to upload. Upload speed is your connection, not the tool. For a multi-gigabyte video, remember the audio is extracted locally first, so the actual upload is much smaller than the file on disk. If you’re on a weak connection, a wired link or simply leaving it to run will get there.
Background noise throughout the recording. Steady noise (air conditioning, traffic, room hum) lowers accuracy across the whole file. Light volume normalization usually helps more than denoising, which tends to damage speech. If a segment is badly affected, the timestamps still let you find it and listen back.
The transcript looks like it ended early. On Hushscript a transcript covers the full file; there’s no silent truncation. If the text seems short, check the final timestamp against the recording’s real length. Long stretches of silence or music simply produce few words, which can read as “missing” but isn’t.
Export a long transcript: which format
For a multi-hour session, the format you pick changes how usable the transcript is.
- TXT is the most portable and the right input for feeding into an LLM to summarize or analyze a long conversation.
- SRT carries utterance-level timestamps, so you can jump to any moment without scrolling. If your reason for transcribing is to find specific moments, this is the format.
- DOCX keeps speaker labels and timestamps in a Word-compatible file, best for sharing a transcript with people who need to comment on it.
- JSON gives you the raw structure (speaker, start, end, text) for piping into another tool.
Every export is included on every plan, with no watermark and no separate fee to unlock subtitles.
For long recordings billed by the minute, pay-as-you-go transcription avoids a recurring monthly plan. You pay for the hours you transcribe, with a high rolling safety ceiling and active-job controls for unusually heavy batches. The two most common long-recording jobs have their own walkthroughs: how to transcribe a podcast covers multi-guest episodes and show notes, and how to transcribe a lecture covers classroom audio quirks like reverb and audience questions from across a room. And if your material is a stack of separate files rather than one long recording, transcribing a folder of recordings covers running them through as a batch.
Independent sources and standards
Hushscript consulted these independent, non-competing references. They explain research, standards, or platform behavior and do not endorse Hushscript.
Sources reviewed: