Skip to main content

How to Transcribe a Long Recording (2+ Hours)

Author: Published: Last reviewed:

A long recording makes every limit visible: maximum duration, browser preparation, upload reliability, minute balance, and queue safeguards. If you have a 3-hour interview, a half-day workshop, or a board meeting to turn into text, check those boundaries before the upload rather than after a failed job.

Hushscript takes uploads up to 10 hours each, with no file-size limit – large files are prepared right in your browser. There is no plan-based daily allowance; a high 40-hour rolling-24-hour safety ceiling and active-job safeguards protect the service. A long file isn’t a special case here; it’s the same workflow as a two-minute clip, just left to finish on its own.

What you need before you start

You need three things, and you probably already have all three.

A quick note on the free minutes, because long recordings burn through them fast: you get 30 free minutes once. The quickest way to unlock them is to add a card (a $1 hold validates it and is released right away, never charged), and the minutes land immediately. If you’d rather use another payment method available in your country, the 30 minutes arrive with your first purchase instead. A card isn’t required; it’s just the fastest route. Thirty minutes is enough to transcribe a short segment and judge the accuracy before you commit a full pack to a three-hour file.

What to verify before a long upload

Four checks prevent most surprises:

Confirm those boundaries before you upload. Hushscript states them here so a long job can be planned from the source duration rather than discovered by trial and error.

Transcribe a long recording, step by step

There’s no long-recording mode to switch on. The process is identical to a short clip; it just runs longer in the background.

  1. Open the converter and drop your file. Go to audio to text and drag the recording onto the upload area, or click to browse. Any common audio or video format is accepted. If it’s a video, the audio is extracted in your browser at this point, so the heavy video file never leaves your machine.
  2. Check the 30-second preview. Before you sign up, Hushscript prepares a short clip in your browser, sends that clip for transcription, and shows you the speaker-labeled output. The full recording stays local at this stage. Read the preview, confirm the speakers are split sensibly and the words are right, and you will know whether the audio is clean enough for the full run.
  3. Sign in to transcribe the rest. If the preview looks right, sign up and the full file uploads. This is where your minute balance matters: a 150-minute file spends 150 minutes. Make sure the pack you have covers the duration.
  4. Let the job finish on its own. Processing time varies with file length and service load, but the job runs unattended. You do not have to keep the tab in focus; come back when it is done.
  5. Relabel the speakers. The transcript arrives with generic labels like Speaker A and Speaker B. In the editor, rename Speaker A to a real name once and it updates everywhere in the document, so a three-hour interview reads with the right names throughout.
  6. Export in the format you need. Download in any of 21 formats (TXT, SRT, DOCX, PDF, and subtitle and editor-timeline formats among them) plus a password-protected .husharchive, with no watermark on any of them. SRT is the one to reach for on long files, because its timestamps let you jump straight to any moment instead of scrolling.

The result sits in your dashboard with speaker labels, timestamps, and the full export menu, and it stays there for you to come back to.

A worked example: a 2.5-hour recorded interview

Here’s how that looks with a real shape of file. Say you’ve recorded a 2-hour-30-minute interview as a single MP3: two people, one quiet, one loud, recorded over a video call.

You drop the MP3 on the upload area. The 30-second preview shows two speakers already separated in the opening sample. You sign in; the 150-minute file uploads and the job starts. When it completes, the result is one continuous document with timestamps running unbroken from 00:00:00 to 02:30:00.

The raw output reads like this:

[00:00:04] Speaker A: Thanks for making the time. Can we start with how the project began?
[00:00:11] Speaker B: Of course. So it really started back in early 2023, when...

You rename Speaker A to the interviewer and Speaker B to the subject once, and every line updates. Then you export twice: a TXT to paste into a summarizer for a first-pass digest, and an SRT so that when you write up a direct quote you can click the timestamp and hear exactly how it was said. Total cost: 150 minutes off your balance, one upload, no splitting, no truncation, no second attempt.

Keep speakers labeled across a long session

Diarization on a 3-hour recording is genuinely harder than on a 10-minute clip: voices drift, pauses stretch out, and background noise comes and goes. Hushscript runs the labeling over the entire transcript at once, so the labels stay consistent end to end, but the recording itself can help or hurt that.

At recording time:

Audio quality:

Once the transcript is back, relabeling Speaker A with a real name is a one-time edit that propagates through the whole document. For a deeper look at how the labeling itself works, see speaker identification.

Split the file, or upload it whole?

The short answer for almost everyone: upload it whole. Splitting a long recording into parts is the thing to avoid, not the thing to do, and there are only two situations where it’s necessary.

Upload as one file when the recording is 10 hours or under, which covers the overwhelming majority of interviews, lectures, meetings, and panels – there’s no file-size ceiling to work around. One file means one continuous set of timestamps and one consistent set of speaker labels. Splitting breaks both: each part restarts its clock at zero, and the same person can get a different label in part two than in part one, leaving you to stitch and re-map by hand.

Speech systems may still divide long audio internally. Recent IWSLT 2026 research on long-form speech processing compared several segmentation strategies and found that the choice affected robustness. That is an engine-level concern, not a reason to cut your source by hand: keeping one upload preserves a continuous user-facing transcript while the processing pipeline handles its own internal boundaries.

Split only when the recording is genuinely over 10 hours, like a multi-day conference recording. In that case, cut on a natural silence (a session break) rather than mid-sentence, and keep a note of the offset so you can renumber timestamps afterward. If you’re cutting an MP3 before upload, prefer constant bitrate (CBR) over variable bitrate (VBR), since VBR can introduce sync drift in some editors.

A multi-gigabyte video file is not a problem. Because Hushscript extracts the audio in your browser before uploading, the source video’s size is not the upload size, and there’s no cap on how large that source file can be – only the much smaller audio stream travels, as long as your browser and device can handle preparing it.

Best settings for accuracy on long audio

The single biggest factor in accuracy is speech clarity, not sample rate or file format. For long sessions specifically:

Troubleshooting common long-file problems

Accuracy dips in the back half of the recording. This is almost always falling audio levels or rising room noise late in a long session, not a length limit. Check the original recording around the timestamps where errors cluster; if the level dropped, normalize and re-run, and use the 30-second preview on a later segment to confirm before spending the minutes again.

Two people keep getting merged into one speaker. Overlapping speech and crosstalk are the usual cause. There’s no perfect fix after the fact, but a recording with more separation between voices (a per-channel track, or mics further apart) diarizes far more cleanly next time. In the transcript you have now, you can correct the occasional mislabel manually in the editor.

A large file is slow to upload. Upload speed is your connection, not the tool. For a multi-gigabyte video, remember the audio is extracted locally first, so the actual upload is much smaller than the file on disk. If you’re on a weak connection, a wired link or simply leaving it to run will get there.

Background noise throughout the recording. Steady noise (air conditioning, traffic, room hum) lowers accuracy across the whole file. Light volume normalization usually helps more than denoising, which tends to damage speech. If a segment is badly affected, the timestamps still let you find it and listen back.

The transcript looks like it ended early. On Hushscript a transcript covers the full file; there’s no silent truncation. If the text seems short, check the final timestamp against the recording’s real length. Long stretches of silence or music simply produce few words, which can read as “missing” but isn’t.

Export a long transcript: which format

For a multi-hour session, the format you pick changes how usable the transcript is.

Every export is included on every plan, with no watermark and no separate fee to unlock subtitles.


For long recordings billed by the minute, pay-as-you-go transcription avoids a recurring monthly plan. You pay for the hours you transcribe, with a high rolling safety ceiling and active-job controls for unusually heavy batches. The two most common long-recording jobs have their own walkthroughs: how to transcribe a podcast covers multi-guest episodes and show notes, and how to transcribe a lecture covers classroom audio quirks like reverb and audience questions from across a room. And if your material is a stack of separate files rather than one long recording, transcribing a folder of recordings covers running them through as a batch.

Independent sources and standards

Hushscript consulted these independent, non-competing references. They explain research, standards, or platform behavior and do not endorse Hushscript.

Sources reviewed:

Frequently asked questions

Is there a file size limit for long recordings?

No. Recordings can run up to 10 hours long, with no file-size limit – large files are prepared right in your browser. There is no plan-based daily or monthly allowance, but a high 40-hour rolling-24-hour safety ceiling and active-job safeguards apply.

Do I have to split a long recording into parts?

No. A 3-hour or 8-hour file goes up as one upload and comes back as one transcript with continuous timestamps. You only need to split if the recording itself runs over 10 hours.

Will speaker labels stay consistent across a 3-hour recording?

Yes. Speaker diarization runs over the whole transcript at once, not chunk by chunk, so each person keeps the same label from the first minute to the last.

Can I upload a long video file instead of audio?

Yes. The audio is extracted from the video in your browser before anything is sent, so only the audio reaches the server. The original video never uploads, which usually turns a multi-gigabyte lecture capture into a much smaller audio upload.

How long does a long file take to transcribe?

Processing time varies with recording length and service load. Long jobs run unattended, so you can leave the page and return when the transcript is complete.

What export formats work best for a long transcript?

TXT for feeding into other tools, SRT for jumping to a timestamp, DOCX for sharing and annotation, and JSON for the raw speaker-and-timing data. All export formats are included with no watermark.

Do my minutes expire if I only transcribe occasionally?

Minutes expire after 180 days without a completed transcription or minute purchase. Either activity resets the window, and we email you before expiry.

Start with 30 free minutes

Start – 30 free minutes