Skip to main content

Convert audio to text, speaker-labeled and private

Hushscript is an audio to text converter that transcribes, separates every speaker, and hands back an editable transcript in 21 export formats. Drop a file – or merge a few into one – preview the first 30 seconds free, then sign up for 30 free minutes to convert the rest. Pay-as-you-go after that, no subscription.

Drop an audio or video file to preview a transcript

First 30 seconds · free · no account · your file stays in your browser

Choose a file

Or paste a direct audio link – Dropbox, Google Drive and OneDrive share links work too. How link import works

How it works

  1. 01

    Drop your file

    Use a common audio or video format up to 10 hours, with no file-size limit – large video and audio are prepared right in your browser. Drop one recording or merge related files.

  2. 02

    Preview the first 30 seconds

    See the speaker-labeled output before you sign up. No account needed for the preview.

  3. 03

    Sign up to transcribe it all

    Add a payment method for 30 free minutes – instant when you validate a card (a $1 hold, released right away), or with your first purchase using another method. Then pay-as-you-go.

Most transcription tools hand back one unbroken block of text. Hushscript converts audio to text and separates each speaker as it goes, so an interview or meeting reads like the conversation it actually was, not a wall of text you have to untangle yourself. Rename “Speaker 1” to a real name in a click, fix a misheard word in the editor, and export in whichever of 21 formats your next step needs. It’s a full audio to text converter, not just a transcript generator: convert, clean up, and ship.

How to convert audio to text

Drop your file, preview the first 30 seconds free, then sign up to transcribe the rest. Compatible audio uploads as-is. For video, Hushscript copies out a compatible audio track without re-encoding and converts locally only when required, so the video itself never uploads. Recordings can run up to 10 hours, with no file-size limit – large video and audio are prepared right in the browser. A handful of options sharpen the result before you start. Pin the source language instead of trusting auto-detect, or leave it on for mixed or unpredictable audio. Turn off speaker detection for a single-voice recording like a dictation or voice memo, or turn on multichannel if your file already separates voices onto their own audio channels, common with call recordings, so each channel transcribes as its own speaker. Add up to 1,000 keyterms (names, brands, acronyms) to bias recognition toward the words you actually said, or write a custom prompt of up to 2,000 characters for context the audio alone can’t give the engine. Recording a clinical interview? Medical mode switches to a model built for clinical language; it doubles the billed minutes for that file, and only when you turn it on.

Teach it your vocabulary

Keyterms and a custom prompt work once, for a single file. Saved dictionaries and prompt presets work every time you transcribe. Build a dictionary of the names, products, and jargon specific to your work, import it in bulk from a CSV or TSV file, and organize entries into categories with spelling variants attached to each canonical term. Save a prompt preset the same way, so the context you’d otherwise retype for every file is there by default. Mark a dictionary or preset as your default and it’s preselected automatically the next time you transcribe, so a standing setup doesn’t need reselecting every time. Both attach to a job alongside any one-off keyterms or prompt you add for that file alone, so a standing glossary and a same-day addition work together. They shape what the speech engine hears. Insights run later against the resulting transcript, so better vocabulary improves the source text but does not configure the Insights job.

Fix it fast

A transcript rarely comes back perfect, so the editor is built for quick correction, not a full re-transcription. Undo any change, find and replace a misheard word across the whole transcript at once, rename a speaker so “Speaker 1” becomes a name, tag sections for later, or revert to the original if an edit goes wrong. When you choose Generate Insights, one optional free run creates the complete revision-bound result: summaries, actions, decisions, open questions, quotes, chapters, speaker topics, subject clusters, export suggestions, terminology and profanity findings, and cleanup. Review cleanup in Edit, restore any suggestion you do not want, and apply only the changes you approve.

Export anywhere

Download your transcript as TXT, SRT, VTT, CSV, TSV, Markdown, HTML, RTF, XML, JSON, DOCX, PDF, XLSX, ODS, ODT, ASS, TTML, SBV, LRC, FCPXML, or EDL. That’s 21 formats total, every one clean and unwatermarked. Not sure which one your next step needs? This guide walks through picking a transcript format. Need several at once? Prebuilt archive packs bundle the formats a job actually needs: Subtitle pack for caption files, Research pack for text and data formats, Office pack for docs, Editor pack for editors like Final Cut, Premiere, and Resolve, or All formats for everything at once. Every archive ships as a password-protected .husharchive with a manifest and checksums, or as a standard AES-encrypted ZIP if the tool on the other end doesn’t open .husharchive files. Received a .husharchive from someone else? Open, verify, and extract it right in your browser, with no extra software to install.

Transcribe in batches

One file at a time isn’t always the job. Stage several files at once and they sit in a queue until you hit start. Nothing transcribes the moment you drop it, so you can review and adjust options per file first, or start them all together. Related recordings (a split interview, a multi-part call, backup files from a recorder) can be merged into a single ordered transcript instead of staying as separate files you’d have to stitch together yourself. Your balance determines cost; active-job and service safeguards control how much work runs at once. If you’re clearing a whole directory of recordings, here’s how to transcribe a folder of recordings in one pass.

What it costs

Hushscript is pay-as-you-go, not a subscription: buy minutes in a pack and spend them on the projects you have. Packs run $1.99 for 45 minutes, $5.99 for 5 h (300 min), $11.99 for 15 h (900 min), and $19.99 for 30 h (1,800 min). The largest pack works out to about $0.67 an hour, the best rate on the table. Unused minutes carry forward instead of resetting monthly; a completed transcription or minute purchase resets the 180-day inactivity window. Checkout shows the price in your local currency automatically, with no manual conversion and no surprise on your card statement. New accounts get 30 free minutes when a payment method is on file: a card’s temporary $1 hold is released immediately, never charged. From there, pay as you go.

Looking for a free audio to text converter with nothing to buy at all? The free, browser-based tools convert formats and extract audio from video: they run client-side, upload nothing, and need no account. They just don’t transcribe. Turning speech into text is the part that runs on paid, top-tier AI, which is why it’s the part that costs.

Why Hushscript

Private by design

Your video never uploads – the audio is extracted in your browser – and your audio is deleted the moment the transcript is ready.

No subscription

Pay-as-you-go: buy minutes only when you need them. Nothing recurring, nothing to cancel.

Free speaker labels

Every transcript separates who said what, automatically – not a paid add-on.

Frequently asked questions

Which formats can I convert?

Any common audio or video file. Drop it in and we convert and transcribe it. If something isn't supported, email support@hushscript.com and we'll add it.

Is it free?

Hushscript isn't free. We use top-tier AI for transcription, so it's pay-as-you-go. You do get 30 free minutes to try: instantly when you add and validate a card (a $1 hold, released right away, never charged), or with your first purchase if you use another payment method.

Do I need a card?

No. A card isn't required, but it's the quickest route to your 30 free minutes: a $1 hold validates it, then releases right away. With an alternative payment method available in your country, your 30 free minutes arrive with your first purchase.

Do you keep my audio?

No. It's deleted the moment your transcript is ready, and the speech engine retains nothing. You can delete a transcript in one click.

Are my transcripts encrypted?

Yes. Transcripts are encrypted at rest, so if our storage were ever leaked, your words would be unreadable ciphertext rather than readable text.

How long and how large can a file be?

Up to 10 hours long, with no file-size limit – large video and audio are prepared right in your browser. Your prepaid balance and service safeguards apply.

What can I export?

TXT, SRT, VTT, CSV, TSV, Markdown, HTML, RTF, XML, JSON, DOCX, PDF, XLSX, ODS, ODT, ASS, TTML, SBV, LRC, FCPXML, or EDL: 21 formats, plus prebuilt archive packs and a password-protected .husharchive or AES-encrypted ZIP. No watermark on any of them.

Which languages can it transcribe?

It detects the language automatically and transcribes around 99 languages. The Languages page lists every one.

Is there a free audio to text converter?

The free, browser-based tools convert formats and extract audio from video, with nothing uploaded and no account needed, but they don't transcribe speech into text. That part runs on paid AI: 30 free minutes unlocked by card verification (temporary $1 hold, released immediately, never charged), then pay as you go.

What file formats can I convert to text?

Common audio (MP3, M4A, WAV, FLAC, OGG, Opus) and common video (MP4, MOV, WebM, MKV). If it's a video, the audio is extracted in your browser before anything uploads; the video itself never leaves your device. Missing a format you need? Email support@hushscript.com and we'll add it.

Start with 30 free minutes

Start – 30 free minutes