Skip to main content

Blog

Research Transcription for Interviews and Focus Groups

Author: Published: Last reviewed:

A recording is not data yet. It becomes data once somebody has decided what gets written down, how speakers are marked, which identifying details come out, and what the file looks like when it reaches your coding software. Those four decisions are research transcription, and every one of them is methodological rather than clerical.

Making them before the first interview is far cheaper than making them at interview fourteen. Here is what each decision changes, what belongs in your methods section, and where automatic transcription genuinely helps a study instead of just making the folder look finished.

Choose the transcript type before the first session

There is no neutral transcript. Every convention makes some things visible and hides others, so choose for the analysis you actually intend to run.

Transcript type What it keeps Fits
Verbatim Fillers, false starts, pauses, overlaps Conversation and discourse analysis
Intelligent verbatim The words, minus fillers and repairs Thematic analysis, grounded theory
Clean Edited for grammar and readability Quotes for publication and reports
Denaturalized Standardized grammar across the corpus Cross-case comparison, large samples

The expensive mistake is not choosing badly, it is choosing twice. A corpus half verbatim and half cleaned cannot be coded consistently, and re-transcribing twenty interviews to repair that is a month nobody has budgeted. Write the convention down in advance, including how you will mark laughter, interruptions, and inaudible stretches.

What machines get right, and what they miss

Automatic transcription is good enough to be the first draft of nearly any study, and not good enough to be the last. Pew Research Center compared human and machine transcription on recorded sermons and found the machines kept pace on clear, continuous speech while falling behind on proper names, specialist terms, and passages where a speaker shifted tone or volume. Those are the passages qualitative researchers tend to quote.

So the working pattern is a draft you did not type, followed by a review pass you did. Read along with the audio once, correct names and terminology, fix speaker attribution across the crosstalk, and stop there. The review is the part that costs real time, and the part that cannot be skipped.

Speaker labels are already pseudonyms

Automatic speaker separation splits a recording by voice and labels the results Speaker 1, Speaker 2, and so on. It recognizes voices, not people, which is quietly useful: the label is a pseudonym by default and stays one unless somebody types a real name into it.

Keep it that way. Rename Speaker 1 to P07 or to a role label once and the change carries through the whole transcript and every export, so a participant ID scheme is applied consistently instead of by hand in forty places. Speaker labels are free on every Hushscript transcript with no cap on the number of voices, which is what a six-person focus group needs.

The honest limit is crosstalk. When participants talk over each other the split needs a human pass, and the timecodes are what tell you which minutes to re-listen to.

Put the transcription method in the paper

Reviewers increasingly ask how the text was produced, not only what it says. The APA’s Journal Article Reporting Standards for qualitative research ask authors to describe how data were recorded and transformed, which makes transcription something you report rather than assume.

One methods paragraph answering these is usually enough:

De-identify while transcribing, not afterwards

Identifying detail arrives mid-sentence: an employer, a ward, a street, a court date, the one colleague who would recognize the story. Stripping it later means reading the entire corpus a second time, so fold it into the transcription pass instead.

Doing that inside the transcript rather than in exported copies also means the pseudonymized version is the only one that ever circulates. Find-and-replace fixes a recurring place name across the document in one action, and saved dictionaries carry the study’s names, places, and technical vocabulary into recognition so fewer of them come back wrong to begin with.

What you can say about the platform matters as much as the redaction. Hushscript deletes the audio the moment each transcript is ready and never uses it to train models; transcripts are encrypted at rest under a retention window you set, from 7 to 365 days with per-transcript overrides; deletion flows return receipts. Audio from the EU, EEA, Switzerland, and the UK is processed in the EU. For a recording that should not leave a stored transcript at all, private transcription hands back a password-protected .husharchive file instead.

Export for the software that receives it

A transcript rarely stays where it was made. It moves into NVivo, ATLAS.ti, MAXQDA, or a spreadsheet, and later into a manuscript. Plain text is where speaker labels and timecodes go to die, so pick the format by its destination.

Hushscript exports 21 formats with no watermark and bundles the relevant ones as a Research pack. Timecodes surviving the export is what lets a quote in your findings point back at the second it was spoken, which is the difference between a citable transcript and a long document.

Multilingual fieldwork without a vendor per country

A comparative study should not need a different transcription supplier in every site. Recognition covers roughly 99 languages with automatic detection, with the strongest accuracy on a flagship set of 18 language variants, so one workflow carries the whole sample.

Translation is a separate decision from transcription, and the order matters: transcribe in the language spoken, then translate. Each translation target costs 25% of the recording duration and puts the original and your working language in tabs of the same document, each exportable on its own, so you quote the participant’s words and code in the language you think in. Report who translated and how disagreements were settled, because two translators will not render the same phrase identically.

Costing it against the grant line

Manual transcription is priced in researcher hours, and twenty hour-long interviews is the sort of number that quietly consumes a term. Automatic transcription moves that into a budget line small enough to stop being a constraint on sample size.

Hushscript sells prepaid minutes rather than a subscription, so fieldwork months cost what they use and writing months cost nothing. On the largest pack, $49.99 for 100 hours, twenty hour-long interviews come to roughly $10 and a 90-minute focus group to about $0.75. Minutes stay valid for 365 days and any transcription or purchase resets that window, which suits a study that pauses for ethics approval. New accounts get 30 free minutes, released either by a $1 card check that is authorized and immediately released and never charged, or by your first purchase if you pay another way.

Where to start

Use the last interview you actually recorded, not a clean test file. Put it through research transcription and read the first 5 minutes that come back: that shows you how speaker identification behaves on your participants, your room, and your accents, which is the only accuracy figure that means anything for your study.

Then write the convention into your protocol while there are still thirty interviews left to apply it to.

Independent sources and standards

Hushscript consulted these independent, non-competing references. They explain research, standards, or platform behavior and do not endorse Hushscript.

Sources reviewed:

Frequently asked questions

Is automatic transcription accurate enough for qualitative research?

As a first draft, yes. As a final transcript, not on its own. Machine recognition keeps up with clear, continuous speech and falls behind on proper names, specialist vocabulary, and the moments a speaker changes register, which are exactly the moments researchers quote. The workable pattern is an automatic draft plus one review pass against the audio.

How do participants stay anonymous in the transcript?

Speaker labels carry no identity. Automatic separation knows voices, not people, so the labels arrive as Speaker 1 and Speaker 2 and stay pseudonyms unless you type a real name into one. Rename them once to your participant ID scheme and that is what the transcript and every export shows.

What can I tell an ethics committee about the platform?

The facts, which are short. Hushscript deletes the audio the moment each transcript is ready and never uses it to train models. Transcripts are encrypted at rest under a retention window you set, from 7 to 365 days with per-transcript overrides, and deletion returns a receipt. Audio from the EU, EEA, Switzerland, and the UK is processed in the EU. These are capabilities to check against your protocol, not a compliance certification.

Which export format works with NVivo or ATLAS.ti?

Use a structured format rather than plain text, because plain text drops speaker labels and timecodes. CSV, TSV, and XLSX keep both as columns for spreadsheet and software-based coding, and JSON suits scripted pipelines. Hushscript exports 21 formats with no watermark and bundles the useful ones as a Research pack.

What does transcribing a study cost?

On the largest prepaid pack, $49.99 for 100 hours, twenty hour-long interviews come to roughly $10 and a 90-minute focus group to about $0.75. There is no subscription and no academic tier: minutes are prepaid, spent only when a file transcribes, and stay valid for 365 days with any transcription or purchase resetting that window.

Can I check it on my own recordings before committing the corpus?

Yes. The first 5 minutes of a file come back speaker-labeled without an account, which is enough to judge the speaker split on your actual participants, room, and accents. Transcribing the full corpus needs an account.

How should I handle interviews in more than one language?

Transcribe in the language spoken, then translate. Recognition covers roughly 99 languages with automatic detection, and adding a translation target costs 25% of the recording duration and puts the original and your working language in tabs of one document. Code-switching inside a single sentence is still the hard case, and a bilingual reviewer is the fix.

Start with 30 free minutes

A $1 hold confirms your card and releases immediately — you're never charged, and 30 free minutes land right away.

Start – 30 free minutes