Blog
Research Transcription for Interviews and Focus Groups
Author: Hushscript Published: Last reviewed:
A recording is not data yet. It becomes data once somebody has decided what gets written down, how speakers are marked, which identifying details come out, and what the file looks like when it reaches your coding software. Those four decisions are research transcription, and every one of them is methodological rather than clerical.
Making them before the first interview is far cheaper than making them at interview fourteen. Here is what each decision changes, what belongs in your methods section, and where automatic transcription genuinely helps a study instead of just making the folder look finished.
Choose the transcript type before the first session
There is no neutral transcript. Every convention makes some things visible and hides others, so choose for the analysis you actually intend to run.
| Transcript type | What it keeps | Fits |
|---|---|---|
| Verbatim | Fillers, false starts, pauses, overlaps | Conversation and discourse analysis |
| Intelligent verbatim | The words, minus fillers and repairs | Thematic analysis, grounded theory |
| Clean | Edited for grammar and readability | Quotes for publication and reports |
| Denaturalized | Standardized grammar across the corpus | Cross-case comparison, large samples |
The expensive mistake is not choosing badly, it is choosing twice. A corpus half verbatim and half cleaned cannot be coded consistently, and re-transcribing twenty interviews to repair that is a month nobody has budgeted. Write the convention down in advance, including how you will mark laughter, interruptions, and inaudible stretches.
What machines get right, and what they miss
Automatic transcription is good enough to be the first draft of nearly any study, and not good enough to be the last. Pew Research Center compared human and machine transcription on recorded sermons and found the machines kept pace on clear, continuous speech while falling behind on proper names, specialist terms, and passages where a speaker shifted tone or volume. Those are the passages qualitative researchers tend to quote.
So the working pattern is a draft you did not type, followed by a review pass you did. Read along with the audio once, correct names and terminology, fix speaker attribution across the crosstalk, and stop there. The review is the part that costs real time, and the part that cannot be skipped.
Speaker labels are already pseudonyms
Automatic speaker separation splits a recording by voice and labels the results Speaker 1, Speaker 2, and so on. It recognizes voices, not people, which is quietly useful: the label is a pseudonym by default and stays one unless somebody types a real name into it.
Keep it that way. Rename Speaker 1 to P07 or to a role label once and the change carries through the whole transcript and every export, so a participant ID scheme is applied consistently instead of by hand in forty places. Speaker labels are free on every Hushscript transcript with no cap on the number of voices, which is what a six-person focus group needs.
The honest limit is crosstalk. When participants talk over each other the split needs a human pass, and the timecodes are what tell you which minutes to re-listen to.
Put the transcription method in the paper
Reviewers increasingly ask how the text was produced, not only what it says. The APA’s Journal Article Reporting Standards for qualitative research ask authors to describe how data were recorded and transformed, which makes transcription something you report rather than assume.
One methods paragraph answering these is usually enough:
- Which convention. Verbatim, intelligent verbatim, clean, or denaturalized, and why it suits the analysis.
- Who produced the text. Researcher, transcription service, automatic recognition, or an automatic draft with researcher review.
- What quality control ran. Who checked against the audio, on what sample, and how disagreements were resolved.
- How language was handled. Dialect, code-switching, and who translated when the analysis language differs from the interview language.
- What was removed. Pseudonymization rules, redacted places and dates, generalized employers and job titles.
De-identify while transcribing, not afterwards
Identifying detail arrives mid-sentence: an employer, a ward, a street, a court date, the one colleague who would recognize the story. Stripping it later means reading the entire corpus a second time, so fold it into the transcription pass instead.
Doing that inside the transcript rather than in exported copies also means the pseudonymized version is the only one that ever circulates. Find-and-replace fixes a recurring place name across the document in one action, and saved dictionaries carry the study’s names, places, and technical vocabulary into recognition so fewer of them come back wrong to begin with.
What you can say about the platform matters as much as the redaction. Hushscript deletes the audio the moment each transcript is ready and never uses it to train models; transcripts are encrypted at rest under a retention window you set, from 7 to 365 days with per-transcript overrides; deletion flows return receipts. Audio from the EU, EEA, Switzerland, and the UK is processed in the EU. For a recording that should not leave a stored transcript at all, private transcription hands back a password-protected .husharchive file instead.
Export for the software that receives it
A transcript rarely stays where it was made. It moves into NVivo, ATLAS.ti, MAXQDA, or a spreadsheet, and later into a manuscript. Plain text is where speaker labels and timecodes go to die, so pick the format by its destination.
- Coding software: CSV, TSV, or XLSX, which keep speaker and timecode as columns.
- Scripted pipelines: JSON, where the structure is already parsed for you.
- Writing and appendices: DOCX or PDF.
- Archiving: plain text beside a metadata file recording the convention you used.
Hushscript exports 21 formats with no watermark and bundles the relevant ones as a Research pack. Timecodes surviving the export is what lets a quote in your findings point back at the second it was spoken, which is the difference between a citable transcript and a long document.
Multilingual fieldwork without a vendor per country
A comparative study should not need a different transcription supplier in every site. Recognition covers roughly 99 languages with automatic detection, with the strongest accuracy on a flagship set of 18 language variants, so one workflow carries the whole sample.
Translation is a separate decision from transcription, and the order matters: transcribe in the language spoken, then translate. Each translation target costs 25% of the recording duration and puts the original and your working language in tabs of the same document, each exportable on its own, so you quote the participant’s words and code in the language you think in. Report who translated and how disagreements were settled, because two translators will not render the same phrase identically.
Costing it against the grant line
Manual transcription is priced in researcher hours, and twenty hour-long interviews is the sort of number that quietly consumes a term. Automatic transcription moves that into a budget line small enough to stop being a constraint on sample size.
Hushscript sells prepaid minutes rather than a subscription, so fieldwork months cost what they use and writing months cost nothing. On the largest pack, $49.99 for 100 hours, twenty hour-long interviews come to roughly $10 and a 90-minute focus group to about $0.75. Minutes stay valid for 365 days and any transcription or purchase resets that window, which suits a study that pauses for ethics approval. New accounts get 30 free minutes, released either by a $1 card check that is authorized and immediately released and never charged, or by your first purchase if you pay another way.
Where to start
Use the last interview you actually recorded, not a clean test file. Put it through research transcription and read the first 5 minutes that come back: that shows you how speaker identification behaves on your participants, your room, and your accents, which is the only accuracy figure that means anything for your study.
Then write the convention into your protocol while there are still thirty interviews left to apply it to.
Independent sources and standards
Hushscript consulted these independent, non-competing references. They explain research, standards, or platform behavior and do not endorse Hushscript.
Sources reviewed: