How To Transcribe Research Interviews: Tools And Method (2026)

The short answer: check what you already have first — Word’s Transcribe feature handles uploaded audio and separates speakers, and Teams, Zoom and Google Meet transcribe live calls. For manual transcription, or when interview audio must not leave your computer, use oTranscribe — free, open source, and entirely local. Before uploading anything, check what your consent form and ethics approval actually allow.

Updated August 2026

Transcribing interviews is the slow part of qualitative research. A one-hour recording can take four hours or more to transcribe by hand, and teachers who take on a research component — a master’s project, an action-research study, a department evaluation — usually discover this after recording the interviews.

The good news is that this got dramatically easier. The bad news is that most guides, including earlier versions of this one, still describe a 2022-era workaround involving file conversion and a free demo tool. That is no longer the shortest path, and for interviews with human subjects it may not be the appropriate one either.

Before anything else: what your consent form allows

This is the step almost every transcription guide skips, and it is the one that can invalidate a study.

Automatic transcription usually means uploading a recording of an identifiable human being to a third-party company’s servers. If your participants signed a consent form, and particularly if your project went through an ethics or institutional review board, there will be terms governing where that recording may go, who may process it, and how long it may be retained.

  • Check the consent wording before choosing a tool, not after. “Recordings will be stored securely and accessed only by the research team” is not obviously compatible with uploading to a consumer transcription service.
  • Check whether your institution has an approved tool. Many universities and districts license a specific service precisely so that a data-processing agreement is already in place. Using the approved one is usually less work than justifying another.
  • Prefer local processing for sensitive material. If the audio involves children, health, safeguarding or employment matters, a tool that never uploads the file removes the question entirely.
  • Pseudonymise early. Replace names in the transcript as you go rather than planning to do it later.

None of this means automated transcription is off-limits. It means the tool choice is partly a compliance decision, and it is far cheaper to make it before you record than to explain it afterwards.

What is transcription?

Transcription is the process of converting speech into text so it can be read, coded and analysed. In research it is what turns interviews, focus groups and recorded discussions into data you can actually work with.

The three styles, and when each is right

Style What it captures Use it when
Verbatim Everything — stutters, pauses, exclamations, non-verbal sounds How something was said matters: discourse analysis, legal records, hesitation as data
Edited The words, minus stutters, false starts and filler Most educational research — readable, and faithful to wording and structure
Intelligent The meaning, with sentences tidied or restructured Reporting and summaries — ⚠️ not for quoting participants in findings

Decide the style before you start, and record it in your method. Switching halfway produces a corpus you cannot analyse consistently — and if you plan to quote participants, intelligent transcription has already edited the words you would be quoting.

Recording well is worth more than any tool

Accuracy in automated transcription is decided mostly at the recording stage. Accent and audio quality affect results more than the choice of software, and no tool recovers detail that was never captured.

  • Record somewhere quiet, with hard surfaces and echo minimised.
  • Put every phone in the room on silent — including the participants’.
  • Place the recording device close enough to reach everyone, not just the interviewer.
  • Ask participants to speak one at a time; overlapping speech is where automated tools fail hardest.
  • When someone says a name — a person, a school, a place — repeat it back. It aids both the transcript and your later accuracy check.
  • Record a few seconds of silence at the start; some tools use it to calibrate.
  • Check the recording is actually running before you begin the real questions.

Tools for transcribing research interviews

1. Start with what your institution already gives you

Most teachers and researchers already have a capable transcription tool and do not know it.

Word’s Transcribe feature converts speech to text and — importantly for interviews — separates each speaker, labelling them “Speaker 1”, “Speaker 2” and so on, which you can rename and correct. Per Microsoft’s documentation it accepts .wav, .mp4, .m4a and .mp3, and on a Microsoft 365 subscription allows up to 300 minutes of uploaded audio per month (a Copilot licence raises that substantially). In Word on the web it works in the new Microsoft Edge and Chrome.

That single feature obsoletes the file-conversion step older guides describe: it takes mp3 and m4a directly, so there is nothing to convert.

Teams, Zoom and Google Meet can all transcribe a call as it happens. If the interview is remote, transcribing it live is far less work than recording and processing afterwards — though the same consent questions apply, and live transcripts still need an accuracy pass.

2. oTranscribe — for manual transcription and anything private

oTranscribe is a free web app built specifically to take the pain out of transcribing recorded interviews by hand. It is the tool to reach for when the audio is sensitive, when the automated output needs heavy correction, or when accuracy matters more than speed.

The feature that matters most for research: your audio file and transcript never leave your computer. It runs in the browser and processes locally, which sidesteps the entire consent-and-data-processing question that cloud transcription raises.

What it does:

  • Combines the audio player and the text editor, so there is no switching between a media player and a document.
  • Pause, rewind and fast-forward from the keyboard without moving your hands off it — the function a foot pedal used to serve.
  • Interactive timestamps you can insert and click to jump back to that moment in the audio. For coding quotes later, this is the single biggest time-saver.
  • Automatically saves to your browser’s storage as you type.
  • Exports to Markdown, plain text and Google Docs.
  • Supports video files with an integrated player.
  • Customisable keyboard shortcuts.

Its limitations, stated plainly: it works on desktop computers only, and because your work is saved in the browser’s own storage, clearing your browser data will delete it. Export to a file at the end of every session — treat the browser copy as a working draft, never as your archive.

It is open source under the MIT licence, created by Elliot Bentley, and is a project of the MuckRock Foundation — which is a reasonable indicator that it is not about to disappear or start charging.

3. Azure Speech to Text — when you need the API

Microsoft’s speech service, now part of Azure Speech in Foundry Tools, handles real-time and batch speech-to-text at scale.

Correcting what this guide used to say. Earlier versions described Azure Speech to Text as “completely free to use” with a demo page that accepted only WAV files. That is out of date. It is a cloud service requiring an Azure account, with a free tier (F0) of 5 audio hours per month for standard real-time transcription and pay-as-you-go beyond it. The WAV-only limitation belonged to the old demo interface, not the service.

For most teachers doing a handful of interviews, this is more setup than the job needs — Word’s Transcribe covers the same ground with no account to configure. Azure earns its place when you are processing many hours, need programmatic access, or want custom vocabulary for technical terms.

Converting audio formats, if you still need to

Far less often necessary now that the mainstream tools accept mp3 and m4a directly. If a tool does demand a specific format, CloudConvert handles audio conversion in the browser. ⚠️ Note that converting means uploading — so the consent question above applies to the converter too, not just the transcription tool.

A realistic workflow

  1. Check consent and institutional policy before choosing a tool.
  2. Record carefully — this determines your accuracy ceiling.
  3. Generate a first-pass transcript with whatever your institution provides, if the material is suitable for cloud processing.
  4. Correct it against the audio in oTranscribe, inserting timestamps at anything you may quote. Automated output always needs this pass; treat it as a draft, never a transcript.
  5. Pseudonymise as you correct.
  6. Export and store properly — out of browser storage, into wherever your data-management plan says it belongs.

The correction pass is not optional. Automated transcription is good and getting better, but it mishears names, technical vocabulary and anything said over the top of someone else — which in interview data is often the interesting part.

When it is worth paying someone else

Doing it yourself is the default for a small study, and for good reason: you hear the data properly, which is genuinely useful for analysis. But there is a point where paying is the rational choice, and teachers running a research project alongside a full timetable often pass it without noticing.

Paying tends to make sense when:

  • You have more than a handful of hours of audio and a fixed deadline.
  • The recordings are clean and one-to-one — the conditions where a service performs well and quotes accurately.
  • Your project has any funding at all. Transcription is one of the few research costs that converts money directly into time.
  • The alternative is transcribing at midnight, where accuracy falls and the correction pass gets skipped.

Doing it yourself tends to be better when:

  • The audio is sensitive and sending it out creates a consent problem you would have to resolve first.
  • The recordings are messy — heavy accents, overlapping focus-group speech, poor audio. You will be correcting a paid transcript so heavily that you may as well have done it.
  • The study is small, or immersion in the data matters to your analytical approach.

If you do buy transcription, check these before ordering: whether they will sign a data-processing agreement your institution accepts; where the audio is stored and for how long; whether human transcribers or only software touch the file; whether pricing is per audio-hour or per turnaround tier; what their stated accuracy figure actually means and on what audio quality; and whether they will follow your transcription style rather than their house one. A cheap transcript in the wrong style is not a saving.

What automated transcription gets wrong

Knowing the failure patterns tells you where to spend your correction time, rather than re-reading everything with equal attention.

  • Proper nouns. Names of people, schools, places and programmes are the most frequent errors and the most damaging, because they are exactly what you will search the transcript for later. This is why repeating names aloud during the interview pays off twice.
  • Subject vocabulary. Curriculum terms, assessment acronyms and anything discipline-specific get replaced with a common word that sounds similar. A transcript that says “formative” where the participant said “summative” inverts the meaning entirely.
  • Overlapping speech. When two people talk at once, most tools pick one and drop the other, usually without any marker that something was lost. In interview data the interruption is often the interesting moment.
  • Numbers and years. “Fifteen” and “fifty”, “2015” and “2050” are common substitutions, and they read as perfectly plausible sentences.
  • Accents and quiet speakers. Accuracy drops for speakers the model has heard less of, and for anyone sitting furthest from the microphone. Worth knowing whose sections need the closest checking.
  • Confident nonsense. Automated transcripts rarely leave gaps. They produce a fluent, grammatical sentence that is wrong — which is far harder to spot than an obvious blank.

A spot-check that catches most of it

Before trusting a transcript, pick three one-minute stretches — one near the start, one from the densest part of the discussion, one where speakers overlapped — and check them word by word against the audio. If those are clean, the rest probably is. If they are not, you know to do a full pass, and you have found out early rather than at the analysis stage.

Then check every proper noun and every number in the whole transcript, regardless. Those two categories carry most of the risk and take minutes to scan for.

Focus groups are harder than interviews

Everything above gets more difficult with more than two voices in the room, and it is worth planning differently rather than being surprised.

  • Speaker attribution degrades. Automated speaker separation copes with two people well and with six poorly, especially when voices are similar or people interrupt.
  • Draw a seating map at the start and note who sits where. Matching “Speaker 3” to a real participant afterwards is otherwise guesswork.
  • Ask participants to say their name the first time they speak. Slightly awkward for thirty seconds, and it makes attribution possible for the whole recording.
  • Consider recording from two positions if the group is large. A second device at the other end of the table rescues the quiet end.
  • Budget more correction time — a focus group frequently takes longer to transcribe than the equivalent minutes of one-to-one interview.

Getting the transcript ready for analysis

The transcript is not the output. What you actually need is something you can code, search and quote from with confidence, and a few decisions at transcription time make the analysis stage far less painful.

  • Insert timestamps as you correct, at minimum wherever a passage might be quoted. When you later need to check a quote or hear the tone it was said in, a timestamp is the difference between five seconds and twenty minutes.
  • Keep speaker labels consistent across every transcript in the study — the same participant should have the same pseudonym everywhere. Fixing this later across a dozen files is miserable.
  • Export to plain text or Markdown if the transcript is heading into analysis software. Heavily formatted documents tend to import badly, and formatting carries no analytical value.
  • Keep the corrected transcript and the raw automated output separately, at least until analysis is finished. If a quote looks surprising, being able to see what the tool originally produced tells you quickly whether it is a real finding or a transcription artefact.
  • Record your transcription conventions in your method — which style you used, how you marked pauses or inaudible passages, how you pseudonymised. Examiners and reviewers ask, and it is much easier to write down while you are doing it.

What changed since this guide was written

Then Now
Convert your file to WAV first Mainstream tools accept mp3, m4a and mp4 directly
“Azure Speech to Text is completely free” Free tier of 5 audio hours per month, then paid; needs an Azure account
Speaker separation was a premium feature Built into Word’s Transcribe and most meeting platforms
Privacy rarely discussed The main constraint on tool choice for interview data

Frequently asked questions

Is oTranscribe still available?

Yes. It is free, open source under the MIT licence, and actively maintained as a project of the MuckRock Foundation. It runs on desktop browsers only, and your audio and transcript stay on your own computer.

What is the fastest way to transcribe an interview?

Generate a first-pass transcript automatically, then correct it against the audio rather than typing from scratch. If you have a Microsoft 365 subscription, Word’s Transcribe feature is usually the shortest route because it takes common audio formats directly and separates speakers. Budget time for correction either way.

Can I upload interview recordings to a transcription service?

Only if your consent form and any ethics approval allow it. Uploading means sending a recording of an identifiable person to a third party. Check your institution’s approved tools first, and for sensitive material prefer a tool that processes locally.

How long does transcribing an interview take?

Typed by hand, a one-hour recording commonly takes four hours or more depending on audio quality and how many speakers overlap. Correcting an automated draft is considerably faster, but it is not instant — plan for real time, not zero.

Do I need verbatim transcription for educational research?

Usually not. Edited transcription — the words, minus stutters and filler — is the normal choice, and is readable while staying faithful to what participants actually said. Choose verbatim when hesitation or delivery is itself part of your data, and avoid intelligent transcription for anything you intend to quote.

JH

Josh Hutcheson — Editor, PriorityLearn

Josh researches, writes, and updates the answers on PriorityLearn, checking each one against current tools, official sources, and real school policies — and flagging what varies by state or district. About PriorityLearn →

Scroll to Top