← Back to blog

Focus Group Transcription: How to Transcribe a Focus Group (and What to Expect from AI)

To transcribe a focus group well, record it with a speaker plan (a central microphone, a backup recorder, a seating chart and a round of introductions), run the recording through an AI tool for a speaker-separated first draft, then review it against the audio yourself, fixing crosstalk and assigning each voice to a participant code. This guide explains why group audio is so hard, which verbatim style to choose, how to label and anonymize participants, how long it really takes, what AI can and cannot do, and what consent and GDPR require.

Eight adults seated around a round table in a bright meeting room, with a table microphone in the center and a video camera on a tripod in the corner

Contents

  1. Why are focus groups the hardest audio to transcribe?
  2. How should you record a focus group to make transcription easier?
  3. Full verbatim or clean verbatim for a focus group?
  4. How do you label and anonymize speakers?
  5. How to transcribe a focus group step by step
  6. How long does focus group transcription take, and what does it cost?
  7. What can AI do with crosstalk, and what can't it?
  8. Consent, confidentiality and GDPR for focus group recordings
  9. How the transcript feeds thematic analysis
  10. Common focus group transcription mistakes
  11. Where AudiosTranscribe fits
  12. Sources and further reading
  13. Frequently asked questions

Why are focus groups the hardest audio to transcribe?

Focus groups are hard to transcribe because several people talk to each other rather than to one interviewer. Voices overlap, short reactions fly across the table, participants sit at different distances from the microphone, and the transcriber has to work out not only what was said but who said it.

The whole point of the method is interaction, so the richest moments are often the messiest on the recording: two participants disagreeing at once, a burst of laughter, someone finishing another person's sentence. Julia Bailey's methods paper on transcribing notes that transcription takes at least 3 hours per hour of talk, and up to 10 hours per hour when a fine level of detail is needed. Group recordings, with their overlapping voices, usually sit towards the slow end.

The difficulties fall into five groups:

The quality of your transcript is therefore decided long before anyone types a word, when you plan the recording.

How should you record a focus group to make transcription easier?

Record a focus group with a good central microphone, a backup recorder, a seating chart and a round of introductions. These habits cost minutes on the day and save hours of guesswork afterwards.

Microphone placement: one mic or several?

One omnidirectional conference microphone in the middle of the table works for a small group seated close together. For a long table or soft-spoken participants, add a second recorder at the other end, or individual lapel microphones if your setup allows: separate channels make attribution far easier but add a synchronization step. Either way, run a short sound check with the actual group.

For online focus groups on Zoom, Teams or Meet, use the platform's recording or capture your computer's system audio, and ask participants to wear headsets and mute when not speaking. A 2024 methods paper by Eftekhari on speech recognition in qualitative research recommends good quality audio, headsets to minimize echo, less overlapping speech and a quiet environment to limit transcription errors.

Video, seating chart and introductions

Tip: when two people talk at once, a good moderator briefly names the next speaker ("Let's hear from P4 first, then P6"). It keeps the discussion fair and leaves a verbal marker on the recording that makes attribution much easier.

Full verbatim or clean verbatim for a focus group?

Choose clean (intelligent) verbatim if you plan a thematic or content analysis, and full verbatim if the interaction itself is your data. For most applied, market and UX projects clean verbatim is enough; studies of how opinions form in the group, or of how people talk, need full verbatim with overlap marked.

Your analysisRecommended styleWhat to keep in the transcript
Thematic or content analysisClean verbatimEvery contribution and speaker label; fillers and false starts removed
Market or UX research findingsClean verbatimQuotable statements, clear attribution, time stamps to find clips
Analysis of group interactionFull verbatimOverlaps, interruptions, laughter, agreement and disagreement markers
Conversation or discourse analysisFull verbatim with notationPauses, overlap onsets, emphasis, following a convention such as Jefferson

Even with clean verbatim, do not strip the short reactions that carry meaning in a group. "Yeah, same" from three participants is evidence of shared experience, and the Eftekhari paper lists square brackets as a common convention for marking overlapping talk or emotion. For the two styles in depth, see our guide to verbatim vs clean verbatim transcription; for transcription across qualitative methods, see our guide to transcription in qualitative research.

How do you label and anonymize speakers?

Label every turn with a stable code, such as M for the moderator and P1 to P8 for participants, and keep the key that links codes to real people in a separate, secured file. Codes are easier to anonymize than names and stay consistent across groups.

In a group, participants also name each other ("As Sarah said..."): replace those names with the matching code too.

Sample focus group transcript (clean verbatim, illustrative):

[00:21:08] M: What made you stop using the app after the first week? P3: Honestly, the notifications. It felt like it was nagging me. P5: [overlapping] Same, I turned them off on day two. P3: Right, and then I just forgot it existed. P1: I had the opposite experience, actually. I liked the reminders. [Several]: [laughter] M: Tell me more about that, P1.

What makes it analysis-ready: a time stamp, a distinct moderator label, participant codes instead of names, an overlap marker where P5 cut in, and a group reaction honestly attributed to several people.

How to transcribe a focus group step by step

The fastest reliable workflow is an AI first pass followed by a careful human review: the machine types, you listen and judge.

1

Plan the recording

Choose your microphones, a backup recorder and, if relevant, a camera. Assign participant codes and prepare the seating chart.

2

Record with introductions and a seating chart

Confirm consent to recording, run introductions so every voice is heard once on its own, and have the note-taker fill in the seating chart and flag key moments.

3

Run an AI first pass

Upload the recording to a transcription tool that separates voices. You get a time-stamped draft with generic labels (Speaker 1, Speaker 2...). Treat it as a draft, never as the final transcript.

4

Map voices to participant codes

Use the introductions, the seating chart and any video to match each generic label to a participant, then rename the labels M, P1, P2 and so on.

5

Review against the audio, crosstalk first

Listen back while reading, prioritizing overlaps, quiet speakers, jargon and proper nouns. Fix attributions, split merged turns and mark what you cannot resolve as [inaudible] with a time stamp.

6

Anonymize and add a header

Replace names and identifying details consistently, log the replacements, and add a header (group ID, date, moderator, number of participants, convention).

7

Export for analysis

Export as .docx or .txt with one speaker label per turn for NVivo, ATLAS.ti or MAXQDA, and keep an unedited master copy.

How long does focus group transcription take, and what does it cost?

By hand, expect several hours of work per hour of recording, and more for a busy group with overlapping voices. With an AI first pass, the typing disappears but the review does not.

What the published figures say

3 to 10 h
Manual transcription per hour of talk, depending on the level of detail (Bailey, Family Practice, 2008)
1.5 to 3.5 h
Checking time per automatic transcript (Teams live transcription, 2021-22) in a study of one-to-one interviews averaging 65 minutes (Eftekhari, 2024)

The second figure comes from individual interviews, not focus groups. With more voices and more overlap, plan for a longer review, and time your own first transcript before you budget the rest of the project.

On cost, doing it yourself costs only your time, which the figures above show is considerable. If you request quotes from a human transcription service, check whether the price changes with the number of speakers, the verbatim level, speaker identification and turnaround, since focus groups tend to need all of them. An AI tool plus your own review saves typing time, not thinking time.

Watch out: do not budget focus group review time from a one-to-one interview. Run one group through your full workflow first, measure how long the review really took, then multiply.

What can AI do with crosstalk, and what can't it?

AI handles the clear, one-at-a-time stretches of a focus group quickly and well, but overlapping speech remains its weak spot: expect missing words, merged turns and misattributed speakers where people talk at once.

This is not a limitation of one product. The review of speaker diarization research by Park and colleagues describes speech recognition on overlapping speech as one of the main challenges in meeting transcription, reports that studies of meeting recordings observed an average of 12% to 15% of speaker overlap, and notes that many traditional diarization systems focused only on non-overlapping regions. A lively focus group is rarely an easier case.

In practice, an AI transcription tool will:

The honest rule: AI turns focus group transcription from a typing job into a reviewing job. You still need to listen, especially to the overlapping passages, which often hold the most interesting interaction.

Consent, confidentiality and GDPR for focus group recordings

Recording a focus group needs every participant's informed agreement, a warning that confidentiality within the group cannot be guaranteed and, in the EU and UK, a lawful basis and safeguards for processing the recordings.

The University of Connecticut IRB guidance on focus groups sums up the specifics: disclose the use of recording devices in the consent form and consent discussion, ask each member again at the session whether they agree to be recorded, tell participants that what is said should not be shared outside the group and that the confidentiality of what they say cannot be guaranteed, and ask them not to use names when the session is recorded.

Data protection is a separate question from research ethics. The UK regulator's guidance on the research provisions of the UK GDPR points out that consent to take part in a study is distinct from consent as a lawful basis for processing, and that in most research cases consent is not the most appropriate lawful basis, partly because it can be withdrawn at any time. In a group recording, think in advance about what a withdrawal would mean in practice, and agree the lawful basis with your data protection officer or ethics committee.

Finally, check where your transcription tool processes and stores the audio: a recording of eight voices is personal data about eight people. A processor hosted in Europe that deletes the audio after processing keeps the data trail short for European studies.

How the transcript feeds thematic analysis

A focus group transcript is analyzed at two levels: what was said (themes across the group) and how it was said (agreement, disagreement, how views shifted in discussion). Accurate speaker labels are what make both possible.

For thematic analysis, the approach developed by Virginia Braun and Victoria Clarke is the usual reference point: a family of methods for exploring and interpreting patterned meaning across a dataset. With focus groups:

Researcher seen from behind facing a wall of colored sticky notes grouped into clusters, with printed transcript pages and highlighters spread on a table
Once the transcript is reviewed and labeled, coding can start: grouping contributions into candidate themes across the whole group.

Common focus group transcription mistakes

  1. Relying on one recorder. A single device with a flat battery or a muffled corner loses a session you cannot rerun.
  2. Skipping introductions. Without a clean voice sample per participant, matching voices to codes becomes guesswork.
  3. Publishing the AI draft. An unreviewed draft is fine for your own notes, never for coding or quotes.
  4. Guessing speakers. Mark uncertain turns as uncertain instead of forcing an attribution.
  5. Deleting group reactions in the name of clean verbatim, and with them the evidence of consensus.
  6. Mixing conventions across groups. Write your transcription rules down before the first session and stick to them.
  7. Forgetting names spoken aloud during anonymization, including names of colleagues, employers and places.

Where AudiosTranscribe fits

AudiosTranscribe provides the AI first pass described above: a time-stamped transcript with voices separated into distinct speakers, ready for your review.

What focus group researchers get:

  • Automatic voice separation: turns are split by speaker; you decide who is who and rename labels (M, P1, P2...) in one click
  • Upload or record: upload an audio or video file from the web app in any browser, or record an online focus group on Zoom, Teams or Meet with the Windows desktop app, which captures system audio and microphone without a bot joining the call, with participants' consent; transcription runs after the recording
  • Exports to TXT, Word, PDF and SRT on paid plans, for NVivo, ATLAS.ti or MAXQDA
  • Speaker analytics (talk time, approximate where voices overlap) on the Pro plan, handy for spotting a participant who dominated the discussion
  • Hosted in Europe, GDPR compliant, audio deleted after processing
  • 120 free minutes per month to test the workflow on a real session; research teams can share a workspace on the Team plan (39 € per month for the workspace, VAT not applicable, up to 20 members, 20 hours of transcription pooled per month)

For one-to-one interviews, see our interview transcription page; the pricing page details each plan.

Sources and further reading

Your institution's ethics committee or IRB may set stricter rules, so check its requirements before your first session.

Frequently asked questions

How do you transcribe a focus group?
Record the session with a good central microphone, a backup recorder, a seating chart and a round of introductions. Upload the recording to an AI transcription tool that separates voices to get a time-stamped first draft, map each voice to a participant code using your seating chart and the introductions, then review the whole transcript against the audio, paying most attention to overlapping speech. Finish by anonymizing names and exporting a .docx or .txt file for your analysis software.
How long does it take to transcribe a focus group?
By hand, a methods paper by Julia Bailey in Family Practice (2008) puts transcription at a minimum of 3 hours per hour of talk and up to 10 hours when fine detail is needed. Group recordings, with their overlapping voices, usually sit towards the slow end. With an AI first pass the typing takes minutes, but you still need a review. In one 2024 study of one-to-one interviews transcribed with Teams' live transcription, checking each transcript took 1.5 to 3.5 hours, so plan for more with a busy group.
Can AI transcribe a focus group accurately?
AI transcribes the clear, one-at-a-time parts of a focus group quickly and well, but overlapping speech remains a known weak spot for speech recognition and speaker separation. Expect missed words, merged turns and misattributed speakers where people talk at once, and expect the tool to separate voices without knowing who each person is. Used as a first draft followed by a researcher's review against the audio, AI makes focus group transcription much faster while keeping the researcher in control of accuracy.
Should a focus group transcript be verbatim?
Yes, but choose the level of verbatim to match your analysis. Clean (intelligent) verbatim, which removes fillers and false starts but keeps every contribution and speaker, suits thematic analysis and most market and UX research. Full verbatim, with overlaps, interruptions and laughter marked, is needed when the group interaction itself is what you analyze, as in conversation or discourse analysis.
How do you label speakers in a focus group transcript?
Use short, stable codes: M for the moderator and P1, P2, P3 and so on for participants, prefixed with a group number if you run several sessions (G1-P3). Keep the key linking codes to real people in a separate, secured file. When you cannot tell who spoke, write [P?] or [Several] rather than guessing, because a wrong attribution distorts any analysis by participant.
Do you need consent to record and transcribe a focus group?
Yes. Participants should know before the session that it will be recorded and transcribed, and some IRB guidance asks you to confirm it again with each participant at the start of the session. Because other participants hear everything, you should also tell people that confidentiality within the group cannot be guaranteed. In the EU and UK, consent to take part in research is separate from the lawful basis you rely on under the GDPR, so agree that basis with your ethics committee or data protection officer.
Is there a free way to transcribe a focus group?
You can transcribe by hand for free, at the cost of many hours. AudiosTranscribe's free plan includes 120 minutes of transcription per month, enough to run a full session through an AI first pass with speaker separation and review it in the app (free-plan file-size limits apply). Exports to Word, PDF, TXT and SRT require a paid plan.

Turn your next focus group into a transcript you can review

Upload a recording, get a time-stamped transcript with voices separated, assign your participant codes and start the review. 120 free minutes per month, no credit card, GDPR, EU hosting.

Start for free