← Back to blog

Verbatim Transcription: Full Verbatim vs Clean Verbatim, with Examples

Verbatim transcription means writing down exactly what was said, word for word. The real decision is how much of the speech you keep: every "um", stutter and laugh (full verbatim), only the meaningful words (clean or intelligent verbatim), or a polished rewrite (edited). This guide shows the same passage in all three styles, the notation to use, which style fits research, legal, journalism, meetings and subtitles, and how to review an AI transcript into true verbatim.

Researcher wearing headphones reviewing a transcript on a laptop, with an open notebook, a pen and a cup of tea on the desk

Contents

  1. What is verbatim transcription?
  2. The three transcription styles: full, clean and edited
  3. Verbatim transcription example: one passage, three styles
  4. What each style keeps or removes
  5. Verbatim notation: timestamps, speaker labels and tags
  6. Which verbatim style should you choose?
  7. How AI transcription handles verbatim
  8. How to review an AI transcript into true verbatim
  9. Common verbatim transcription mistakes
  10. Sources and further reading
  11. Frequently asked questions

Short answer: full verbatim keeps everything you can hear (fillers, false starts, stutters, repetitions, laughter, overlaps). Clean verbatim, also called intelligent verbatim, keeps the words and meaning but drops the stumbles. Edited transcription rewrites for readability and is no longer verbatim. Use full verbatim when how something was said matters, clean verbatim when what was said is enough.

What is verbatim transcription?

Verbatim transcription means turning speech into text exactly as it was spoken, word for word, without summarizing or paraphrasing. The transcript follows the speaker's own wording and word order, even when that wording is informal, ungrammatical or unfinished.

The complication is that "exactly as spoken" can mean two different things. Real speech is full of material that never makes it into written language: hesitations, restarts, repeated words, laughter, people talking over each other. Whether that material belongs in the transcript is the whole question behind the terms full verbatim, clean verbatim and "verbatim vs non-verbatim". Researchers have argued about it for decades. A widely cited paper by Oliver, Serovich and Mason (2005, Social Forces) describes the two ends of the spectrum: one where "every utterance is transcribed in as much detail as possible", and one where idiosyncratic elements such as stutters, pauses and non-verbal sounds are removed.

Neither end is "right": each style discards some information to make the text useful for a particular job. The real question is what you need to keep.

The three transcription styles: full, clean and edited

Many transcription providers distinguish three levels, from the most literal to the most polished. The names vary between providers and universities, so check what each one actually does rather than trusting the label.

Full verbatim (true or strict verbatim)

Full verbatim captures every audible word and sound: fillers ("um", "uh", "er"), false starts, stutters, repetitions, backchannels ("mm-hmm", "yeah") and audible non-verbal events such as laughter, sighs or a door slamming. Grammar is left exactly as spoken. Overlapping speech, pauses and passages you cannot hear are marked with tags. It is the slowest style to produce and the hardest to read, but nothing is silently lost.

Clean verbatim (intelligent verbatim)

Clean verbatim keeps the speaker's words, word order and meaning, but removes the noise: fillers, stutters, false starts and repetitions that add nothing. Some style guides also allow light corrections, such as fixing an obvious slip of the tongue, while keeping dialect and slang. It reads naturally and is quick to scan, which is why it is the default for most interviews, meetings and published quotes. You will sometimes see it called "non-verbatim", although the words are still the speaker's own.

Edited transcription

Edited transcription (sometimes called a "clean read" or "polished" transcript) goes a step further: the transcriber fixes grammar, tightens rambling sentences and may reorder phrases so the text reads like prose. It is ideal for blog posts, reports or show notes, but it is an interpretation, not a record. Never quote an edited transcript as someone's exact words.

Verbatim transcription example: one passage, three styles

The easiest way to see the difference is to transcribe the same exchange three times. The dialogue below is invented for illustration: an interviewer asks a participant (coded P03) about a new shift schedule at work.

1. Full verbatim
[00:04:12] Interviewer: So, um, how did the new rota go for, for your team?
[00:04:17] P03: Uh, well, it was, it was- honestly? At first it was a mess. [laughs] Like, n- nobody knew who was, uh, covering the Friday shifts, you know?
[00:04:29] Interviewer: Mm-hmm.
[00:04:30] P03: And then, um [pause] after the second month or so it kind of, it kind of settled down, and [crosstalk]
[00:04:38] Interviewer: [crosstalk] So it got bett-
[00:04:39] P03: Yeah, yeah, it got better. We ain't going back to the [inaudible 00:04:42], put it that way.
2. Clean (intelligent) verbatim
[00:04:12] Interviewer: How did the new rota go for your team?
[00:04:17] P03: Honestly? At first it was a mess. [laughs] Nobody knew who was covering the Friday shifts.
[00:04:30] P03: Then, after the second month or so, it settled down.
[00:04:38] Interviewer: So it got better?
[00:04:39] P03: Yeah, it got better. We ain't going back to the [inaudible 00:04:42], put it that way.
3. Edited transcript
Interviewer: How did the new rota work for your team?
P03: At first it was chaotic: nobody knew who was covering Friday shifts. After about two months it settled down, and the team would not go back to the old system.

Look at what changed. The clean version drops "um", "uh", "you know", the stutter ("n- nobody") and the repeated "it was, it was", and it tidies the overlap into a plain question. It still keeps "Honestly?", the laugh and "ain't", because they carry tone and they are the participant's own voice. The edited version is shorter and easier to read, but "a mess" has become "chaotic", "we ain't going back" has become "the team would not go back to the old system", and the inaudible word has been filled in with a guess. That is fine for a report summary; it is not a quote.

What each style keeps or removes

Here is how the three styles treat the elements that most often cause disagreement between transcribers. House styles differ at the edges, so treat this as the common baseline and write down your own choices.

Element Full verbatim Clean verbatim Edited
Fillers (um, uh, er, like, you know)KeptRemovedRemoved
False starts ("it was- honestly")Kept, with a hyphen at the cut-offRemoved unless they change meaningRemoved
Stutters ("n- nobody")KeptRemovedRemoved
Repetitions ("for, for")KeptRemoved unless used for emphasisRemoved
Backchannels (mm-hmm, yeah)Kept as separate turnsUsually removed, kept if they answer a questionRemoved
Non-verbal cues ([laughs], [sighs])KeptKept only when they change meaningUsually removed
Overlapping speechMarked with [crosstalk] and transcribed where audibleSimplified into clear turnsMerged or dropped
Grammar and slangExactly as spokenAs spoken; obvious slips may be fixedCorrected
Exact quote?YesYes, of the words that matterNo

Verbatim notation: timestamps, speaker labels and tags

Whatever the style, a transcript needs a small, consistent set of conventions so readers know what the brackets mean. There is no single universal standard for everyday transcription; the set below is a common practical baseline. Pick one, write it down, and stick to it.

NotationWhat it meansWhen to use it
Interviewer: / P03:Speaker label at the start of each turnAlways. Use pseudonyms or codes if participants must stay anonymous
[00:04:12]Timestamp (hours:minutes:seconds)At each turn, or at a fixed interval, so you can jump back to the audio
[inaudible 00:12:31]A word or passage you cannot make out, with the time it occursEvery time. Never guess silently
[rota?]A best guess at an unclear wordWhen you are fairly, but not fully, sure
[crosstalk]Two or more people speaking at onceFull verbatim, and in clean verbatim when overlap hides content
[laughs] [sighs] [coughs]Audible non-verbal eventsFull verbatim; clean verbatim when they change meaning
[pause]A noticeable silenceFull verbatim, or anywhere the hesitation matters
bett-A word cut off mid-wayFull verbatim

When plain tags are not enough: Jefferson notation

Conversation analysts need far more detail than "[pause]". The standard system in that field was developed by Gail Jefferson; her Glossary of transcript symbols (2004), published in Conversation Analysis: Studies from the first generation, defines symbols such as [ for the point where overlapping talk begins, = for no gap between turns, (0.5) for a pause timed in tenths of a second, (.) for a very brief interval, colons for a stretched sound, empty parentheses ( ) for talk the transcriber could not make out, and doubled parentheses (( )) for the transcriber's own descriptions. It is a specialist skill; for most interviews and meetings, the tags above are plenty. Our guide to transcription in qualitative research shows a lightly notated research excerpt if you want to see the two approaches combined.

Which verbatim style should you choose?

Choose from the use of the transcript, not from habit: the more your analysis or legal exposure depends on the way people spoke, the closer you stay to full verbatim.

Interview seen from behind the interviewer, with a small handheld digital voice recorder and an open notepad on the table between the two people
The transcription style you need is decided by what you will do with the recording, ideally before the interview even starts.

Qualitative research

For conversation analysis, discourse analysis and other methods that study talk itself, use full verbatim, often with Jefferson notation. For thematic analysis and most interview studies, clean verbatim is usually enough, as long as you keep hesitations that carry meaning. Be careful with the jargon: the labels "naturalized" and "denaturalized" are not used consistently. Oliver and colleagues use "naturalized" for the detailed, every-utterance style, while McMullin (2021, Voluntas) pairs "denaturalized" with full verbatim and "naturalized" with intelligent verbatim. Describe what your transcripts keep in your methods section rather than relying on the label. The University of Bath research data guidance on transcription is a good checklist for those decisions. Focus groups are the hardest case, with more voices and far more overlap; we cover that in our guide to focus group transcription.

Legal and compliance

In legal contexts, "verbatim" means an exact record, with no paraphrase. In US federal courts, 28 U.S.C. § 753 requires court sessions to be "recorded verbatim by shorthand, mechanical means, electronic sound recording, or any other method", subject to Judicial Conference regulations and the judge's approval. Official court transcripts come from that certified process, not from your own transcription. For your own evidential recordings (an investigation interview, a complaint call, a disciplinary hearing), keep full verbatim, mark every inaudible passage and never clean the master copy.

Journalism

A clean transcript is the practical way to find the story, but check every quote against the audio before publishing. Anything that shifts meaning or tone (a sarcastic "great", a long hesitation, a retracted word) has to be heard, not read. Your publication's style guide decides how far a quote may be tidied.

Business meetings and interviews

For meetings, hiring interviews, customer calls and user research, clean verbatim is usually the right choice: readable, searchable, and exact where decisions and commitments are made. Many teams circulate an edited summary and keep the clean transcript as the reference. If you need to transcribe your interviews with separated speakers and timestamps, a dedicated interview transcription workflow saves you most of the typing.

Subtitles and captions

Captions sit between the two styles. The US rule for closed captioning of televised programming, 47 CFR § 79.1, asks captions to match the spoken words in the order spoken "without paraphrasing, except to the extent that paraphrasing is necessary to resolve any time constraints", to mirror slang or grammatical errors that are intentional in the dialogue, and to convey non-verbal information such as the identity of speakers, music, sound effects and audience reaction. Even where that rule does not apply, it is a sound benchmark for subtitles: stay close to verbatim and keep sound cues.

How AI transcription handles verbatim

Automatic speech recognition gives you a fast first draft, but do not assume that draft is full verbatim. Engines generally aim for readable text, so their output often lands closer to clean verbatim: fillers may be dropped or kept inconsistently, stutters and false starts smoothed out, non-verbal cues rarely marked, and overlapping speech merged into a single turn or lost. Behavior varies by tool, so test yours on a short sample first.

McMullin notes that speech recognition may now give researchers "good enough" first drafts, but that voice-to-text software "is also generally less accurate in discerning multiple voices or different accents", and that such transcripts still need to be carefully checked by the researcher. In practice, AI saves you the typing; the listening is still yours.

What typically needs a human pass:

Where AudiosTranscribe fits: you upload a recording (or record it with the Windows desktop app) and get a timestamped transcript in which the voices are separated automatically; you then put a name on each speaker in one click. Treat it as a first draft to review against the audio, not as a finished full-verbatim record. The service is hosted in Europe, GDPR-compliant, and audio is deleted after processing. The free plan includes 120 minutes per month, and TXT, Word, PDF and SRT exports come with the paid plans.

How to review an AI transcript into true verbatim

Turning an AI draft into reliable verbatim is a structured listening job. Six steps keep it fast and consistent.

1

Write your style sheet first

Decide full or clean verbatim, and list your tags ([inaudible 00:00:00], [crosstalk], [laughs], [pause]), speaker label format and timestamp interval. One short page is enough, and it keeps several reviewers consistent.

2

Fix the speakers before the words

Name each voice, then skim the whole transcript for turns attributed to the wrong person, especially short replies like "yes" or "mm-hmm". Correcting labels first stops you from re-reading everything later.

3

Listen through for the words

Play the audio at a slightly reduced speed and follow the text. Restore fillers, stutters, false starts and repetitions if you need full verbatim, and correct misheard words, names, numbers and negations in any style.

4

Add the tags

On the same pass or a second one, mark non-verbal cues, pauses, overlapping speech and every passage you cannot make out, with its timestamp. Replace any word the engine invented over unclear audio with [inaudible] or a bracketed best guess.

5

Re-check the hard spots

Go back to each flagged timestamp with headphones, and ask a colleague to listen to anything that still sounds ambiguous. Leave a tag rather than a guess if nobody can agree.

6

Save a master, then derive copies

Keep the reviewed verbatim transcript as your master file. If you need a clean or anonymized version for coding, reporting or publication, create it as a copy, so you can always return to exactly what was said.

Common verbatim transcription mistakes

  1. Not choosing a style up front. Switching between full and clean verbatim halfway through a project makes transcripts impossible to compare.
  2. Cleaning away meaning. Hedges like "I guess" or "sort of", or a long hesitation, can be evidence. When in doubt, keep them: deleting later is easier than going back to the audio.
  3. Correcting the speaker's voice. Dialect, slang and grammar belong to the speaker. Clean verbatim removes noise, not personality.
  4. Guessing silently. An unmarked guess looks exactly like a verified word. Use [inaudible 00:12:31] or [word?] every time.
  5. Trusting speaker labels blindly. A quote attributed to the wrong person is worse than a typo.
  6. Quoting from an edited transcript. Edited text is a paraphrase; go back to the verbatim master or the audio.
  7. Over-notating. Jefferson symbols in a transcript meant for thematic coding or a meeting record only slow readers down.

Sources and further reading

Frequently asked questions

What is verbatim transcription?
Verbatim transcription is a word-for-word written record of what was said in a recording. In its strictest form, full verbatim, it keeps every filler, false start, repetition and audible non-verbal cue such as laughter. Clean verbatim (also called intelligent verbatim) keeps the speaker's words and meaning but removes the fillers and stumbles. Edited transcription goes further and rewrites for readability, so it is no longer verbatim.
What is the main difference between clean verbatim and full verbatim?
Full verbatim records how something was said: every "um", stutter, false start, repetition, pause and non-verbal sound. Clean verbatim records what was said: the same words in the same order, minus the fillers and stumbles that carry no meaning. Full verbatim suits conversation analysis and legal or evidential uses; clean verbatim suits thematic analysis, business interviews and published quotes.
What is not included in full verbatim?
Full verbatim includes every spoken word and audible sound, but it is still plain text: it does not normally record intonation, stress, speed or exact pause lengths. That detail belongs to specialist notation such as the Jefferson system used in conversation analysis. Full verbatim also never corrects grammar or adds words the speaker did not say.
What does a verbatim transcript look like?
A verbatim transcript is laid out as turns of talk, each starting with a speaker label (for example Interviewer or P03) and often a timestamp such as [00:04:12]. Square brackets mark anything that is not a word: [laughs], [crosstalk], [pause] or [inaudible 00:12:31]. In full verbatim, fillers like "um" and cut-off words like "bett-" stay in the text.
Is it okay to use non-verbatim transcription?
"Non-verbatim" usually means clean verbatim, and sometimes an edited transcript. Clean verbatim is fine when you care about the content rather than the way it was spoken, as with meeting notes or most business interviews. Avoid edited or summarized text for legal or evidential material, for conversation or discourse analysis, and for any quote you publish as someone's exact words without checking the audio.
What does verbatim mean in legal terms?
In legal settings, verbatim means an exact record of what was said, with no summarizing or paraphrasing. In US federal courts, for example, 28 U.S.C. § 753 requires court sessions to be recorded verbatim, by shorthand, mechanical means, electronic sound recording or another approved method. A transcript for official court use has to come from the process your court recognizes, not from an unreviewed AI draft.
Can AI do verbatim transcription?
AI gives you a fast first draft that often reads closer to clean verbatim than to full verbatim: fillers, false starts and overlaps may be smoothed out, and non-verbal cues are rarely marked. To reach true full verbatim, someone has to listen back, restore what was dropped, mark inaudible passages and overlaps, and check the speaker labels.

Start from a first draft, not a blank page

Upload your recording, get a timestamped transcript with the voices separated, then review it into the verbatim style you need. 120 free minutes per month, no credit card, hosted in Europe.

Start for free