Short answer: full verbatim keeps everything you can hear (fillers, false starts, stutters, repetitions, laughter, overlaps). Clean verbatim, also called intelligent verbatim, keeps the words and meaning but drops the stumbles. Edited transcription rewrites for readability and is no longer verbatim. Use full verbatim when how something was said matters, clean verbatim when what was said is enough.
What is verbatim transcription?
Verbatim transcription means turning speech into text exactly as it was spoken, word for word, without summarizing or paraphrasing. The transcript follows the speaker's own wording and word order, even when that wording is informal, ungrammatical or unfinished.
The complication is that "exactly as spoken" can mean two different things. Real speech is full of material that never makes it into written language: hesitations, restarts, repeated words, laughter, people talking over each other. Whether that material belongs in the transcript is the whole question behind the terms full verbatim, clean verbatim and "verbatim vs non-verbatim". Researchers have argued about it for decades. A widely cited paper by Oliver, Serovich and Mason (2005, Social Forces) describes the two ends of the spectrum: one where "every utterance is transcribed in as much detail as possible", and one where idiosyncratic elements such as stutters, pauses and non-verbal sounds are removed.
Neither end is "right": each style discards some information to make the text useful for a particular job. The real question is what you need to keep.
The three transcription styles: full, clean and edited
Many transcription providers distinguish three levels, from the most literal to the most polished. The names vary between providers and universities, so check what each one actually does rather than trusting the label.
Full verbatim (true or strict verbatim)
Full verbatim captures every audible word and sound: fillers ("um", "uh", "er"), false starts, stutters, repetitions, backchannels ("mm-hmm", "yeah") and audible non-verbal events such as laughter, sighs or a door slamming. Grammar is left exactly as spoken. Overlapping speech, pauses and passages you cannot hear are marked with tags. It is the slowest style to produce and the hardest to read, but nothing is silently lost.
Clean verbatim (intelligent verbatim)
Clean verbatim keeps the speaker's words, word order and meaning, but removes the noise: fillers, stutters, false starts and repetitions that add nothing. Some style guides also allow light corrections, such as fixing an obvious slip of the tongue, while keeping dialect and slang. It reads naturally and is quick to scan, which is why it is the default for most interviews, meetings and published quotes. You will sometimes see it called "non-verbatim", although the words are still the speaker's own.
Edited transcription
Edited transcription (sometimes called a "clean read" or "polished" transcript) goes a step further: the transcriber fixes grammar, tightens rambling sentences and may reorder phrases so the text reads like prose. It is ideal for blog posts, reports or show notes, but it is an interpretation, not a record. Never quote an edited transcript as someone's exact words.
Verbatim transcription example: one passage, three styles
The easiest way to see the difference is to transcribe the same exchange three times. The dialogue below is invented for illustration: an interviewer asks a participant (coded P03) about a new shift schedule at work.
[00:04:12] Interviewer: So, um, how did the new rota go for, for your team? [00:04:17] P03: Uh, well, it was, it was- honestly? At first it was a mess. [laughs] Like, n- nobody knew who was, uh, covering the Friday shifts, you know? [00:04:29] Interviewer: Mm-hmm. [00:04:30] P03: And then, um [pause] after the second month or so it kind of, it kind of settled down, and [crosstalk] [00:04:38] Interviewer: [crosstalk] So it got bett- [00:04:39] P03: Yeah, yeah, it got better. We ain't going back to the [inaudible 00:04:42], put it that way.
[00:04:12] Interviewer: How did the new rota go for your team? [00:04:17] P03: Honestly? At first it was a mess. [laughs] Nobody knew who was covering the Friday shifts. [00:04:30] P03: Then, after the second month or so, it settled down. [00:04:38] Interviewer: So it got better? [00:04:39] P03: Yeah, it got better. We ain't going back to the [inaudible 00:04:42], put it that way.
Interviewer: How did the new rota work for your team? P03: At first it was chaotic: nobody knew who was covering Friday shifts. After about two months it settled down, and the team would not go back to the old system.
Look at what changed. The clean version drops "um", "uh", "you know", the stutter ("n- nobody") and the repeated "it was, it was", and it tidies the overlap into a plain question. It still keeps "Honestly?", the laugh and "ain't", because they carry tone and they are the participant's own voice. The edited version is shorter and easier to read, but "a mess" has become "chaotic", "we ain't going back" has become "the team would not go back to the old system", and the inaudible word has been filled in with a guess. That is fine for a report summary; it is not a quote.
What each style keeps or removes
Here is how the three styles treat the elements that most often cause disagreement between transcribers. House styles differ at the edges, so treat this as the common baseline and write down your own choices.
| Element | Full verbatim | Clean verbatim | Edited |
|---|---|---|---|
| Fillers (um, uh, er, like, you know) | Kept | Removed | Removed |
| False starts ("it was- honestly") | Kept, with a hyphen at the cut-off | Removed unless they change meaning | Removed |
| Stutters ("n- nobody") | Kept | Removed | Removed |
| Repetitions ("for, for") | Kept | Removed unless used for emphasis | Removed |
| Backchannels (mm-hmm, yeah) | Kept as separate turns | Usually removed, kept if they answer a question | Removed |
| Non-verbal cues ([laughs], [sighs]) | Kept | Kept only when they change meaning | Usually removed |
| Overlapping speech | Marked with [crosstalk] and transcribed where audible | Simplified into clear turns | Merged or dropped |
| Grammar and slang | Exactly as spoken | As spoken; obvious slips may be fixed | Corrected |
| Exact quote? | Yes | Yes, of the words that matter | No |
Verbatim notation: timestamps, speaker labels and tags
Whatever the style, a transcript needs a small, consistent set of conventions so readers know what the brackets mean. There is no single universal standard for everyday transcription; the set below is a common practical baseline. Pick one, write it down, and stick to it.
| Notation | What it means | When to use it |
|---|---|---|
| Interviewer: / P03: | Speaker label at the start of each turn | Always. Use pseudonyms or codes if participants must stay anonymous |
| [00:04:12] | Timestamp (hours:minutes:seconds) | At each turn, or at a fixed interval, so you can jump back to the audio |
| [inaudible 00:12:31] | A word or passage you cannot make out, with the time it occurs | Every time. Never guess silently |
| [rota?] | A best guess at an unclear word | When you are fairly, but not fully, sure |
| [crosstalk] | Two or more people speaking at once | Full verbatim, and in clean verbatim when overlap hides content |
| [laughs] [sighs] [coughs] | Audible non-verbal events | Full verbatim; clean verbatim when they change meaning |
| [pause] | A noticeable silence | Full verbatim, or anywhere the hesitation matters |
| bett- | A word cut off mid-way | Full verbatim |
When plain tags are not enough: Jefferson notation
Conversation analysts need far more detail than "[pause]". The standard system in that field was developed by Gail Jefferson; her Glossary of transcript symbols (2004), published in Conversation Analysis: Studies from the first generation, defines symbols such as [ for the point where overlapping talk begins, = for no gap between turns, (0.5) for a pause timed in tenths of a second, (.) for a very brief interval, colons for a stretched sound, empty parentheses ( ) for talk the transcriber could not make out, and doubled parentheses (( )) for the transcriber's own descriptions. It is a specialist skill; for most interviews and meetings, the tags above are plenty. Our guide to transcription in qualitative research shows a lightly notated research excerpt if you want to see the two approaches combined.
Which verbatim style should you choose?
Choose from the use of the transcript, not from habit: the more your analysis or legal exposure depends on the way people spoke, the closer you stay to full verbatim.
Qualitative research
For conversation analysis, discourse analysis and other methods that study talk itself, use full verbatim, often with Jefferson notation. For thematic analysis and most interview studies, clean verbatim is usually enough, as long as you keep hesitations that carry meaning. Be careful with the jargon: the labels "naturalized" and "denaturalized" are not used consistently. Oliver and colleagues use "naturalized" for the detailed, every-utterance style, while McMullin (2021, Voluntas) pairs "denaturalized" with full verbatim and "naturalized" with intelligent verbatim. Describe what your transcripts keep in your methods section rather than relying on the label. The University of Bath research data guidance on transcription is a good checklist for those decisions. Focus groups are the hardest case, with more voices and far more overlap; we cover that in our guide to focus group transcription.
Legal and compliance
In legal contexts, "verbatim" means an exact record, with no paraphrase. In US federal courts, 28 U.S.C. § 753 requires court sessions to be "recorded verbatim by shorthand, mechanical means, electronic sound recording, or any other method", subject to Judicial Conference regulations and the judge's approval. Official court transcripts come from that certified process, not from your own transcription. For your own evidential recordings (an investigation interview, a complaint call, a disciplinary hearing), keep full verbatim, mark every inaudible passage and never clean the master copy.
Journalism
A clean transcript is the practical way to find the story, but check every quote against the audio before publishing. Anything that shifts meaning or tone (a sarcastic "great", a long hesitation, a retracted word) has to be heard, not read. Your publication's style guide decides how far a quote may be tidied.
Business meetings and interviews
For meetings, hiring interviews, customer calls and user research, clean verbatim is usually the right choice: readable, searchable, and exact where decisions and commitments are made. Many teams circulate an edited summary and keep the clean transcript as the reference. If you need to transcribe your interviews with separated speakers and timestamps, a dedicated interview transcription workflow saves you most of the typing.
Subtitles and captions
Captions sit between the two styles. The US rule for closed captioning of televised programming, 47 CFR § 79.1, asks captions to match the spoken words in the order spoken "without paraphrasing, except to the extent that paraphrasing is necessary to resolve any time constraints", to mirror slang or grammatical errors that are intentional in the dialogue, and to convey non-verbal information such as the identity of speakers, music, sound effects and audience reaction. Even where that rule does not apply, it is a sound benchmark for subtitles: stay close to verbatim and keep sound cues.
How AI transcription handles verbatim
Automatic speech recognition gives you a fast first draft, but do not assume that draft is full verbatim. Engines generally aim for readable text, so their output often lands closer to clean verbatim: fillers may be dropped or kept inconsistently, stutters and false starts smoothed out, non-verbal cues rarely marked, and overlapping speech merged into a single turn or lost. Behavior varies by tool, so test yours on a short sample first.
McMullin notes that speech recognition may now give researchers "good enough" first drafts, but that voice-to-text software "is also generally less accurate in discerning multiple voices or different accents", and that such transcripts still need to be carefully checked by the researcher. In practice, AI saves you the typing; the listening is still yours.
What typically needs a human pass:
- Fillers, stutters and false starts left out by the engine, if you need full verbatim
- Non-verbal cues: laughter, sighs, long pauses
- Overlapping speech, often the weakest part of an AI draft
- Inaudible passages, which an engine may fill with a plausible guess instead of flagging
- Names, numbers, jargon and negations, where one wrong word flips the meaning
- Speaker labels, especially with similar voices or short interjections
Where AudiosTranscribe fits: you upload a recording (or record it with the Windows desktop app) and get a timestamped transcript in which the voices are separated automatically; you then put a name on each speaker in one click. Treat it as a first draft to review against the audio, not as a finished full-verbatim record. The service is hosted in Europe, GDPR-compliant, and audio is deleted after processing. The free plan includes 120 minutes per month, and TXT, Word, PDF and SRT exports come with the paid plans.
How to review an AI transcript into true verbatim
Turning an AI draft into reliable verbatim is a structured listening job. Six steps keep it fast and consistent.
Write your style sheet first
Decide full or clean verbatim, and list your tags ([inaudible 00:00:00], [crosstalk], [laughs], [pause]), speaker label format and timestamp interval. One short page is enough, and it keeps several reviewers consistent.
Fix the speakers before the words
Name each voice, then skim the whole transcript for turns attributed to the wrong person, especially short replies like "yes" or "mm-hmm". Correcting labels first stops you from re-reading everything later.
Listen through for the words
Play the audio at a slightly reduced speed and follow the text. Restore fillers, stutters, false starts and repetitions if you need full verbatim, and correct misheard words, names, numbers and negations in any style.
Add the tags
On the same pass or a second one, mark non-verbal cues, pauses, overlapping speech and every passage you cannot make out, with its timestamp. Replace any word the engine invented over unclear audio with [inaudible] or a bracketed best guess.
Re-check the hard spots
Go back to each flagged timestamp with headphones, and ask a colleague to listen to anything that still sounds ambiguous. Leave a tag rather than a guess if nobody can agree.
Save a master, then derive copies
Keep the reviewed verbatim transcript as your master file. If you need a clean or anonymized version for coding, reporting or publication, create it as a copy, so you can always return to exactly what was said.
Common verbatim transcription mistakes
- Not choosing a style up front. Switching between full and clean verbatim halfway through a project makes transcripts impossible to compare.
- Cleaning away meaning. Hedges like "I guess" or "sort of", or a long hesitation, can be evidence. When in doubt, keep them: deleting later is easier than going back to the audio.
- Correcting the speaker's voice. Dialect, slang and grammar belong to the speaker. Clean verbatim removes noise, not personality.
- Guessing silently. An unmarked guess looks exactly like a verified word. Use [inaudible 00:12:31] or [word?] every time.
- Trusting speaker labels blindly. A quote attributed to the wrong person is worse than a typo.
- Quoting from an edited transcript. Edited text is a paraphrase; go back to the verbatim master or the audio.
- Over-notating. Jefferson symbols in a transcript meant for thematic coding or a meeting record only slow readers down.
Sources and further reading
- Oliver, D. G., Serovich, J. M. and Mason, T. L. (2005). Constraints and Opportunities with Interview Transcription: Towards Reflection in Qualitative Research. Social Forces, 84(2), 1273-1289.
- McMullin, C. (2021). Transcription and Qualitative Methods: Implications for Third Sector Research. Voluntas, 34(1), 140-153.
- Jefferson, G. (2004). Glossary of transcript symbols with an introduction. In G. H. Lerner (ed.), Conversation Analysis: Studies from the first generation. John Benjamins.
- University of Bath Library. Transcription (Working with data).
- 28 U.S. Code § 753, Reporters (Legal Information Institute, Cornell Law School).
- 47 CFR § 79.1, Closed captioning of televised video programming (Electronic Code of Federal Regulations).