If you’ve been asked to produce a verbatim transcript, or you’re reviewing one a transcriptionist delivered, the real question usually isn’t what verbatim transcription is. It’s what you’re actually supposed to type when someone stutters, trails off mid-sentence, or two people start talking over each other.
This guide covers the practical rules for typing a verbatim transcript: what to capture, how to represent it, and how to format the result, with real audio-to-text examples for the most common situations.
Verbatim transcription means typing what’s said as it’s said, not just the general idea. That includes the words themselves, plus filler words, false starts, stutters, repetitions, pauses, and relevant nonverbal sounds like laughter or a cough. The rules below describe full, or true, verbatim: the most detailed version of a transcript. Where a lighter “clean verbatim” style would render something differently, we’ll note it.
Full verbatim means transcribing every word spoken, not a paraphrase or a cleaned-up version of it. Skipping a word is an editorial decision about what mattered, and that’s not the transcriptionist’s call to make on a verbatim project. In practice, this means typing exactly what’s said, including grammatically incorrect sentences and incomplete thoughts, rather than smoothing it into something that reads more naturally.
Words like “um,” “uh,” “like,” and “you know” get left out of most other transcript styles, but in verbatim work they stay in. Fillers can signal hesitation, uncertainty, or emphasis, which matters when the transcript is used to analyze how someone communicates, not just what they said. Type them exactly where they occur: “I, um, I think that’s fine.”
A stutter or stumble over a word is typed as it sounds, not corrected. If a speaker says “I th-th-think so,” that’s what goes in the transcript. This matters most in legal and clinical contexts, where how someone spoke can be as relevant as what they said.
When a speaker repeats a word or phrase, whether out of habit, emphasis, or hesitation, the repetition is typed in full rather than trimmed to a single instance. “I really, really think we should wait” stays exactly that way, not “I really think we should wait.”
A false start is a sentence or thought a speaker begins and then abandons or restarts. The abandoned fragment stays in the transcript, typically marked with an ellipsis (some styles use a double hyphen instead), followed by the sentence that replaces it. “I don’t... I mean, I don’t think that’s right” shows how the thought actually developed, not just where it landed.
Short pauses are usually shown with an ellipsis inside a sentence. Longer, more meaningful pauses, several seconds or more, are typically noted separately, for example (pause) or (long pause), sometimes with an approximate duration if it’s relevant. The convention can vary by client, but the underlying rule stays the same: a pause that changes how a statement reads should be visible in the transcript, not silently closed up.
Laughter, sighs, coughs, and similar sounds are noted in brackets when they add meaning or context, for example [laughs] or [sighs]. Not every throat-clear needs to be logged. The judgment call is whether the sound changes how a statement should be read: a sarcastic laugh after a comment changes its meaning, while a mid-sentence cough usually doesn’t.
Every time the speaker changes, the transcript should clearly show it, typically with a speaker label followed by a colon, such as “Interviewer: How did that go?” Consistency matters more than which exact labeling convention you use. Decide on a format, names, roles, or Speaker 1/Speaker 2, before you start, and apply it the same way throughout.
When two people talk at the same time, note it rather than picking one thread and dropping the other. Common conventions include a bracketed note like [overlapping] on each speaker’s line involved, or bracketing the overlapping portion of each line. What matters is that a reader can tell overlapping speech happened, rather than assuming one person simply interrupted and stopped.
When a word or phrase genuinely can’t be made out, don’t guess. Mark it clearly, for example [inaudible] or [inaudible 00:12:47] with a timestamp so it can be revisited later. If you’re fairly confident but not certain, a flagged guess like [unclear: possible word] notes the uncertainty instead of presenting it as fact.
Timestamps help a reader or reviewer match a line of text back to the exact moment in the recording. Common approaches include a timestamp at regular intervals, such as every one or two minutes, at each speaker change, or at specific markers like inaudible sections. Whichever approach you choose, apply it the same way through the entire transcript.
All of the rules above only work if they’re applied the same way from the first line to the last. Pick your conventions for speaker labels, brackets, timestamps, and pause notation before you start transcribing, and stick to them. A transcript that mixes formatting styles partway through is harder to read and harder to trust.
Here’s how some of the most common speech patterns actually look once they’re typed out.
Audio:
"I th-th-think we should wait."
Verbatim transcript:
Participant: I th-th-think we should wait.
The stutter is typed exactly as it sounds rather than smoothed into “I think we should wait.”
Audio:
"So, um, I guess we could look at it tomorrow."
Verbatim transcript:
Participant: So, um, I guess we could look at it tomorrow.
The “um” stays in. In a clean or intelligent verbatim version, it would typically be dropped: “So I guess we could look at it tomorrow.”
Audio:
"I really, really don't think that's a good idea."
Verbatim transcript:
Participant: I really, really don't think that's a good idea.
The repetition is kept in full rather than trimmed to a single “really.”
Audio:
"I don't... I mean, I don't think we should go."
Verbatim transcript:
Participant: I don't... I mean, I don't think we should go.
The abandoned fragment, “I don’t,” stays in, marked with an ellipsis, showing how the thought actually developed.
Audio:
[5-second pause] "I'm not sure how to answer that."
Verbatim transcript:
Participant: (pause) I'm not sure how to answer that.
A pause long enough to carry meaning, hesitation, discomfort, thinking time, is marked rather than closed up as if the speaker answered immediately.
Audio:
[laughs] "No, I don't think so."
Verbatim transcript:
Participant: [laughs] No, I don't think so.
The laugh is included because it changes how the statement reads, more amused than defensive.
Audio:
[sighs] "It's been a long week."
Verbatim transcript:
Participant: [sighs] It's been a long week.
The sigh adds emotional context that a transcript of the words alone would miss.
Audio:
Two participants begin speaking at the same time.
Verbatim transcript:
Participant A: I think we should [overlapping] Participant B: [overlapping] wait, can I say something?
Both lines are marked as overlapping so a reader knows the two statements happened at the same time, rather than assuming one speaker simply cut the other one off cleanly.
Audio:
A word is lost under background noise.
Verbatim transcript:
Participant: I think the [inaudible 00:14:22] was the real issue.
Rather than guessing at the missing word, it’s marked as inaudible with a timestamp so it can be checked against the recording later.
Audio:
A word is mumbled but a likely match can be made out.
Verbatim transcript:
Participant: I think it was [unclear: Tuesday] when we spoke.
Flagging the best guess as uncertain is more accurate than presenting it as fact or leaving a blank.
Audio:
A moderator asks a question and a participant answers.
Verbatim transcript:
Moderator: What made you decide to apply? Participant: Honestly, um, a friend told me about it.
A clear label at each speaker change keeps the exchange easy to follow, even before any names are confirmed.
Verbatim transcription has its own notation on top of standard transcript formatting. The core conventions:
These verbatim-specific conventions sit on top of general transcript formatting, margins, headers, and file structure. For the full formatting picture, see how to format a transcript.
Multiple speakers talking at once. In a group interview, panel, or meeting, more than two people can end up talking simultaneously. Mark each overlapping stretch on every speaker’s line involved, rather than just the first or loudest voice, so the transcript reflects that a genuine pile-up happened.
Verbatim transcription isn’t the right choice for every project. It takes longer to produce and can be harder to read than a cleaned-up transcript. It tends to matter most when the exact wording, hesitation, or delivery of a statement carries meaning; legal proceedings and quoting a source directly are common examples. For a fuller breakdown of when verbatim transcription is appropriate and how it compares to other transcription styles, see our guide on the topic.
Researchers often need verbatim transcripts specifically because speech patterns, pauses, and exact wording can be part of what’s being analyzed, not just background detail. A hesitation before answering a sensitive question, or a phrase a participant repeats for emphasis, can be meaningful data in coding and thematic analysis. For more on verbatim transcription in qualitative research, see our dedicated guide.
Get 100% Human-Powered Verbatim Transcripts With A 99% Accuracy Guarantee.
Every rule above matters more when the transcript itself matters more: a deposition, a research interview that will be coded and quoted, an HR investigation. GMR Transcription produces verbatim transcripts through a 100% human transcription process, so pauses, false starts, overlapping speech, and other verbatim-specific details are captured accurately rather than smoothed over or missed. Contact GMR Transcription to get a project transcribed to these standards.
Filler words like “um,” “uh,” and “like” are typed exactly as they’re spoken, in the position where they occur, rather than removed. They’re left out only in non-verbatim or cleaned-up transcript styles.
A stutter is typed as it sounds, repeating the stuttered sound or syllable rather than smoothing it into the full word, for example, “I th-th-think so.”
Overlapping speech is typically marked with a bracketed note like [overlapping] on each speaker’s line involved, so a reader can tell the statements happened at the same time rather than in sequence.
Mark it clearly, for example [inaudible] or [inaudible 00:12:47] with a timestamp, rather than guessing at what was said. If you’re fairly confident but not certain, a flagged guess like [unclear: word] can be used instead.
A short pause is usually shown with an ellipsis within a sentence. A longer, meaningful pause is typically noted separately, for example (pause) or (long pause), depending on the transcript’s formatting convention.