A quality transcript isn't just accurate, it's easy to scan, consistent from start to finish, and formatted to match how it will actually be used, whether that's a research file, a legal record, or a published interview. This guide covers exactly how to format one: choosing a transcript style, applying speaker labels and timestamps correctly, formatting in Microsoft Word, and citing a transcript in APA style.
Transcript formatting is the set of conventions that determine how a written transcript is structured: how speakers are labeled, how timestamps are applied, how pauses and inaudible sections are marked, and how the document is laid out for readability. The right transcript format depends on what the transcript is for, a legal deposition, a research interview, or a video upload each call for different conventions, but every good transcript follows a format that's consistent from the first line to the last.
Before applying any formatting, decide which of the three common transcript styles the project calls for.
| Style | What it captures | Best for |
|---|---|---|
| Full verbatim | Every word and sound, including fillers, stutters, false starts, repeated speech, and nonverbal noises | Legal proceedings, research where speech patterns matter |
| Semi-verbatim | A middle ground, removes distracting filler and nonverbal noise while keeping the substance of what was said | Interviews and meetings where full detail isn't required but some nuance should stay |
| Intelligent/clean verbatim | Polished standard English, fillers, repetitions, and grammatical errors removed | Business content, publications, anything meant to be read rather than analyzed word by word |
Full verbatim replicates everything in the recording, slang, filler words like "you know," repeated speech, stutters, false starts, and nonverbal noises like a throat clearing or a chair creaking. It's the most detailed transcript style and the standard for legal transcription and research where wording and delivery are part of the data. For a closer look at exactly how to type one, including examples for stutters, fillers, and overlapping speech, see how to type a verbatim transcript.
Semi-verbatim sits between full verbatim and clean verbatim. It removes some of the nonverbal detail and extraneous sounds that can distract from the substance of an interview, without fully polishing the language the way clean verbatim does.
Intelligent, edited, or clean verbatim is the most polished style. It's written in standard English, never captures repetitions or slang, and is fully edited for grammar and punctuation, with irrelevant words removed. It's a common choice for business meetings and published content where readability matters more than capturing every verbal tic.
Whichever style you choose, apply it consistently throughout the document rather than mixing styles partway through. For guidance on which style fits a given project, especially in a research context, see when verbatim transcription is needed.
These are the core transcript formatting guidelines that apply regardless of which style you're working in.
Label each speaker by name or role. If you can't identify who's talking, use a generic label like "Speaker 1." Whichever convention you choose, keep the same label format for that speaker throughout the transcript, don't switch between "Peter" and "Speaker 1" for the same person. Add a colon after the label, then a space, then the transcribed text:
Peter: Hello. Speaker 1: Hello, Peter. Host: Hello.
There are two standard approaches to a transcript timestamp format:
Regular interval insertion, where a timestamp appears after a uniform period, such as every two minutes:
Nancy: This is a timestamp [02:00] example. Tom: Another timestamp [04:00] example.
Timestamp at each speaker change, where a timestamp appears at the start of a new paragraph or whenever the speaker changes:
[02:30] Nancy: Thank you, Tom. [02:34] Tom: You're welcome.
Timestamps matter most for longer recordings, research interviews, and legal transcripts, where someone needs to locate a specific moment in the audio later. Keep the format identical throughout a single transcript.
Mark any section you can't confidently make out as inaudible rather than guessing, and mark two people talking at the same time as crosstalk. Include a timestamp with each:
[inaudible 03:14] [crosstalk 02:28]
Note non-speech sounds in square brackets when they're relevant to the recording, for example [door closing] or [phone ringing]. Skip incidental noise that doesn't add context.
Use U.S. spelling for American English content and UK spelling for British English content, and apply that choice consistently. Check spelling against the language you've chosen, not just what sounds right.
Follow standard grammar rules: capitalize the first letter of names, places, organizations, and job titles when they precede a name. If the project specifies a style guide like APA or MLA, follow that guide's capitalization rules for headings and titles.
Break long speech into paragraphs, roughly 400 to 500 characters is a reasonable guideline, rather than running it as one continuous block. Add headings, subheadings, and page numbers for longer transcripts so readers can navigate the document without scanning line by line.
Times New Roman and Calibri at 11 or 12 points are common, readable choices for a transcript formatted in a word processor. Single or 1.5 line spacing both work; whichever you choose, keep it consistent across the whole document.
Here's a short transcript template that puts several of these conventions to work together, use it as a starting model for your own transcripts:
[00:00:04] Interviewer: Thanks for joining me today. Can you tell me about your role? [00:00:09] Participant: Sure. I've been the project lead for about two years now. [00:00:15] Interviewer: And what's changed the most in that time? [00:00:18] Participant: Honestly, [inaudible 00:20] the team size, mostly. [00:00:24] Participant: We started with three people. [door closing] Sorry, someone just came in. [00:00:29] Interviewer: No problem. You were saying, about the team? [00:00:32] Participant: Right, we're at twelve now.
Notice the consistent speaker labels, the timestamp-at-speaker-change format, the bracketed inaudible tag with its own timestamp, and the bracketed background sound, all applied the same way throughout.
Microsoft Word doesn't enforce a single required transcript format, but a common structure includes a title page with the source file name, media file name or ID, the duration of the recording, and the date it was transcribed. The body is typically set in a standard, readable font (Times New Roman or Calibri at 11 or 12 points work well), with page numbers in the footer for longer documents.
Word for the web also includes a built-in Transcribe feature under some Microsoft 365 plans, which can generate a first-pass transcript with automatic speaker labels and timestamps. Treat that output as a draft: automated transcripts still need to be checked against the original recording for accuracy, speaker identification, and correct handling of unclear audio before you'd consider them finished.
Interview transcripts follow the same core rules above, with a few things worth setting up before you start:
For a full breakdown of how to represent stutters, false starts, overlapping speech, and unclear audio in an interview transcript, see how to type a verbatim transcript.
The APA rule that matters most here is whether the interview is recoverable, meaning whether someone else could actually access it.
Personal interviews you conducted yourself (a research interview, a private conversation) are treated as personal communications. They're cited in-text only and are not included in the reference list, because there's no way for a reader to retrieve and verify them. The in-text format is:
(F. Last Name, personal communication, Month Day, Year)
Published or broadcast interviews that readers can actually access, a podcast episode, a YouTube video, a published transcript, a newspaper interview, are recoverable sources. These get both an in-text citation and a full reference list entry, formatted according to the type of source they appeared in. For example, an interview published as a podcast episode is cited using the podcast episode reference format, with the host listed as the author:
Host Last Name, F. M. (Host). (Year, Month Day). Episode title [Audio podcast episode]. Podcast Name. URL
Because citation guidance can be updated, confirm the current rule against APA Style's official page on personal communications before finalizing a reference list, particularly for edge cases like interviews posted on social media or intranet resources.
Formatting a transcript by hand is manageable for a short recording, but it gets time-consuming fast on longer projects, especially ones that require consistent speaker identification, precise timestamps, a specific verbatim style, or custom formatting to match a template. GMR Transcription's transcripts are produced by 100% U.S.-based human transcriptionists, who can apply the formatting conventions in this guide, and any project-specific requirements, consistently across a document. Explore transcription services or contact GMR Transcription to discuss formatting for your specific project.
There isn't one universal standard, the right format depends on the transcript's purpose. Most transcripts share a few conventions: consistent speaker labels, a defined timestamp format, marked inaudible sections, and consistent paragraph breaks and spacing throughout.
Use the speaker's name or role if known, or a generic label like "Speaker 1" if not, followed by a colon and a space before their words. Keep the same label for each speaker throughout the entire transcript.
Either at a regular interval (such as every two minutes) or at each speaker change, both are standard approaches. Whichever you choose, apply it the same way throughout the document.
Use a bracketed tag with a timestamp, such as [inaudible 03:14], rather than guessing at the word.
It depends on whether the interview is recoverable by other readers. A personal interview you conducted is cited in-text only as a personal communication and left out of the reference list. A published or broadcast interview, like a podcast or YouTube video, gets a full reference list entry formatted according to that source type.
Either can work; what matters most is picking one and applying it consistently. Single or 1.5 line spacing is common for readability in longer transcripts.