How to Record and Transcribe a D&D Session on Discord
The full workflow for an online game: what to record with, why per-speaker tracks make such a difference, how to handle the player whose mic is a disaster, and what your options actually are for turning four hours of audio into something you can search.
· 7 min read
What to use to record a Discord D&D session
If your game runs in a Discord voice channel, the usual answer is Craig — a recording bot you invite to your server that joins the voice channel and records each person to their own separate audio track. It is free to use, it has been a mainstay for tabletop groups for years, and the multi-track output is the single most useful thing you can have when it comes time to transcribe.
Set it up once and the per-session routine is about thirty seconds of work: run the join command before you start, run the stop command when you finish, then download the file from the link Craig sends you.
Everything below is the detail — consent, formats, mic problems, alternatives if Craig will not work for your group, and what to do with the audio afterwards.
Why a screen recorder is the wrong tool for this
Discord's own client is not built for this — there is no reliable built-in way to record a voice call and end up with usable audio. What people reach for instead is screen sharing plus a screen-capture tool, or a general desktop recorder pointed at the Discord window. Both work in the sense that you end up with a file. Both are the wrong shape for the job.
A screen or desktop capture gives you one mixed stream: everyone's voice, your music bot, your notification pings, and whatever else your speakers were doing, all flattened into a single track you can never pull apart again. When two players talk over each other — which is most of a combat round — the transcription has to guess at overlapping speech, and it guesses badly.
You also end up carrying a video file around for audio you actually want, which makes it larger, slower to upload, and more annoying to archive.
The other problem is reliability. A screen recorder runs on your machine, which means it dies when your machine does. A bot sitting in the voice channel is recording server-side: if your client crashes, it keeps capturing everyone else and picks you back up when you rejoin, instead of losing the whole session.
How to record a session with Craig, step by step
Craig lives at craig.chat. You invite it to your Discord server the same way you invite any bot, which requires permission to add bots to that server — if it is not your server, ask whoever runs it.
Once it is in, the flow per session is:
- Invite Craig to your server ahead of time, not five minutes before the game. Do the setup on a quiet afternoon so session night is just a command.
- Join the voice channel your group plays in, then run Craig's join command. Type a forward slash in any text channel and Discord will show you the commands the bot accepts, so you do not have to memorize the exact wording.
- Confirm it is actually recording before you start playing. Craig confirms when it has joined and started capturing; check that the bot is present in the voice channel list and say a test line.
- Play your session. Craig records each speaker to a separate track as they talk. People joining late get their own track too.
- Run the stop command when you are done. Craig will send you a link — usually by direct message — to a download page for that recording.
- Download it promptly. Recordings are kept for a limited window, not forever, and the exact window depends on the bot's current policy and whether you support the project. Treat the link as time-limited and pull the file down the same night or the next morning.
Why per-speaker tracks matter so much for transcription
This is the reason to bother with a bot at all. Craig records each participant's audio stream individually, so a five-person table produces five separate files that all start at the same moment.
With one mixed track, crosstalk is unrecoverable. When your rogue interrupts your cleric mid-sentence, the transcription model hears one muddy waveform and produces one mangled line, or drops the quieter speaker entirely. Every D&D table does this constantly — the reactions, the jokes, the three people all calling out what they want to do.
With separate tracks, there is nothing to untangle. Each file contains one clean voice with silence in between. Transcription accuracy goes up, and you get speaker attribution for free, because the file name tells you who was talking.
It also means one bad audio source is contained. A player with a buzzing mic damages exactly one track instead of contaminating the whole recording.
Which download format to pick
Craig's download page offers several options: lossless multi-track audio, smaller compressed multi-track audio, a single mixed-down track, and formats meant to open directly in an audio editor.
A reasonable default is to grab two things. Take a multi-track download as your archive copy — that is the version with the most information in it, and you cannot recreate it later once the link expires. Then take a compressed or single-track version as your working copy for whatever you are going to transcribe with.
Size matters more than people expect. Lossless audio of a four-hour session across five tracks gets large fast, and most transcription services have an upload limit. Compressed audio at a sensible bitrate for speech makes little practical difference to a transcription model and is a fraction of the size.
If you plan to feed the audio to a tool that takes one file at a time, the mixed-down single track is the pragmatic choice — you lose per-speaker separation, but you skip the work of stitching five transcripts back together in timestamp order. If your tool accepts multiple files, or you are running transcription yourself, keep them separate.
One thing to watch: a multi-track download usually arrives as an archive containing the individual tracks, not as a single playable file. Unzip it before you try to upload it anywhere.
Ask your players before you record anything
Do this. It is not a formality and it is not optional socially.
Say it out loud at the top of the session, in plain terms: you want to record the audio so you can write recaps and keep track of the campaign, it is stored privately, it is not going on the internet, and anyone who is not comfortable can say so now. Then actually accept a no.
Craig posts a message when it starts recording, and players can see the bot sitting in the voice channel list, so a recording is not a secret in practice. Announcing it anyway is what keeps people relaxed. A table that has forgotten there is a bot in the channel talks freely; a table that discovers a recording afterwards does not talk freely again.
There is also a legal dimension depending on where everyone lives, since your players may be in different jurisdictions with different rules about recording conversations. Getting explicit agreement from everyone at the start solves that too.
Two practical courtesies: tell people you will delete a recording if they ask, and pause or stop if someone needs to say something off the record.
Fixing the player with a noisy mic or push-to-talk problems
Craig receives whatever Discord transmits, which means each player's own audio settings are baked into their track before the bot ever sees it. You cannot fix that in post.
Push-to-talk is the big one. It gates the mic, so the first syllable of a sentence is often cut off when a player starts talking before the key is fully down. On the recording this shows up as clipped words and dropped sentence starts, which the transcription then turns into nonsense. Ask players on push-to-talk to press the key a beat early and release it a beat late, or to switch to voice activity for game nights if their room is quiet enough.
Aggressive noise suppression causes a subtler version of the same problem. It can chew the ends off quiet speech and mute someone who is speaking softly for a character. If a player's track comes back with holes in it, have them turn suppression down a notch.
Headphones for everyone, always. Without them, the DM's voice comes out of a player's speakers and back into their mic, so your voice appears smeared and echoing on every track. Nothing else improves the recording this much for free.
If one player's track is genuinely unusable — constant fan hum, a mic clipping into distortion — you still have four clean tracks. Ask them to reposition the mic away from the fan, move it off the desk so it stops picking up typing, or use their phone's earbuds as a stopgap. Cheap headset mics placed correctly beat expensive mics placed badly.
Alternatives if Craig is not an option
Some groups do not play in Discord, or cannot add bots to the server they use. The options in rough order of usefulness:
Zoom. When you record locally, it can save a separate audio file for each participant — an option you enable in the recording settings before the call. Availability varies by plan and client version, so check it on a test call rather than on session night. If you already run the game in Zoom, turn that on before you go looking for anything else.
OBS Studio. Free, capable, records desktop audio and your mic, and can be configured to write them to separate audio tracks. It will not separate individual Discord users from each other, so you get one track containing all your players and one containing you. That is still better than a fully mixed file, since it keeps the DM's voice separate from the table's.
Audacity with a loopback device. Free and completely local. Audacity records from one input device at a time, so you route your system audio into a virtual input — VB-Audio Cable on Windows, BlackHole on macOS, the PulseAudio monitor source on Linux — and capture your mic through the same device. You get one mixed track, it takes a bit of audio-routing patience the first time, and it works forever after that.
A phone voice recorder on the desk. Genuinely fine as a backup, and the right answer if your table is hybrid with some players in the room. Start it, put it face down, forget it. It costs nothing and it is the thing that rescues the night your main recording turns out to be four hours of silence.
Turning the recording into a searchable transcript
Audio alone does not solve the problem you were trying to solve. Nobody scrubs through a four-hour file to remember the name of the innkeeper.
You can do the transcription yourself with Whisper, OpenAI's open-source speech model, and this is a real option worth taking seriously. It is free, it runs entirely on your own machine so nothing leaves your computer, and the quality is good. Community builds like whisper.cpp and faster-whisper are considerably quicker than the original release and run acceptably on a normal laptop.
The honest caveats: it needs setup. You are installing a Python environment or a command-line binary, picking a model size, and learning enough of the flags to point it at a file. On CPU, transcribing a long session with the larger models takes hours — you start it and come back later. On a single mixed file it has no idea who is speaking unless you add a separate speaker-diarization step, which is another install and another set of problems; per-speaker tracks sidestep that entirely, at the cost of running the job once per player. And it will confidently mangle every proper noun in your setting, because Thordak Ironforge is not a word. You can improve that by feeding it an initial prompt containing your campaign's names, which helps and which you have to remember to maintain.
Then there is the part nobody warns you about. When it finishes, you have a text file containing tens of thousands of words with no paragraphs, no structure, no headings, and no distinction between a critical plot revelation and ten minutes of arguing about pizza. It is searchable in the sense that Ctrl+F works, which is a real improvement over audio, and it is not the same thing as knowing what happened in your campaign.
Getting from that wall of text to something useful — a recap you can read to the table, a list of every NPC they have met, an answer to which session they first heard the name of the cult — is a second job. Whether you do that by hand in Obsidian or Notion, feed chunks of the transcript to an AI assistant, or use something built for it, budget for it as its own step rather than assuming the transcript is the finish line.
Related guides
- How to record your D&D session so the audio is actually usable
Where to put the mic at a real table, how to record Discord and Zoom games, cheap gear that helps, and what good enough session audio actually means.
- How to Take Notes During a D&D Session Without Wrecking the Game
A working note system for DMs: what to capture mid-session, what can wait, the one-line-per-beat method, player scribes, and the 10-minute post-game dump.
- How to Write a D&D Session Recap Your Players Will Actually Read
A short, practical method for D&D session recaps: previously-on framing, what to cut, how to hide reveals, a worked example, and a skeleton you can copy.