How to Record and Transcribe a D&D Session on Discord

The full workflow for an online game: what to record with, why per-speaker tracks make such a difference, how to handle the player whose mic is a disaster, and what your options actually are for turning four hours of audio into something you can search.

· 9 min read

What to use to record a Discord D&D session

If your game runs in a Discord voice channel, the usual answer is Craig — a recording bot you invite to your server that joins the voice channel and records each person to their own separate audio track. It is free to use, it has been a mainstay for tabletop groups for years, and the multi-track output is the single most useful thing you can have when it comes time to transcribe.

Set it up once and the per-session routine is about thirty seconds of work: run the join command before you start, run the stop command when you finish, then download the file from the link Craig sends you.

Everything below is the detail — consent, formats, mic problems, alternatives if Craig will not work for your group, and what to do with the audio afterwards. For the setup in full — every command, the format decision and the free-tier limits — see the dedicated Craig guide.

Why a screen recorder is the wrong tool for this

Discord's own client is not built for this — there is no reliable built-in way to record a voice call and end up with usable audio. What people reach for instead is screen sharing plus a screen-capture tool, or a general desktop recorder pointed at the Discord window. Both work in the sense that you end up with a file. Both are the wrong shape for the job.

A screen or desktop capture gives you one mixed stream: everyone's voice, your music bot, your notification pings, and whatever else your speakers were doing, all flattened into a single track you can never pull apart again. When two players talk over each other — which is most of a combat round — the transcription has to guess at overlapping speech, and it guesses badly.

You also end up carrying a video file around for audio you actually want, which makes it larger, slower to upload, and more annoying to archive.

The other problem is reliability. A screen recorder runs on your machine, which means it dies when your machine does. A bot sitting in the voice channel is recording server-side: if your client crashes, it keeps capturing everyone else and picks you back up when you rejoin, instead of losing the whole session.

Recording with Craig, in one paragraph

Invite the bot from craig.chat ahead of time, run its join command in your voice channel at the top of the session, confirm it announces itself, and run the stop command at the end — it sends you a download link that expires, so pull the file down the same night. Craig records each speaker to a separate track as they talk, including anyone who joins late.

That is the whole per-session flow. The setup detail — server permissions, the exact commands, the six-hour and seven-day limits with sources, the Giarc backup bot, and what each download option is actually for — lives on the dedicated Craig guide, which owns that ground so this page does not have to repeat it.

Why per-speaker tracks matter so much for transcription

This is the reason to bother with a bot at all. Craig records each participant's audio stream individually, so a five-person table produces five separate files that all start at the same moment.

With one mixed track, crosstalk is unrecoverable. When your rogue interrupts your cleric mid-sentence, the transcription model hears one muddy waveform and produces one mangled line, or drops the quieter speaker entirely. Every D&D table does this constantly — the reactions, the jokes, the three people all calling out what they want to do.

With separate tracks, there is nothing to untangle. Each file contains one clean voice with silence in between. Transcription accuracy goes up, and you get speaker attribution for free, because the file name tells you who was talking.

It also means one bad audio source is contained. A player with a buzzing mic damages exactly one track instead of contaminating the whole recording.

Which download to take

Two things, and you are done. Take a multi-track download as your archive copy — it holds the most information and cannot be recreated once the link expires — and make one single mixed file as your working copy for transcription. On Craig's free tier the mixed-down downloads sit behind the patron tier, so the free route to that working copy is the Audacity project: open it and export a single mixed file in a couple of clicks.

Two mechanical notes: the multi-track download arrives as a ZIP of separate per-speaker files, not one playable file, so unzip it before uploading anywhere (a ZIP is also not the same thing as a recorder that writes several audio tracks into one container, like OBS — DM Scribe mixes those automatically on upload). And per-format advice — FLAC vs AAC vs the reduced-quality WAV — is on the Craig guide.

Ask your players before you record anything

Do this. It is not a formality and it is not optional socially.

Say it out loud at the top of the session, in plain terms: you want to record the audio so you can write recaps and keep track of the campaign, it is stored privately, it is not going on the internet, and anyone who is not comfortable can say so now. Then actually accept a no.

Craig posts a message when it starts recording, and players can see the bot sitting in the voice channel list, so a recording is not a secret in practice. Announcing it anyway is what keeps people relaxed. A table that has forgotten there is a bot in the channel talks freely; a table that discovers a recording afterwards does not talk freely again.

There is also a legal dimension depending on where everyone lives, since your players may be in different jurisdictions with different rules about recording conversations. Getting explicit agreement from everyone at the start solves that too.

Two practical courtesies: tell people you will delete a recording if they ask, and pause or stop if someone needs to say something off the record.

Fixing the player with a noisy mic or push-to-talk problems

Craig receives whatever Discord transmits, which means each player's own audio settings are baked into their track before the bot ever sees it. You cannot fix that in post.

Push-to-talk is the big one. It gates the mic, so the first syllable of a sentence is often cut off when a player starts talking before the key is fully down. On the recording this shows up as clipped words and dropped sentence starts, which the transcription then turns into nonsense. Ask players on push-to-talk to press the key a beat early and release it a beat late, or to switch to voice activity for game nights if their room is quiet enough.

Aggressive noise suppression causes a subtler version of the same problem. It can chew the ends off quiet speech and mute someone who is speaking softly for a character. If a player's track comes back with holes in it, have them turn suppression down a notch.

Headphones for everyone, always. Without them, the DM's voice comes out of a player's speakers and back into their mic, so your voice appears smeared and echoing on every track. Nothing else improves the recording this much for free.

If one player's track is genuinely unusable — constant fan hum, a mic clipping into distortion — you still have four clean tracks. Ask them to reposition the mic away from the fan, move it off the desk so it stops picking up typing, or use their phone's earbuds as a stopgap. Cheap headset mics placed correctly beat expensive mics placed badly.

Alternatives if Craig is not an option

Some groups do not play in Discord, or cannot add bots to the server they use. The options in rough order of usefulness:

Zoom. When you record locally, it can save a separate audio file for each participant — an option you enable in the recording settings before the call. Availability varies by plan and client version, so check it on a test call rather than on session night. If you already run the game in Zoom, turn that on before you go looking for anything else.

OBS Studio. Free, capable, records desktop audio and your mic, and can be configured to write them to separate audio tracks. It will not separate individual Discord users from each other, so you get one track containing all your players and one containing you. That is still better than a fully mixed file, since it keeps the DM's voice separate from the table's.

Audacity with a loopback device. Free and completely local. Audacity records from one input device at a time, so you route your system audio into a virtual input — VB-Audio Cable on Windows, BlackHole on macOS, the PulseAudio monitor source on Linux — and capture your mic through the same device. You get one mixed track, it takes a bit of audio-routing patience the first time, and it works forever after that.

A phone voice recorder on the desk. Genuinely fine as a backup, and the right answer if your table is hybrid with some players in the room. Start it, put it face up in the middle, forget it. It costs nothing and it is the thing that rescues the night your main recording turns out to be four hours of silence.

Turning the recording into a searchable transcript

Audio alone does not solve the problem you were trying to solve. Nobody scrubs through a four-hour file to remember the name of the innkeeper.

You can do the transcription yourself with Whisper, OpenAI's open-source speech model, and this is a real option worth taking seriously. It is free, it runs entirely on your own machine so nothing leaves your computer, and the quality is good. Community builds like whisper.cpp and faster-whisper are considerably quicker than the original release and run acceptably on a normal laptop.

The honest caveats: it needs setup. You are installing a Python environment or a command-line binary, picking a model size, and learning enough of the flags to point it at a file. On CPU, transcribing a long session with the larger models takes hours — you start it and come back later. On a single mixed file it has no idea who is speaking unless you add a separate speaker-diarization step, which is another install and another set of problems; per-speaker tracks sidestep that entirely, at the cost of running the job once per player. And it will confidently mangle every proper noun in your setting, because Thordak Ironforge is not a word. You can improve that by feeding it an initial prompt containing your campaign's names, which helps and which you have to remember to maintain.

Then there is the part nobody warns you about. When it finishes, you have a text file containing tens of thousands of words with no paragraphs, no structure, no headings, and no distinction between a critical plot revelation and ten minutes of arguing about pizza. It is searchable in the sense that Ctrl+F works, which is a real improvement over audio, and it is not the same thing as knowing what happened in your campaign.

Getting from that wall of text to something useful — a recap you can read to the table, a list of every NPC they have met, an answer to which session they first heard the name of the cult — is a second job. Whether you do that by hand in Obsidian or Notion, feed chunks of the transcript to an AI assistant, or use something built for it, budget for it as its own step rather than assuming the transcript is the finish line.

Related guides

All guides