Logging and Monitoring a Discord Bot in Production
Set up logs you can actually use, metrics that reveal problems early, and alerts that tell you your Discord bot is in trouble before your users do.
On this page
When a bot misbehaves in production, you usually find out from a user message like “the bot isn’t working”. Without good logs, you then spend an hour guessing. With them, you spend five minutes reading. This article covers how to log, what to measure and how to get alerted, without building an enterprise observability stack.
What good logs look like
A useful log line answers four questions: when it happened, how serious it is, where in the bot it happened and what context surrounded it.
Compare:
Error: Missing Permissions
with:
2026-05-02T14:31:07Z ERROR [cmd:ban] guild=81234 user=55621 target=99812 DiscordAPIError[50013]: Missing Permissions
The second line tells you which command failed, in which server, for which users, and the exact Discord error code. You can reproduce it or explain it to the server’s admins immediately.
Use log levels
Levels let you filter noise:
- error — something failed and a user was affected, or data may be wrong.
- warn — something unexpected that the bot recovered from: a rate limit, a reconnect, a retry.
- info — normal milestones: startup, ready, joined or left a guild, commands registered.
- debug — detail for troubleshooting, off in production by default.
Control the level with an environment variable (LOG_LEVEL=info) so you can turn on debug output temporarily without a code change.
Structured logging
For anything beyond a small bot, log structured data. In Node.js, pino is fast and outputs JSON:
const pino = require('pino');
const log = pino({ level: process.env.LOG_LEVEL ?? 'info' });
log.info({ guild: interaction.guildId, cmd: interaction.commandName }, 'command used');
log.error({ err, guild: interaction.guildId, cmd: interaction.commandName }, 'command failed');
In Python, the standard logging module is enough; include context in the message or use extra fields with a JSON formatter:
import logging
logging.basicConfig(
level=os.environ.get("LOG_LEVEL", "INFO"),
format="%(asctime)s %(levelname)s [%(name)s] %(message)s",
)
log = logging.getLogger("bot.commands")
log.error("ban failed guild=%s target=%s", interaction.guild_id, target.id, exc_info=True)
exc_info=True includes the full traceback — without it, you get the message but not where it came from.
Log to standard output
Write logs to stdout and stderr rather than to files your bot manages. On Kerit Cloud, output appears in the panel’s live console, and you never have to worry about a log file filling your disk. If you do write files, rotate them — an unrotated log is one of the most common causes of a full disk.
What not to log
Never log tokens, API keys, database URLs or full config objects. Be careful with message content and personal data: log IDs rather than usernames and message text where you can. If your bot has users in regions with privacy laws, less logged personal data means less to worry about.
Metrics that reveal problems early
Logs tell you what happened. Metrics tell you how the bot is doing over time. A handful of numbers cover most of what matters:
| Metric | Why it matters | Warning sign |
|---|---|---|
| Gateway latency | Network health, event-loop stalls | Sustained spikes or climbing trend |
| Memory (RSS) | Leaks, cache growth | Rises steadily and never levels off |
| CPU | Blocking code, heavy commands | Pinned near your plan’s limit |
| Guild count | Growth, sharding planning | Approaching 2,500 unsharded |
| Commands per minute | Usage, abuse | Sudden drops (bot broken) or spikes (abuse) |
| Command errors per minute | Bugs, API issues | Any sustained increase |
| Command duration (p95) | Slow database or APIs | Approaching 3 seconds |
The panel already graphs CPU, memory and network for your server. For the bot-specific numbers, the simplest approach is a periodic summary log line:
setInterval(() => {
log.info({
ping: client.ws.ping,
guilds: client.guilds.cache.size,
rssMB: Math.round(process.memoryUsage().rss / 1e6),
cmds: counters.commands,
errors: counters.errors,
}, 'health');
counters.commands = counters.errors = 0;
}, 5 * 60_000);
Every five minutes you get a snapshot. When something goes wrong, you can scroll back and see exactly when latency jumped or memory started climbing.
For larger bots, expose these as Prometheus metrics and graph them in Grafana. The pattern is the same one we describe for Lavalink in monitoring Lavalink with Prometheus and Grafana.
Alerts: find out before your users do
Metrics are useless if nobody looks at them. Set up a few alerts that reach you wherever you are.
Uptime monitoring
An external monitor checks your bot from outside and alerts you when it stops responding. Give your bot a tiny HTTP endpoint that returns 200 only when the client is ready:
const http = require('node:http');
http.createServer((req, res) => {
const ok = client.isReady() && client.ws.ping < 1000;
res.writeHead(ok ? 200 : 503).end(ok ? 'ok' : 'unhealthy');
}).listen(process.env.SERVER_PORT || 8080);
Point a free uptime service at it. Details, including a no-HTTP alternative, are in setting up uptime monitoring.
Error alerts to Discord
Send errors to a private channel in your own server using a webhook. It’s free and you’ll see problems on your phone:
async function alert(text) {
await fetch(process.env.ALERT_WEBHOOK_URL, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ content: text.slice(0, 1900) }),
}).catch(() => {});
}
Rate-limit your alerts — group repeated errors and send at most one message per minute — or a burst of failures will spam the channel and hit Discord’s webhook limits. Keep the webhook URL in an environment variable; anyone who has it can post to that channel.
Startup notifications
Send a short message when the bot starts. Frequent unexpected startup messages mean the bot is crash-looping, which you’ll notice immediately.
Using logs to debug
When something breaks:
- Find the first error. Scroll to the earliest error around the incident. Later errors are often consequences.
- Match it to context. Guild, user and command fields tell you whether it’s one server’s configuration or a bug affecting everyone.
- Check the health lines. Did memory, latency or error rates change just before?
- Check what changed. A deploy just before the problem is the most likely cause — roll back with
git revertand push, then investigate calmly.
Reading logs and debugging crashes goes deeper into interpreting stack traces from Node.js, Python and Java.
Summary
Log with timestamps, levels and context — command, guild and user IDs — to standard output, and never log secrets. Track a small set of metrics: latency, memory, CPU, guilds, command volume, errors and duration. Then add alerts: an uptime monitor on a health endpoint, error notifications to a private Discord channel, and startup messages that reveal crash loops. Together they turn “the bot isn’t working” into a five-minute fix.