Discord Bots

Auto-Restart and Crash Recovery for Discord Bots

Why bots crash, how automatic restarts keep them online, and how to handle errors so a single bad event never takes your bot down for long.

On this page
  1. The four ways a bot goes offline
  2. How automatic restarts work
  3. Handling errors so crashes stay rare
  4. Diagnosing a crash loop
  5. Catching hangs
  6. Make restarts safe
  7. Summary

Every bot crashes eventually. A library update changes a return type, an API returns something unexpected, a user finds an input you never tested, or the server runs out of memory. What separates a reliable bot from an unreliable one isn’t whether it crashes — it’s how quickly it comes back and whether it learns anything from the crash.

This article covers the kinds of failures bots experience, how automatic restarts work, and the error handling that keeps restarts rare.

The four ways a bot goes offline

1. The process exits. An unhandled exception, an explicit process.exit(), or a fatal error in a native module ends the process. The bot disappears from the member list within a minute.

2. The process is killed. The operating system stops the process because it exceeded its memory limit. There’s often no stack trace — the process simply vanishes.

3. The connection drops but the process lives. Network blips and Discord-side restarts close the gateway connection. Good libraries reconnect and resume automatically, replaying missed events. You’ll see a reconnect in the logs, but the bot recovers on its own.

4. The process hangs. The bot is technically running but stops responding — a deadlock, an infinite loop, or blocking code that stalls the event loop. This is the hardest failure to notice because nothing crashes.

Automatic restarts solve the first two. Your library solves the third. The fourth needs monitoring.

How automatic restarts work

A supervisor — sometimes called a watchdog — starts your bot and waits. When the process exits for any reason, the supervisor starts it again. On your own server you’d configure this with systemd or PM2 (see keeping Node.js apps alive with PM2). On Kerit Cloud it’s built in: every bot runs under a watchdog that restarts it within seconds after an exception, an out-of-memory kill or a hardware fault, and the incident shows up in the live console so you can see what happened.

A restart takes a few seconds, and Discord libraries reconnect quickly, so most users never notice a single crash. The danger is a crash loop — the bot crashes on startup, restarts, crashes again, and never becomes usable.

Handling errors so crashes stay rare

Restarts are a safety net, not a strategy. The goal is for one bad command or event to fail on its own without taking the whole process down.

Node.js / discord.js

Wrap each command handler in try/catch and reply with a friendly error:

client.on(Events.InteractionCreate, async (interaction) => {
  if (!interaction.isChatInputCommand()) return;
  try {
    await client.commands.get(interaction.commandName)?.execute(interaction);
  } catch (err) {
    console.error(`[${interaction.commandName}] guild=${interaction.guildId}`, err);
    const msg = { content: 'That command failed — the error has been logged.', ephemeral: true };
    if (interaction.deferred || interaction.replied) await interaction.followUp(msg).catch(() => {});
    else await interaction.reply(msg).catch(() => {});
  }
});

Add process-level handlers as a last line of defence:

process.on('unhandledRejection', (reason) => {
  console.error('Unhandled rejection:', reason);
});

process.on('uncaughtException', (err) => {
  console.error('Uncaught exception — exiting:', err);
  process.exit(1); // let the supervisor restart a clean process
});

The asymmetry is deliberate. A stray rejected promise usually leaves the process healthy, so logging it is enough. After an uncaught exception, Node.js may be in an inconsistent state, and the official guidance is to exit and restart. Since you have a supervisor, exiting is cheap.

Also listen for client errors, which otherwise surface as unhandled error events:

client.on(Events.Error, (err) => console.error('Client error:', err));
client.on(Events.ShardDisconnect, (event, id) => console.warn(`Shard ${id} disconnected (${event.code})`));

Python / discord.py

discord.py catches exceptions inside event handlers and passes them to on_error, and app command failures to the tree’s error handler. Override them so errors are logged with context:

@bot.tree.error
async def on_app_command_error(interaction: discord.Interaction, error: app_commands.AppCommandError):
    log.exception("Command %s failed", interaction.command and interaction.command.name, exc_info=error)
    if interaction.response.is_done():
        await interaction.followup.send("That command failed.", ephemeral=True)
    else:
        await interaction.response.send_message("That command failed.", ephemeral=True)

Background tasks started with asyncio.create_task() are a common source of silent failures: if the task raises and nobody awaits it, the exception is only reported when the task is garbage-collected. Use discord.ext.tasks loops, which log errors, or attach a done callback that logs exceptions.

Diagnosing a crash loop

When a bot restarts over and over, the cause is almost always in the first few seconds of startup:

  • A missing environment variable. The token, database URL or an API key isn’t set on the server.
  • A missing dependency. Something isn’t listed in package.json or requirements.txt.
  • A bad deploy. The latest commit has a syntax error. Roll back with git revert and push.
  • Out of memory at startup. Large bots load a lot of data on connect. If the process dies before “ready” with no stack trace, check the memory graph and see how much RAM a Discord bot needs.
  • A corrupt data file. A JSON file half-written during a previous crash fails to parse on every start. Validate files on load and fall back to a backup — or move to a database.

Scroll back to the first error in the console, not the last. The last one is often a consequence of the first.

Catching hangs

A hung bot doesn’t exit, so the watchdog can’t help. Two simple defences:

Log a heartbeat. Every few minutes, log a line with gateway latency and memory. If the log goes quiet, the process is stuck.

Use external monitoring. Expose a tiny HTTP health endpoint that reports whether the client is ready, and point an uptime monitor at it — or monitor the bot’s presence. Setting up uptime monitoring walks through both approaches.

Most hangs in Node.js come from synchronous work blocking the event loop; in Python, from calling blocking libraries like requests or time.sleep() inside async code. Replace them with async equivalents (fetch or undici; aiohttp and asyncio.sleep).

Make restarts safe

Because restarts will happen, design for them:

  • Keep state in a database, not memory. Anything that must survive — ticket state, economy balances, reminders — belongs in storage. See building a ticket bot that survives restarts.
  • Make startup idempotent. Don’t re-send welcome messages, re-register commands or duplicate scheduled jobs every time the bot starts.
  • Shut down cleanly. Close database connections on SIGTERM so a restart never interrupts a write. See graceful shutdowns.
  • Reschedule on start. Reminders and timed punishments should be stored with their due time and rescheduled when the bot comes back.

Summary

Bots go offline by exiting, being killed, losing their connection or hanging. A watchdog restarts crashed processes in seconds — built in on Kerit Cloud — and Discord libraries handle reconnects. Your job is to catch errors per command so crashes stay rare, exit cleanly after truly fatal errors, watch for hangs, and keep state somewhere that survives a restart.