Hosting an AI-Powered Discord Bot: Memory, Keys and Costs
Practical hosting advice for AI Discord bots — protecting API keys, controlling costs, handling slow responses within Discord's limits, and sizing memory.
On this page
AI-powered bots — chat assistants, summarisers, image generators, moderation helpers — are some of the most popular bots being built today. They also break in ways ordinary bots don’t: responses take longer than Discord’s interaction deadline, API bills spike overnight, and conversation history quietly eats memory. This article covers what’s different about hosting them.
Most AI bots don’t run the model themselves
The first question is where inference happens.
Calling a hosted model API (from providers such as Anthropic, OpenAI or Google) is how most bots work. The bot sends a request, waits, and posts the answer. The bot itself needs modest resources — it’s mostly waiting on network calls — but it needs careful handling of keys, costs and latency.
Running a model locally means loading the model’s weights into memory and computing responses on your server. Even small language models need several gigabytes of RAM, and without a GPU, generation is slow. For most Discord bots, a hosted API is far more practical. Local models make sense for small, specialised models (embeddings, classifiers) or when you have dedicated hardware.
The rest of this article assumes the common case: a bot calling a hosted API.
Protect your API keys like your bot token
An AI provider key is effectively a credit card. If it leaks, someone else can run up charges against your account until you notice. Treat it exactly like your Discord token:
- Keep it in an environment variable, never in code — on Kerit Cloud, environment variables are stored encrypted and kept out of logs.
- Set a spending limit in your provider’s dashboard. It’s the single best protection against a leaked key or a runaway loop.
- Use a separate key for development and production, so you can revoke one without breaking the other.
- Never echo raw API errors to users; some include request details you’d rather not share.
Our guide to keeping your bot token safe applies to every secret, not just the token.
Handle slow responses within Discord’s limits
Discord requires an interaction to be acknowledged within three seconds. Model responses often take longer. Always defer first:
await interaction.deferReply();
const answer = await generate(prompt); // may take 5–30 seconds
await interaction.editReply(answer.slice(0, 2000));
After deferring, you have up to 15 minutes to edit or follow up. Two more practical points:
- Messages are capped at 2,000 characters (4,096 in an embed description). Split long answers into several follow-ups, or send them as a file.
- Streaming output can be shown by editing the reply as text arrives, but edit at most every second or two — editing on every token will hit rate limits. See handling Discord API rate limits.
Set a timeout on model calls, and reply with a friendly message if it expires, rather than leaving the “thinking…” indicator forever.
Control costs before they control you
AI usage is billed per token (roughly, per word) for both the prompt and the response. Costs grow in three ways people underestimate:
Conversation history. Sending the whole conversation with every message means the prompt grows with each turn. A long chat can cost many times more per message than the first one. Keep a sliding window of recent messages, or summarise older context.
Popularity. A bot that’s cheap in one server becomes expensive in five hundred. Budget per server, not per bot.
Abuse. People will paste huge texts, run the command in loops, or use your bot as a free proxy to the model.
Defences that work:
- Per-user and per-guild rate limits — for example, 20 requests per user per hour.
- Input length caps — reject or truncate prompts beyond a sensible size.
- Output length caps — set the maximum tokens per response in the API call.
- Smaller models for simple tasks — classification, short answers and moderation rarely need the largest model.
- Caching — identical questions (help text, FAQs) can reuse earlier answers.
- Premium tiers — many AI bots limit free usage and fund heavier use through subscriptions.
Track tokens used per guild in your database. When the bill arrives, you’ll know exactly where it came from.
Memory: where AI bots differ
The bot process doesn’t hold the model, but it can still use more memory than a typical bot:
- Conversation state — history per channel or user. Keep it in a database or Redis with an expiry, not an ever-growing in-memory map.
- Large responses and attachments — images, files and long texts held while processing.
- Concurrency — many requests in flight at once, each holding buffers while waiting on the API.
- Embeddings and vector search — bots that search documents may load indexes into memory.
On Kerit Cloud, the bot workload matrix places AI bots in the Supreme to Elite range — 8–12 GB of RAM and 400–600% CPU — for bots with heavy memory needs and request bursts, with premium support. Smaller AI bots that only call an API and keep history in a database often run comfortably on Pro or Extreme. Measure before you buy: see how much RAM a Discord bot needs.
Outbound networking and dedicated IPs
AI bots make many outbound HTTPS requests. Some enterprise APIs require you to register the IP addresses your requests come from. On shared infrastructure that address may change or be shared; the Supreme and Elite Discord bot plans include a dedicated IP, which makes allow-listing straightforward.
Safety and responsibility
A bot that generates text or images on demand needs guardrails:
- Follow your AI provider’s usage policies and Discord’s developer policies.
- Filter or moderate outputs in public channels, especially in servers with younger members.
- Respect age-restricted channels for any mature content — and generally avoid it.
- Be clear with users that responses are AI-generated and can be wrong.
- Think about what you store. Conversation logs can contain personal information; keep only what you need, and for as long as you need it.
A production checklist
- [ ] API keys in environment variables, with a provider spending limit
- [ ] Defer every interaction before calling the model
- [ ] Timeouts on model calls, with a friendly fallback message
- [ ] Responses split to fit Discord’s message limits
- [ ] Conversation history windowed or summarised, stored with an expiry
- [ ] Per-user and per-guild rate limits and input caps
- [ ] Token usage tracked per guild
- [ ] Output moderation where appropriate
Summary
Most AI Discord bots call a hosted model API, so the bot needs modest compute but careful engineering: keys protected and capped, every interaction deferred, responses split to fit Discord’s limits, and costs controlled with history windows, rate limits and usage tracking. Keep conversation state in a database with expiry, size memory for concurrency rather than the model, and add guardrails so the bot is safe for the communities it serves.