Running Multiple Lavalink Nodes and Balancing Load
Why and how to run several Lavalink nodes — capacity, failover and regions — with node selection, penalties, moving players between nodes and keeping configs in sync.
On this page
A single Lavalink node takes a music bot a long way. Eventually, though, one node becomes a risk: if it goes down, every song in every server stops, and a single machine can only serve so many players. Running several nodes solves both problems. This article explains how multi-node setups work and how to run one well.
Three reasons to add a node
- Capacity. Spread players across nodes so none is overloaded.
- Resilience. If one node fails or restarts for an update, others keep playing and new players go elsewhere.
- Geography. Nodes in different regions give listeners around the world a short audio path — see choosing a Lavalink region.
Even bots that fit comfortably on one node benefit from a second for resilience.
How clients use several nodes
Lavalink nodes don’t know about each other. The client library in your bot holds connections to all of them and decides which node gets each new player. Once a player is created on a node, it stays there unless the client moves it.
All the major clients support multiple nodes:
- Shoukaku (discord.js) — pass several nodes;
getIdealNode()picks the least loaded. - lavalink-client (discord.js) — a node manager with selection strategies.
- Wavelink (discord.py) — pass several nodes to
wavelink.Pool.connect(); the pool chooses a node for each player.
A Shoukaku example:
const nodes = [
{ name: 'us-east', url: 'us-east.example:2333', auth: process.env.LL_US_PASSWORD },
{ name: 'eu-central', url: 'eu.example:2333', auth: process.env.LL_EU_PASSWORD },
{ name: 'in-mumbai', url: 'mumbai.example:2333', auth: process.env.LL_IN_PASSWORD },
];
const shoukaku = new Shoukaku(new Connectors.DiscordJS(client), nodes, {
resume: true,
resumeTimeout: 60,
moveOnDisconnect: true, // move players to another node if theirs disconnects
});
And Wavelink:
nodes = [
wavelink.Node(identifier="us-east", uri=os.environ["LL_US_URI"], password=os.environ["LL_US_PASSWORD"]),
wavelink.Node(identifier="eu-central", uri=os.environ["LL_EU_URI"], password=os.environ["LL_EU_PASSWORD"]),
]
await wavelink.Pool.connect(nodes=nodes, client=bot)
How “least loaded” is decided
Every minute, each Lavalink node sends a stats message: players, playing players, CPU load, memory and frame statistics. Clients turn these into a penalty score and choose the node with the lowest penalty. Typical ingredients:
- the number of playing players,
- CPU load, weighted heavily as it approaches 100%,
- frame deficits and nulled frames, which indicate a struggling node.
The effect is that a node which is busy, CPU-bound or stuttering gets fewer new players automatically. You don’t need to implement this yourself, but it helps to know why a client picks the node it does.
Choosing by region
Load-based selection ignores geography. For global bots, pick the node by region first and load second. The approach:
- When your bot joins a voice channel, it receives a voice server update from Discord. The endpoint usually indicates the voice region.
- Map that region to your nearest node (for example, voice servers in India → your Mumbai node).
- If that node is down or overloaded, fall back to the least-loaded node overall.
Some libraries accept a custom node resolver or region hints; if yours doesn’t, you can select the node yourself before creating the player.
When a node goes down
Plan for a node disappearing — during an update, a crash or a network problem.
- Session resuming. If a node’s WebSocket drops briefly and comes back within the resume timeout, players continue as if nothing happened.
- Moving players. If a node is truly gone, players must be recreated on another node. Shoukaku’s
moveOnDisconnectdoes this automatically; with other clients, you can move a player yourself (Wavelink 3 offersplayer.switch_node(node)). The track resumes from roughly its last known position on the new node. - Telling users. A short “Music moved to a backup server” message is better than silence.
Test failover deliberately: stop one node during a quiet period and confirm that players move and new players go to the remaining nodes.
Keep nodes consistent
Players move between nodes, so nodes should behave the same way:
- Same Lavalink version and same plugin versions — a track that resolves on one node should resolve on the others.
- Same sources enabled and matching
application.ymlsettings for buffers and quality. - Unique, strong passwords per node stored in your bot’s environment variables.
- Consistent filters and plugin endpoints, so features like SponsorBlock work everywhere — see SponsorBlock for Lavalink.
Keep your configuration in a git repository and deploy the same file to every self-managed node. With managed nodes, ask for the same plugin set on each.
Rolling updates
Multiple nodes let you update without downtime:
- Update one node (new Lavalink or plugin version) while the others serve players.
- Its players resume or move; new players go to the healthy nodes.
- Confirm it’s healthy, then update the next.
This is especially valuable for the YouTube plugin, which needs frequent updates — see the YouTube source plugin.
Monitoring several nodes
Watch each node individually — an average can hide one node in trouble. Useful per-node numbers: playing players, CPU load, heap usage and frame deficits. Log the stats your client receives, or scrape each node’s Prometheus endpoint into one Grafana dashboard with a panel per node. See monitoring Lavalink with Prometheus and Grafana.
One big node or several small ones?
| One larger node | Several smaller nodes | |
|---|---|---|
| Simplicity | Easiest | More moving parts |
| Resilience | Single point of failure | One failure affects only its players |
| Regions | One location | Near listeners in each region |
| Updates | Brief interruption | Rolling, no downtime |
| Efficiency | Lowest overhead | Each node carries JVM baseline memory |
A common path: grow one node until it’s comfortably sized, then add a second in another region for resilience and geography.
On Kerit Cloud, managed plans are available in 14 regions and upgrades are live, so you can mix a larger primary node with smaller regional ones. Self-managed plans also let you run several Lavalink instances on one server with different ports and configs. See Lavalink plans compared.
Summary
Run multiple Lavalink nodes for capacity, resilience and geography. Your client library connects to all of them, scores them by load and picks one per player; add region-aware selection for global bots. Enable session resuming and automatic player moves so a failed node doesn’t stop the music, keep versions and plugins identical across nodes, update one node at a time, and monitor each node individually.