Monitoring Lavalink With Prometheus and Grafana
Monitor a Lavalink node properly — enabling Prometheus metrics, scraping them securely, building a Grafana dashboard and alerting on frame deficits, CPU and memory.
On this page
A Lavalink node can be “up” and still sound terrible. Listeners hear stutters long before a process crashes, so monitoring a node means watching how well it delivers audio, not just whether it’s running. This guide sets up Prometheus and Grafana for Lavalink and covers what to graph and alert on.
Two ways to get numbers
The stats endpoint
Lavalink always exposes live statistics:
curl -s -H "Authorization: your-password" http://your-node:2333/v4/stats
It returns players, playing players, uptime, memory, CPU and frame statistics. Your client library also receives these every minute over the WebSocket. For small setups, logging them from your bot is enough — see how much RAM Lavalink needs for what each field means.
The Prometheus endpoint
For history, graphs and alerts, Lavalink can expose metrics in Prometheus format. Enable it in application.yml:
metrics:
prometheus:
enabled: true
endpoint: /metrics
Restart the node and check it:
curl -s http://your-node:2333/metrics | head -40
You’ll see Lavalink’s own metrics (players, playing players and related counters) alongside standard JVM metrics such as heap usage, garbage collection time and thread counts. Metric names can differ between versions, so browse your endpoint’s output and use the names you see there.
Securing the metrics endpoint
The metrics endpoint is served on the same port as the rest of Lavalink. Check whether your version requires the node password on it: if curl returns metrics without an Authorization header, anyone who can reach the port can read them. Metrics don’t allow controlling the node, but they reveal usage details you may not want public.
The simplest protection is the firewall — allow your Prometheus server’s IP to reach port 2333 and block everyone else:
sudo ufw allow from PROMETHEUS_IP to any port 2333 proto tcp
If your version does require the password, Prometheus’ standard authorization option sends a Bearer token, which doesn’t match Lavalink’s raw-password header. Recent Prometheus versions can send custom headers per scrape job; alternatively, run Prometheus on the same private network or behind a small reverse proxy that adds the header.
Scraping with Prometheus
A minimal prometheus.yml job:
scrape_configs:
- job_name: lavalink
scrape_interval: 30s
metrics_path: /metrics
static_configs:
- targets:
- "us-node.example:2333"
- "eu-node.example:2333"
labels:
service: lavalink
Give each node a recognisable target or label so you can tell them apart on dashboards. With several nodes, always graph them individually — an average hides the one node that’s struggling. See running multiple Lavalink nodes.
You can run Prometheus and Grafana on a small VPS, in Docker alongside other tools, or use a hosted Grafana service that accepts Prometheus remote-write.
What to put on the dashboard
A useful Lavalink dashboard needs only a handful of panels:
| Panel | Why it matters |
|---|---|
| Playing players | The main driver of load. Compare with capacity. |
| Total players | Players that exist, including paused and idle ones. |
| Node CPU / Lavalink CPU | High and rising means filters, resampling or too many players. |
| JVM heap used vs max | Close to max means the heap is too small. |
| GC pause time | Spikes line up with audio stutters. |
| Frame deficit | Frames that should have been sent but weren’t — the best proxy for “does it sound bad?” |
| Uptime | Unexpected resets reveal crashes or restarts. |
If your Lavalink version doesn’t export frame statistics to Prometheus, log frameStats from /v4/stats in your bot instead — it’s worth having either way.
Example queries (adjust metric names to match your endpoint):
# Heap usage as a percentage
100 * jvm_memory_bytes_used{area="heap"} / jvm_memory_bytes_max{area="heap"}
# Time spent in garbage collection per second
rate(jvm_gc_collection_seconds_sum[5m])
Alerts worth having
Alert on things that affect listeners, and keep the list short so alerts stay meaningful:
- Node down — the scrape target is unreachable for a couple of minutes.
- Frame deficits sustained — deficits above a small threshold for 5+ minutes mean audible stutter.
- Heap above ~90% for 10 minutes — out-of-memory risk.
- CPU saturated for 10 minutes — the node can’t keep up at peak.
- Unexpected restart — uptime reset outside a maintenance window.
Send alerts somewhere you’ll see them: a private Discord channel through a webhook works well for small teams.
Correlate before you act
Graphs are most useful side by side. When listeners report stutters:
- Deficits + GC spikes → heap or collector tuning. See JVM tuning for Lavalink.
- Deficits + CPU pinned → too many players or filters for the CPU; lower
resamplingQualityoropusEncodingQuality, or upgrade. - Deficits with healthy CPU and GC → network or distance; see choosing a Lavalink region.
- Playing players climbing toward your usual peak → plan an upgrade or another node before the next busy evening.
Monitor the bot too
The node is half the system. Track the bot’s view as well: node connection events, track exceptions per minute (a spike usually means a source broke), and the time from /play to playback starting. Logging and monitoring a Discord bot covers the bot side.
Managed vs self-managed monitoring
On Kerit Cloud’s self-managed Lavalink plans, you have root access: enable the Prometheus endpoint and scrape it with your own Grafana setup for full visibility into CPU, memory, players and track load. On managed plans, the node is monitored for you around the clock, and you can still watch the stats your client receives.
Summary
Monitor Lavalink by what listeners hear, not just whether the process is up. Enable the Prometheus endpoint, restrict who can reach it, scrape each node separately, and build a small dashboard of playing players, CPU, heap, GC pauses and frame deficits. Alert on node-down, sustained deficits, high heap and saturated CPU, and read the graphs together to find whether stutters come from GC, CPU or the network.