Lavalink

Monitoring Lavalink With Prometheus and Grafana

Monitor a Lavalink node properly — enabling Prometheus metrics, scraping them securely, building a Grafana dashboard and alerting on frame deficits, CPU and memory.

On this page
  1. Two ways to get numbers
  2. Securing the metrics endpoint
  3. Scraping with Prometheus
  4. What to put on the dashboard
  5. Alerts worth having
  6. Correlate before you act
  7. Monitor the bot too
  8. Managed vs self-managed monitoring
  9. Summary

A Lavalink node can be “up” and still sound terrible. Listeners hear stutters long before a process crashes, so monitoring a node means watching how well it delivers audio, not just whether it’s running. This guide sets up Prometheus and Grafana for Lavalink and covers what to graph and alert on.

Two ways to get numbers

The stats endpoint

Lavalink always exposes live statistics:

curl -s -H "Authorization: your-password" http://your-node:2333/v4/stats

It returns players, playing players, uptime, memory, CPU and frame statistics. Your client library also receives these every minute over the WebSocket. For small setups, logging them from your bot is enough — see how much RAM Lavalink needs for what each field means.

The Prometheus endpoint

For history, graphs and alerts, Lavalink can expose metrics in Prometheus format. Enable it in application.yml:

metrics:
  prometheus:
    enabled: true
    endpoint: /metrics

Restart the node and check it:

curl -s http://your-node:2333/metrics | head -40

You’ll see Lavalink’s own metrics (players, playing players and related counters) alongside standard JVM metrics such as heap usage, garbage collection time and thread counts. Metric names can differ between versions, so browse your endpoint’s output and use the names you see there.

Securing the metrics endpoint

The metrics endpoint is served on the same port as the rest of Lavalink. Check whether your version requires the node password on it: if curl returns metrics without an Authorization header, anyone who can reach the port can read them. Metrics don’t allow controlling the node, but they reveal usage details you may not want public.

The simplest protection is the firewall — allow your Prometheus server’s IP to reach port 2333 and block everyone else:

sudo ufw allow from PROMETHEUS_IP to any port 2333 proto tcp

If your version does require the password, Prometheus’ standard authorization option sends a Bearer token, which doesn’t match Lavalink’s raw-password header. Recent Prometheus versions can send custom headers per scrape job; alternatively, run Prometheus on the same private network or behind a small reverse proxy that adds the header.

Scraping with Prometheus

A minimal prometheus.yml job:

scrape_configs:
  - job_name: lavalink
    scrape_interval: 30s
    metrics_path: /metrics
    static_configs:
      - targets:
          - "us-node.example:2333"
          - "eu-node.example:2333"
        labels:
          service: lavalink

Give each node a recognisable target or label so you can tell them apart on dashboards. With several nodes, always graph them individually — an average hides the one node that’s struggling. See running multiple Lavalink nodes.

You can run Prometheus and Grafana on a small VPS, in Docker alongside other tools, or use a hosted Grafana service that accepts Prometheus remote-write.

What to put on the dashboard

A useful Lavalink dashboard needs only a handful of panels:

Panel Why it matters
Playing players The main driver of load. Compare with capacity.
Total players Players that exist, including paused and idle ones.
Node CPU / Lavalink CPU High and rising means filters, resampling or too many players.
JVM heap used vs max Close to max means the heap is too small.
GC pause time Spikes line up with audio stutters.
Frame deficit Frames that should have been sent but weren’t — the best proxy for “does it sound bad?”
Uptime Unexpected resets reveal crashes or restarts.

If your Lavalink version doesn’t export frame statistics to Prometheus, log frameStats from /v4/stats in your bot instead — it’s worth having either way.

Example queries (adjust metric names to match your endpoint):

# Heap usage as a percentage
100 * jvm_memory_bytes_used{area="heap"} / jvm_memory_bytes_max{area="heap"}

# Time spent in garbage collection per second
rate(jvm_gc_collection_seconds_sum[5m])

Alerts worth having

Alert on things that affect listeners, and keep the list short so alerts stay meaningful:

  • Node down — the scrape target is unreachable for a couple of minutes.
  • Frame deficits sustained — deficits above a small threshold for 5+ minutes mean audible stutter.
  • Heap above ~90% for 10 minutes — out-of-memory risk.
  • CPU saturated for 10 minutes — the node can’t keep up at peak.
  • Unexpected restart — uptime reset outside a maintenance window.

Send alerts somewhere you’ll see them: a private Discord channel through a webhook works well for small teams.

Correlate before you act

Graphs are most useful side by side. When listeners report stutters:

  • Deficits + GC spikes → heap or collector tuning. See JVM tuning for Lavalink.
  • Deficits + CPU pinned → too many players or filters for the CPU; lower resamplingQuality or opusEncodingQuality, or upgrade.
  • Deficits with healthy CPU and GC → network or distance; see choosing a Lavalink region.
  • Playing players climbing toward your usual peak → plan an upgrade or another node before the next busy evening.

Monitor the bot too

The node is half the system. Track the bot’s view as well: node connection events, track exceptions per minute (a spike usually means a source broke), and the time from /play to playback starting. Logging and monitoring a Discord bot covers the bot side.

Managed vs self-managed monitoring

On Kerit Cloud’s self-managed Lavalink plans, you have root access: enable the Prometheus endpoint and scrape it with your own Grafana setup for full visibility into CPU, memory, players and track load. On managed plans, the node is monitored for you around the clock, and you can still watch the stats your client receives.

Summary

Monitor Lavalink by what listeners hear, not just whether the process is up. Enable the Prometheus endpoint, restrict who can reach it, scrape each node separately, and build a small dashboard of playing players, CPU, heap, GC pauses and frame deficits. Alert on node-down, sustained deficits, high heap and saturated CPU, and read the graphs together to find whether stutters come from GC, CPU or the network.