Prometheus and Grafana bare-metal monitoring architecture
Prometheus collects numeric time-series metrics by scraping HTTP endpoints, usually every 15 to 60 seconds. Grafana queries Prometheus and turns those metrics into dashboards, charts, tables, and alerts. Node Exporter exposes Linux host metrics such as CPU usage, memory pressure, disk space, filesystem I/O, load average, and network throughput.
For production monitoring, run Prometheus and Grafana on a dedicated bare-metal server if monitoring is business-critical, retention exceeds 30 days, or you expect more than 100 monitored systems. A VPS is suitable for a small lab, a few websites, or fewer than 20 low-volume targets. Bare metal provides predictable disk I/O and avoids noisy-neighbor problems during Prometheus compaction.
This setup uses one monitoring server and multiple monitored Linux servers:
- Prometheus: stores metrics and evaluates alert rules on port 9090.
- Node Exporter: exposes host metrics on port 9100.
- Grafana: provides dashboards on port 3000 behind Nginx.
- Alertmanager: groups and routes alerts on port 9093.
- Nginx: publishes Grafana securely over HTTPS.
Prerequisites and server sizing
Use a clean Ubuntu Server 24.04 LTS installation with a static public IP address. The commands below require a non-root user with sudo access. Open TCP ports 22, 80, and 443 to the internet only if needed. Keep Prometheus port 9090, Alertmanager port 9093, and Node Exporter port 9100 private whenever possible.
- Operating system: Ubuntu Server 24.04 LTS x86_64
- Minimum monitoring server: 4 CPU cores, 8 GB RAM, 160 GB NVMe SSD
- Recommended production server: 8 CPU cores, 32 GB RAM, 960 GB NVMe SSD
- Scrape interval used here: 15 seconds
- Prometheus retention used here: 30 days
- Required DNS record: an A or AAAA record such as
grafana.example.com
Disk capacity is the limiting resource for most Prometheus deployments. A 15-second scrape interval produces 5,760 samples per metric each day. Lower-cardinality metrics such as CPU, memory, and disk statistics are inexpensive; labels containing user IDs, request IDs, IP addresses, container hashes, or URLs with dynamic paths can consume storage and memory rapidly.
Production sizing by monitored host count
For 30-day retention, a 15-second scrape interval, and standard Node Exporter metrics, size the monitoring server as follows.
Takeaway: Start with 8 CPU cores, 32 GB RAM, and 960 GB NVMe for around 100 Linux servers; increase disk capacity before reducing your retention period.
| Monitored Linux hosts | vCPU / CPU cores | RAM | Disk | Monthly bandwidth |
|---|---|---|---|---|
| 1-20 hosts | 4 cores | 8 GB | 160 GB NVMe | 1 TB |
| 21-150 hosts | 8 cores | 32 GB | 960 GB NVMe | 3 TB |
| 151-500 hosts | 16 cores | 64 GB | 2 × 1.92 TB NVMe in RAID 1 | 8 TB |
These figures cover operating-system metrics and a moderate number of application metrics. Add capacity for MySQL, PostgreSQL, Redis, mail-server, game-server, Kubernetes, CI/CD, streaming, or blackbox probe metrics. For more than 500 hosts, multi-year retention, or high-cardinality application telemetry, use Prometheus federation or a long-term metrics platform such as Thanos, Mimir, or VictoriaMetrics.
Step 1: Prepare the bare-metal server
Update packages, set the timezone, install required utilities, and enable the uncomplicated firewall. Replace UTC with your preferred timezone if required.
sudo apt update && sudo apt -y upgrade
sudo timedatectl set-timezone UTC
sudo apt install -y curl wget tar gnupg2 ca-certificates apt-transport-https nginx ufw
sudo ufw allow OpenSSH
sudo ufw allow 'Nginx Full'
sudo ufw enable
sudo ufw status
Create a dedicated directory for Prometheus data. NVMe storage is preferred because Prometheus continuously writes a write-ahead log and periodically compacts time-series blocks.
sudo mkdir -p /var/lib/prometheus
sudo mkdir -p /etc/prometheus/rules
sudo useradd --no-create-home --shell /usr/sbin/nologin prometheus
sudo chown -R prometheus:prometheus /var/lib/prometheus /etc/prometheus
Step 2: Install Prometheus
Download Prometheus 2.53.4 for Linux AMD64, install its binaries, and create a systemd service. Confirm the release version at the Prometheus GitHub releases page before upgrading later; keep version changes deliberate and tested.
cd /tmp
PROM_VERSION=2.53.4
wget https://github.com/prometheus/prometheus/releases/download/v${PROM_VERSION}/prometheus-${PROM_VERSION}.linux-amd64.tar.gz
tar -xzf prometheus-${PROM_VERSION}.linux-amd64.tar.gz
sudo install -m 0755 prometheus-${PROM_VERSION}.linux-amd64/prometheus /usr/local/bin/prometheus
sudo install -m 0755 prometheus-${PROM_VERSION}.linux-amd64/promtool /usr/local/bin/promtool
sudo cp -r prometheus-${PROM_VERSION}.linux-amd64/consoles /etc/prometheus/
sudo cp -r prometheus-${PROM_VERSION}.linux-amd64/console_libraries /etc/prometheus/
sudo chown -R prometheus:prometheus /etc/prometheus
Create /etc/prometheus/prometheus.yml. Replace 10.20.0.11 and 10.20.0.12 with private IP addresses of monitored hosts. Private network addresses avoid exposing Node Exporter to the public internet.
sudo tee /etc/prometheus/prometheus.yml > /dev/null <<'EOF'
global:
scrape_interval: 15s
evaluation_interval: 15s
external_labels:
environment: production
site: bare-metal-1
alerting:
alertmanagers:
- static_configs:
- targets: ['127.0.0.1:9093']
rule_files:
- /etc/prometheus/rules/*.yml
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ['127.0.0.1:9090']
- job_name: node
static_configs:
- targets:
- '127.0.0.1:9100'
- '10.20.0.11:9100'
- '10.20.0.12:9100'
EOF
sudo chown prometheus:prometheus /etc/prometheus/prometheus.yml
Create and start the Prometheus systemd unit with 30-day retention. The --storage.tsdb.retention.time=30d setting limits retained blocks by age; also use a size limit to prevent monitoring data from filling the server volume.
sudo tee /etc/systemd/system/prometheus.service > /dev/null <<'EOF'
[Unit]
Description=Prometheus Monitoring
Wants=network-online.target
After=network-online.target
[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus \
--storage.tsdb.retention.time=30d \
--storage.tsdb.retention.size=120GB \
--web.listen-address=127.0.0.1:9090
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now prometheus
sudo systemctl status prometheus --no-pager
Step 3: Install Node Exporter on every monitored Linux server
Run this section on the Prometheus server and on each target machine. Node Exporter listens on port 9100. Bind it to a private IP or use firewall rules so only the Prometheus server can connect.
cd /tmp
NODE_VERSION=1.8.2
sudo useradd --no-create-home --shell /usr/sbin/nologin node_exporter
wget https://github.com/prometheus/node_exporter/releases/download/v${NODE_VERSION}/node_exporter-${NODE_VERSION}.linux-amd64.tar.gz
tar -xzf node_exporter-${NODE_VERSION}.linux-amd64.tar.gz
sudo install -m 0755 node_exporter-${NODE_VERSION}.linux-amd64/node_exporter /usr/local/bin/node_exporter
sudo tee /etc/systemd/system/node_exporter.service > /dev/null <<'EOF'
[Unit]
Description=Prometheus Node Exporter
Wants=network-online.target
After=network-online.target
[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter --web.listen-address=:9100
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now node_exporter
curl -s http://127.0.0.1:9100/metrics | head
On each monitored host, allow port 9100 only from the Prometheus server private address. Replace 10.20.0.10 with the monitoring server address.
sudo ufw allow from 10.20.0.10 to any port 9100 proto tcp
sudo ufw status numbered
Step 4: Add alert rules and Alertmanager
Create alerts for an unreachable host, sustained high CPU, low disk space, and low available memory. Alerts should identify an actionable condition rather than every transient spike.
sudo tee /etc/prometheus/rules/host-alerts.yml > /dev/null <<'EOF'
groups:
- name: host-alerts
rules:
- alert: HostDown
expr: up{job="node"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Host {{ $labels.instance }} is unreachable"
- alert: HighCPUUsage
expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 90
for: 10m
labels:
severity: warning
annotations:
summary: "CPU usage exceeds 90% on {{ $labels.instance }}"
- alert: LowDiskSpace
expr: (node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes{fstype!~"tmpfs|overlay"}) * 100 < 10
for: 10m
labels:
severity: critical
annotations:
summary: "Less than 10% disk space remains on {{ $labels.instance }}"
- alert: LowAvailableMemory
expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100 < 10
for: 10m
labels:
severity: warning
annotations:
summary: "Available memory is below 10% on {{ $labels.instance }}"
EOF
sudo chown prometheus:prometheus /etc/prometheus/rules/host-alerts.yml
sudo promtool check config /etc/prometheus/prometheus.yml
sudo systemctl restart prometheus
Install Alertmanager and begin with a local receiver. Replace the receiver configuration with a properly secured SMTP relay, webhook, Slack-compatible endpoint, or PagerDuty integration after validating alert delivery.
cd /tmp
ALERT_VERSION=0.27.0
wget https://github.com/prometheus/alertmanager/releases/download/v${ALERT_VERSION}/alertmanager-${ALERT_VERSION}.linux-amd64.tar.gz
tar -xzf alertmanager-${ALERT_VERSION}.linux-amd64.tar.gz
sudo useradd --no-create-home --shell /usr/sbin/nologin alertmanager
sudo install -m 0755 alertmanager-${ALERT_VERSION}.linux-amd64/alertmanager /usr/local/bin/alertmanager
sudo mkdir -p /etc/alertmanager /var/lib/alertmanager
sudo chown -R alertmanager:alertmanager /etc/alertmanager /var/lib/alertmanager
sudo tee /etc/alertmanager/alertmanager.yml > /dev/null <<'EOF'
route:
receiver: local-log
group_by: ['alertname', 'instance']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receivers:
- name: local-log
EOF
sudo tee /etc/systemd/system/alertmanager.service > /dev/null <<'EOF'
[Unit]
Description=Prometheus Alertmanager
After=network-online.target
[Service]
User=alertmanager
Group=alertmanager
ExecStart=/usr/local/bin/alertmanager --config.file=/etc/alertmanager/alertmanager.yml --storage.path=/var/lib/alertmanager --web.listen-address=127.0.0.1:9093
Restart=always
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now alertmanager
Step 5: Install Grafana
Install Grafana from its signed APT repository. Grafana listens only on localhost because Nginx will handle public HTTPS access.
sudo mkdir -p /etc/apt/keyrings
wget -q -O - https://apt.grafana.com/gpg.key | sudo gpg --dearmor -o /etc/apt/keyrings/grafana.gpg
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt update
sudo apt install -y grafana
sudo sed -i 's/^;http_addr =.*/http_addr = 127.0.0.1/' /etc/grafana/grafana.ini
sudo systemctl enable --now grafana-server
sudo systemctl status grafana-server --no-pager
Log in initially through the local service using an SSH tunnel rather than exposing port 3000:
ssh -L 3000:127.0.0.1:3000 youruser@YOUR_SERVER_IP
Open http://localhost:3000, log in with admin and admin, and change the password immediately. Add a Prometheus data source with URL http://127.0.0.1:9090. Import dashboard ID 1860 for the widely used Node Exporter Full dashboard, then select the Prometheus data source.
Step 6: Publish Grafana through Nginx and HTTPS
Set DNS for grafana.example.com to this server before requesting a certificate. Replace the domain in the configuration and Certbot command.
sudo tee /etc/nginx/sites-available/grafana > /dev/null <<'EOF'
server {
listen 80;
server_name grafana.example.com;
location / {
proxy_pass http://127.0.0.1:3000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
EOF
sudo ln -s /etc/nginx/sites-available/grafana /etc/nginx/sites-enabled/grafana
sudo rm -f /etc/nginx/sites-enabled/default
sudo nginx -t
sudo systemctl reload nginx
sudo apt install -y certbot python3-certbot-nginx
sudo certbot --nginx -d grafana.example.com --redirect --agree-tos -m [email protected]
sudo systemctl status certbot.timer --no-pager
Create individual Grafana users or connect an identity provider; do not share the default administrator account. Disable anonymous access unless dashboards are intentionally public.
Step 7: Verify collection, storage, and alerts
Check that Prometheus sees each target as healthy. A target status of up equals 1; 0 means Prometheus cannot scrape the endpoint.
curl -s http://127.0.0.1:9090/api/v1/query?query=up
curl -s http://127.0.0.1:9090/api/v1/targets | grep -o 'health":"[^"]*' | sort | uniq -c
sudo promtool check rules /etc/prometheus/rules/host-alerts.yml
sudo systemctl --no-pager --full status prometheus node_exporter alertmanager grafana-server nginx
In Grafana Explore, run these PromQL queries:
up{job="node"}
100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
node_memory_MemAvailable_bytes / 1024 / 1024
node_filesystem_avail_bytes{mountpoint="/"} / 1024 / 1024 / 1024
Test the HostDown alert safely by stopping Node Exporter on one non-critical host for more than two minutes, then restore it.
sudo systemctl stop node_exporter
# Wait at least 2 minutes, inspect Alerts in Prometheus, then restore:
sudo systemctl start node_exporter
Troubleshooting common Prometheus and Grafana problems
Target shows DOWN in Prometheus
Check the target service, firewall, and route from the Prometheus server. The following command must return metrics from the monitoring server. If it times out, allow TCP 9100 only from the Prometheus private IP and verify that the target address in prometheus.yml is correct.
curl -v http://10.20.0.11:9100/metrics
sudo systemctl status node_exporter --no-pager
sudo ss -lntp | grep 9100
Prometheus consumes too much disk or RAM
Check active series, disk consumption, and retention settings. Avoid labels with unbounded values. Reducing retention from 30 days to 15 days is safer than allowing the filesystem to become full, but add NVMe capacity for a sustained workload.
curl -s http://127.0.0.1:9090/api/v1/status/tsdb
sudo du -sh /var/lib/prometheus
sudo df -h /var/lib/prometheus
sudo journalctl -u prometheus -n 100 --no-pager
Grafana returns 502 Bad Gateway
A 502 response usually means Grafana is stopped, listening on a different address, or blocked by a configuration error. Check both services and validate Nginx before reloading it.
sudo systemctl status grafana-server nginx --no-pager
sudo ss -lntp | grep 3000
sudo nginx -t
sudo journalctl -u grafana-server -n 100 --no-pager
Grafana has no data
Confirm the Grafana data source URL is http://127.0.0.1:9090, not a public URL blocked by the firewall. In Prometheus, run up first. If no series are returned, fix Prometheus scraping before changing dashboard variables.
Operational practices for a reliable monitoring server
- Back up Grafana dashboards, data sources, alert rules, and configuration files daily. Prometheus data can be recreated, but historical metrics are valuable during incident reviews.
- Monitor the monitoring server itself with Node Exporter, including disk usage, memory, CPU, and RAID health.
- Use RAID 1 NVMe drives for important production monitoring data. RAID is not a backup.
- Patch Ubuntu and update Prometheus, Grafana, Node Exporter, and Alertmanager on a scheduled maintenance window.
- Use private VLANs, VPN access, or firewall allowlists for metrics ports. Do not expose Node Exporter or Prometheus endpoints publicly.
- Set alerts for certificate expiry, backup failures, RAID degradation, and Prometheus TSDB storage growth.