Monitoring ist für echtes Verständnis der Infrastruktur.
Stack
Prometheus, Grafana, Zabbix, Loki, Jaeger.
Beispiel: Prometheus-Scrape
Eine Scrape-Konfiguration fuer node_exporter und ein Disk-Alarm:
# prometheus.yml — scrape node_exporter
scrape_configs:
- job_name: 'nodes'
static_configs:
- targets: ['10.0.0.11:9100','10.0.0.12:9100']
# alerta simpla (rules.yml): disk > 85%
# expr: 100-(node_filesystem_avail_bytes/node_filesystem_size_bytes*100) > 85
node_exporter + Prometheus-Scrape
Sie installieren den Exporter auf Hosts und fuegen ihn zu Prometheus hinzu:
# pe fiecare host
apt -y install prometheus-node-exporter
# prometheus.yml
scrape_configs:
- job_name: nodes
static_configs:
- targets: ['10.0.0.11:9100','10.0.0.12:9100']
promtool check config /etc/prometheus/prometheus.yml
Alarmierungsregeln
Eine Regel, die alarmiert, wenn die Platte fast voll ist:
# rules.yml
groups:
- name: disk
rules:
- alert: DiskAproapePlin
expr: 100 - (node_filesystem_avail_bytes/node_filesystem_size_bytes*100) > 85
for: 10m
labels: { severity: warning }
annotations: { summary: 'Disk > 85% pe {{ $labels.instance }}' }
Grafana als Code (Provisioning)
Datasource und Dashboards versioniert, nicht per Hand geklickt:
# /etc/grafana/provisioning/datasources/prom.yml
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
url: http://localhost:9090
isDefault: true
# dashboards din fisiere JSON in git:
# /etc/grafana/provisioning/dashboards/ -> path catre .json
systemctl restart grafana-server
Zentrale Logs: Loki + promtail
Sie senden Logs an Loki und durchsuchen sie in Grafana:
# promtail (agent pe host) - scrape_configs
scrape_configs:
- job_name: system
static_configs:
- targets: [localhost]
labels: { job: varlogs, __path__: /var/log/*.log }
# in Grafana (Explore): {job="varlogs"} |= "error"