Monitoring-ul nu e doar pentru detectare probleme. E pentru intelegerea reala a infrastructurii — capacity planning, trend analysis, debugging.
Stratificare monitoring
- Infrastructura — CPU, RAM, disk, network (Prometheus node_exporter, Zabbix agent)
- Retea — switch-uri, routere, firewall (SNMP, NetFlow, sFlow)
- Aplicatie — endpoint health, latency P50/P95/P99, error rate
- Business metrics — KPI specifici (RPS, transactions/sec, signups)
- Logs — centralizate (Loki, Graylog, Elastic Stack)
- Distributed tracing — Jaeger pentru microservices
De ce nu doar un instrument
Prometheus excelent pentru metrics time-series, slab pe long-term retention. Zabbix mai bun pe infrastructura clasica si SNMP. Nagios/Check_MK potrivit pentru check-uri sintetice. Combinatia = best of breeds.
Alerting inteligent
Cel mai prost lucru: alert fatigue. Configuram severity levels clare, escalation chains, quiet hours pentru P3/P4, alert grouping cu Alertmanager, auto-resolve cand metric revine la normal.
Capacity planning
Monitoring nu e doar pentru "ce e rau acum", e pentru "ce va fi peste 6 luni". Trends de crestere CPU/RAM/storage cu predictie linear regression. Decizii bugetare informate de date reale.
Ce livram
Stack complet Prometheus + Grafana + Zabbix + Loki, dashboards per audienta (tehnic/operational/executive), alerting cu severity calibration, runbooks per alert tip.
Exemplu: scrape Prometheus
Configurare scrape pentru node_exporter si o alerta de disk:
# prometheus.yml — scrape node_exporter
scrape_configs:
- job_name: 'nodes'
static_configs:
- targets: ['10.0.0.11:9100','10.0.0.12:9100']
# alerta simpla (rules.yml): disk > 85%
# expr: 100-(node_filesystem_avail_bytes/node_filesystem_size_bytes*100) > 85
node_exporter + scrape Prometheus
Instalezi exporterul pe hosturi si il adaugi la Prometheus:
# pe fiecare host
apt -y install prometheus-node-exporter
# prometheus.yml
scrape_configs:
- job_name: nodes
static_configs:
- targets: ['10.0.0.11:9100','10.0.0.12:9100']
promtool check config /etc/prometheus/prometheus.yml
Reguli de alertare
O regula care alerteaza cand discul e aproape plin:
# rules.yml
groups:
- name: disk
rules:
- alert: DiskAproapePlin
expr: 100 - (node_filesystem_avail_bytes/node_filesystem_size_bytes*100) > 85
for: 10m
labels: { severity: warning }
annotations: { summary: 'Disk > 85% pe {{ $labels.instance }}' }
Grafana ca-cod (provisioning)
Datasource si dashboards versionate, nu facute manual:
# /etc/grafana/provisioning/datasources/prom.yml
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
url: http://localhost:9090
isDefault: true
# dashboards din fisiere JSON in git:
# /etc/grafana/provisioning/dashboards/ -> path catre .json
systemctl restart grafana-server
Loguri centralizate: Loki + promtail
Trimiti logurile in Loki si le cauti in Grafana:
# promtail (agent pe host) - scrape_configs
scrape_configs:
- job_name: system
static_configs:
- targets: [localhost]
labels: { job: varlogs, __path__: /var/log/*.log }
# in Grafana (Explore): {job="varlogs"} |= "error"