CRITICAL INFRA
Loading critical CVEs…
ALL EXPLOITED
Loading…

Monitoring infrastructura cu Grafana + Prometheus + Zabbix

Monitoring-ul nu e doar pentru detectare probleme. E pentru intelegerea reala a infrastructurii — capacity planning, trend analysis, debugging.

Stratificare monitoring

De ce nu doar un instrument

Prometheus excelent pentru metrics time-series, slab pe long-term retention. Zabbix mai bun pe infrastructura clasica si SNMP. Nagios/Check_MK potrivit pentru check-uri sintetice. Combinatia = best of breeds.

Alerting inteligent

Cel mai prost lucru: alert fatigue. Configuram severity levels clare, escalation chains, quiet hours pentru P3/P4, alert grouping cu Alertmanager, auto-resolve cand metric revine la normal.

Capacity planning

Monitoring nu e doar pentru "ce e rau acum", e pentru "ce va fi peste 6 luni". Trends de crestere CPU/RAM/storage cu predictie linear regression. Decizii bugetare informate de date reale.

Ce livram

Stack complet Prometheus + Grafana + Zabbix + Loki, dashboards per audienta (tehnic/operational/executive), alerting cu severity calibration, runbooks per alert tip.

Exemplu: scrape Prometheus

Configurare scrape pentru node_exporter si o alerta de disk:

# prometheus.yml — scrape node_exporter
scrape_configs:
  - job_name: 'nodes'
    static_configs:
      - targets: ['10.0.0.11:9100','10.0.0.12:9100']
# alerta simpla (rules.yml): disk > 85%
#   expr: 100-(node_filesystem_avail_bytes/node_filesystem_size_bytes*100) > 85

node_exporter + scrape Prometheus

Instalezi exporterul pe hosturi si il adaugi la Prometheus:

# pe fiecare host
apt -y install prometheus-node-exporter
# prometheus.yml
scrape_configs:
  - job_name: nodes
    static_configs:
      - targets: ['10.0.0.11:9100','10.0.0.12:9100']
promtool check config /etc/prometheus/prometheus.yml

Reguli de alertare

O regula care alerteaza cand discul e aproape plin:

# rules.yml
groups:
- name: disk
  rules:
  - alert: DiskAproapePlin
    expr: 100 - (node_filesystem_avail_bytes/node_filesystem_size_bytes*100) > 85
    for: 10m
    labels: { severity: warning }
    annotations: { summary: 'Disk > 85% pe {{ $labels.instance }}' }

Grafana ca-cod (provisioning)

Datasource si dashboards versionate, nu facute manual:

# /etc/grafana/provisioning/datasources/prom.yml
apiVersion: 1
datasources:
  - name: Prometheus
    type: prometheus
    url: http://localhost:9090
    isDefault: true
# dashboards din fisiere JSON in git:
# /etc/grafana/provisioning/dashboards/ -> path catre .json
systemctl restart grafana-server

Loguri centralizate: Loki + promtail

Trimiti logurile in Loki si le cauti in Grafana:

# promtail (agent pe host) - scrape_configs
scrape_configs:
  - job_name: system
    static_configs:
      - targets: [localhost]
        labels: { job: varlogs, __path__: /var/log/*.log }
# in Grafana (Explore):  {job="varlogs"} |= "error"
Discutam despre proiectul tau →