Monitoring is for real understanding of infrastructure — capacity planning, trend analysis, debugging.
Monitoring stratification
- Infrastructure (Prometheus, Zabbix)
- Network (SNMP, NetFlow)
- Application (latency P50/P95/P99)
- Business metrics
- Centralized logs
- Distributed tracing (Jaeger)
What we deliver
Complete Prometheus + Grafana + Zabbix + Loki stack, dashboards per audience, alerting with severity calibration.
Example: Prometheus scrape
A scrape config for node_exporter and a disk alert:
# prometheus.yml — scrape node_exporter
scrape_configs:
- job_name: 'nodes'
static_configs:
- targets: ['10.0.0.11:9100','10.0.0.12:9100']
# alerta simpla (rules.yml): disk > 85%
# expr: 100-(node_filesystem_avail_bytes/node_filesystem_size_bytes*100) > 85
node_exporter + Prometheus scrape
You install the exporter on hosts and add it to Prometheus:
# pe fiecare host
apt -y install prometheus-node-exporter
# prometheus.yml
scrape_configs:
- job_name: nodes
static_configs:
- targets: ['10.0.0.11:9100','10.0.0.12:9100']
promtool check config /etc/prometheus/prometheus.yml
Alerting rules
A rule that alerts when the disk is nearly full:
# rules.yml
groups:
- name: disk
rules:
- alert: DiskAproapePlin
expr: 100 - (node_filesystem_avail_bytes/node_filesystem_size_bytes*100) > 85
for: 10m
labels: { severity: warning }
annotations: { summary: 'Disk > 85% pe {{ $labels.instance }}' }
Grafana as code (provisioning)
Datasources and dashboards versioned, not clicked by hand:
# /etc/grafana/provisioning/datasources/prom.yml
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
url: http://localhost:9090
isDefault: true
# dashboards din fisiere JSON in git:
# /etc/grafana/provisioning/dashboards/ -> path catre .json
systemctl restart grafana-server
Centralized logs: Loki + promtail
You ship logs to Loki and query them in Grafana:
# promtail (agent pe host) - scrape_configs
scrape_configs:
- job_name: system
static_configs:
- targets: [localhost]
labels: { job: varlogs, __path__: /var/log/*.log }
# in Grafana (Explore): {job="varlogs"} |= "error"