Where they started
The client’s proprietary monitoring solution had reached its limits:
- no way to group signals or model dependencies between services;
- performance problems;
- data volumes that were harder and harder to absorb.
What we delivered
- A complete observability platform: OpenTelemetry for collection, Mimir and VictoriaMetrics for metrics, Loki for logs, Tempo for traces.
- High availability: load balancing and automatic failover, so the tool that watches production never becomes a point of failure itself.
- Service discovery: every new service is monitored automatically, with no manual configuration, across all 1,500 services.
- Custom, Prometheus-compatible tools and probes: synthetic tests and regression tests built into the platform, with their own data volumes.
- Existing tools: legacy tools are connected to the platform rather than left aside.
The turning point
Both solutions were running side by side when the old tool went down. The client switched quickly to the new platform, and never went back.
Outcome
- Incidents: from around fifteen a day to one or two a month.
- On call: once woken several times a night, engineers are now rarely disturbed.
- Cost: a much lower total cost, with no proprietary licences or support.
- Operations: the client runs the platform day to day. We handle its evolution and remain on hand as backup.

