IT service provider

From 15 incidents a day to 1 or 2 a month: open-source observability for 1,500 services

An IT service provider was monitoring 1,500 services and applications with a proprietary tool at the end of its rope. Its teams faced around fifteen incidents a day, and on-call engineers were woken several times a night. We built a highly available, open-source observability platform. Incidents dropped to one or two a month.

Where they started

The client’s proprietary monitoring solution had reached its limits:

  • no way to group signals or model dependencies between services;
  • performance problems;
  • data volumes that were harder and harder to absorb.

What we delivered

  • A complete observability platform: OpenTelemetry for collection, Mimir and VictoriaMetrics for metrics, Loki for logs, Tempo for traces.
  • High availability: load balancing and automatic failover, so the tool that watches production never becomes a point of failure itself.
  • Service discovery: every new service is monitored automatically, with no manual configuration, across all 1,500 services.
  • Custom, Prometheus-compatible tools and probes: synthetic tests and regression tests built into the platform, with their own data volumes.
  • Existing tools: legacy tools are connected to the platform rather than left aside.

The turning point

Both solutions were running side by side when the old tool went down. The client switched quickly to the new platform, and never went back.

Outcome

  • Incidents: from around fifteen a day to one or two a month.
  • On call: once woken several times a night, engineers are now rarely disturbed.
  • Cost: a much lower total cost, with no proprietary licences or support.
  • Operations: the client runs the platform day to day. We handle its evolution and remain on hand as backup.

Other case studies

  • Video software vendor

    Air-gapped Kubernetes, embedded in trucks

    Context
    No prior experience with container orchestration, and an ambitious goal: Kubernetes clusters embedded in trucks that operate anywhere in the world, from jungle to desert, with no connectivity at all.
    What we did
    A tailor-made, fully air-gapped Kubernetes cluster, deployed through versioned and tested Ansible playbooks, and scaling of the client’s application.
    Outcome
    A self-contained platform that runs without any connection, and that the client still builds on today.
    Read the full case study — Air-gapped Kubernetes, embedded in trucks
  • Financial institution

    Secure AWS landing zone

    Context
    Teams moving to the cloud in a scattered way: accounts created one by one, with no central organisation and no least privilege.
    What we did
    An AWS landing zone (Organizations, Control Tower, Terraform), SSO integrated with the corporate identity system, and guardrails that enforce internal rules automatically.
    Outcome
    A transformed security posture, and the cloud foundation of the whole organisation, which keeps welcoming new services.
    Read the full case study — Secure AWS landing zone
  • Healthcare start-up

    On-premise sovereign AI

    Context
    Customer feedback volumes that had become unmanageable, and medical data that must never leave the company.
    What we did
    Open-weight LLMs on custom-orchestrated on-premise hardware, RAG over internal knowledge, and MCP servers built to query business software.
    Outcome
    In production, several hours saved per case, and an MVP delivered at two thirds of the estimated budget.
    Read the full case study — On-premise sovereign AI

Contact

A similar project?

Tell us about your context. We reply within two working days with a first, concrete opinion.

Prefer a direct line?

contact@okko.be