Healthcare start-up

On-premise sovereign AI for a healthcare start-up

Customer feedback volumes that had become unmanageable, medical data that cannot leave the company, and a start-up budget, not an enterprise one. We designed an AI assistant that runs entirely on the client’s own hardware. It is in production, and the MVP was delivered at two thirds of the planned budget.

Where they started

Every piece of customer feedback required lengthy analysis: reading the documents, finding the relevant reports, establishing correlations by hand. Volumes kept growing, and the data, medical and subject to the GDPR, could under no circumstances leave the company.

The challenge

  • No data leaves the building: models must run locally, not in a provider’s cloud.
  • A controlled budget: finding the hardware with the best performance-to-cost ratio, without investing in a GPU server room.
  • Very new territory: open-weight models and the MCP protocol were changing month by month. At times, the question was not “how” but “is this even possible”.

What we delivered

  • Infrastructure: a cluster of on-premise machines, with custom orchestration that spreads requests across machines while managing context size and task scheduling. Every machine is put to work, and usage stays smooth.
  • Models: comparison and selection of open-weight models (Llama, Qwen, Gemma), tested in real conditions with RAG and business tools.
  • RAG: over internal databases and knowledge bases.
  • MCP servers we built ourselves: the assistant queries the company’s business software directly.

Outcome

  • In production: the assistant is used every day.
  • Several hours saved per case: document analysis and correlation work, once manual, are now assisted.
  • Budget: the MVP was delivered at two thirds of the estimated budget, thanks to careful management of hardware resources.
  • Next: the client is considering automating other parts of its business.

Other case studies

  • Video software vendor

    Air-gapped Kubernetes, embedded in trucks

    Context
    No prior experience with container orchestration, and an ambitious goal: Kubernetes clusters embedded in trucks that operate anywhere in the world, from jungle to desert, with no connectivity at all.
    What we did
    A tailor-made, fully air-gapped Kubernetes cluster, deployed through versioned and tested Ansible playbooks, and scaling of the client’s application.
    Outcome
    A self-contained platform that runs without any connection, and that the client still builds on today.
    Read the full case study — Air-gapped Kubernetes, embedded in trucks
  • Financial institution

    Secure AWS landing zone

    Context
    Teams moving to the cloud in a scattered way: accounts created one by one, with no central organisation and no least privilege.
    What we did
    An AWS landing zone (Organizations, Control Tower, Terraform), SSO integrated with the corporate identity system, and guardrails that enforce internal rules automatically.
    Outcome
    A transformed security posture, and the cloud foundation of the whole organisation, which keeps welcoming new services.
    Read the full case study — Secure AWS landing zone
  • IT service provider

    Open-source observability for 1,500 services

    Context
    A proprietary monitoring stack at the end of its rope: no way to group signals or model dependencies, slowness, and around fifteen incidents a day.
    What we did
    A highly available observability platform (OpenTelemetry, Mimir, VictoriaMetrics, Loki, Tempo), service discovery, and custom Prometheus-compatible probes.
    Outcome
    From fifteen incidents a day to one or two a month, and on-call engineers who are hardly ever woken at night.
    Read the full case study — Open-source observability for 1,500 services

Contact

A similar project?

Tell us about your context. We reply within two working days with a first, concrete opinion.

Prefer a direct line?

contact@okko.be