Where they started
Every piece of customer feedback required lengthy analysis: reading the documents, finding the relevant reports, establishing correlations by hand. Volumes kept growing, and the data, medical and subject to the GDPR, could under no circumstances leave the company.
The challenge
- No data leaves the building: models must run locally, not in a provider’s cloud.
- A controlled budget: finding the hardware with the best performance-to-cost ratio, without investing in a GPU server room.
- Very new territory: open-weight models and the MCP protocol were changing month by month. At times, the question was not “how” but “is this even possible”.
What we delivered
- Infrastructure: a cluster of on-premise machines, with custom orchestration that spreads requests across machines while managing context size and task scheduling. Every machine is put to work, and usage stays smooth.
- Models: comparison and selection of open-weight models (Llama, Qwen, Gemma), tested in real conditions with RAG and business tools.
- RAG: over internal databases and knowledge bases.
- MCP servers we built ourselves: the assistant queries the company’s business software directly.
Outcome
- In production: the assistant is used every day.
- Several hours saved per case: document analysis and correlation work, once manual, are now assisted.
- Budget: the MVP was delivered at two thirds of the estimated budget, thanks to careful management of hardware resources.
- Next: the client is considering automating other parts of its business.

