Fastly, Multilingual AI Answers that Stay on Your Infrastructure.
A private answer engine for teams that need speed and cannot send their data to a third party, serving nine languages from a stack they host themselves.
- Industry
- AI
- Services
- AI development, Cloud & DevOps
- Timeline
- 12 weeks to launch
- Engagement
- Fixed scope

The problem
A hosted API was the obvious answer and the legal answer was no: nothing touching customer records could leave the client's infrastructure.The obvious answer was a hosted API, and the legal answer was no. Anything touching customer records had to stay inside infrastructure the client controlled, which ruled out most of what was easy.
What we did
An open model on the client's own cloud with retrieval in front, plus language routing so each query uses the right model.We ran an open model on the client's own cloud with retrieval in front of it, and spent the time on caching and batching rather than on prompt tricks. Language routing picks the right model per request, so a Hindi query does not pay for an English-tuned one.
The result
First token under half a second across nine languages, roughly a third the hosted cost, and no data leaving the tenant.First token in under half a second across nine languages, at roughly a third of what the hosted equivalent would have cost, and no customer data ever left the tenant.
They were the first team who did not try to talk us out of the constraint. They designed around it and it turned out cheaper anyway.
Building Something Like This?
Tell us where you are. We will tell you honestly what it takes.Tell us where you are. We will tell you honestly what it takes and what it costs.
Start a conversation