Case · under NDA
AI Platform Infrastructure Beyond the Model API
Calling a model takes one line of code. Routing across providers, memory, quality evaluation, limits and money take the rest, and that is the platform. This page describes how we built that layer for several consumer AI services. Details are generalised because of NDA.
02Problem
The model call is the easy part
The work spans several consumer AI services. Because of the NDA we do not name the products or describe their domain, but the engineering problems of such systems are similar, and those are what this page is about.
Providers fail and slow down. Request cost and response latency vary significantly across models and providers, which makes model selection part of the product logic. Answer quality drifts after a model swap or a prompt edit. And a user’s money has to add up exactly, even when a call breaks halfway through a stream.
So a separate layer is built around the model: a gateway, memory, evaluation, billing and observability. In that layer the model is just one of the dependencies.
03Architecture
The platform around the LLM
A generalised view. The real topology is more detailed, but the roles of the nodes are preserved.
04Routing
The gateway: routing and fault tolerance
The gateway is the single entry point to every model provider. Product code does not know whose model answered.
A circuit breaker per provider
When a provider starts erroring or slowing down, traffic stops going there, so timeouts do not pile up behind it.
Fallback chain
If a provider is unavailable, the request goes to another model of a comparable class. The user sees a short delay instead of an error.
Tier and intent routing
Simple requests go to cheaper models and harder ones to stronger models. The decision rests on task type and cost.
Graceful degradation
If the preferred model is down, the system offers an available replacement instead of failing as a whole.
Streaming
The answer reaches the client as it is generated, over WebSocket or SSE, so nobody waits for the full text.
05Memory
Memory and context
This is about storing and finding information. It does not involve training or fine-tuning models, only retrieving the right data.
Semantic cache
For requests close in meaning, an earlier answer found in PostgreSQL and pgvector is reused instead of calling a provider again.
Three-layer memory
Short-term conversation context sits in Redis, structured facts in PostgreSQL, and semantic retrieval over embeddings in pgvector.
Session summarisation
A long history is compressed into a summary, and what matters is saved to long-term memory.
Prompt under a token budget
Context is composed from prioritised layers to fit the budget, rather than being cut off arbitrarily.
Retrieval instead of tuning
New knowledge reaches the answer by retrieving relevant passages, while the model itself stays unchanged.
06Evaluation
Quality is measured, not guessed
Swapping a model or editing a prompt without a check is an experiment on users. That is why evaluation lives inside the server.
LLM-as-judge
A judge model scores answers against defined criteria. It is a fast check, best complemented by human review.
Model and prompt comparison
The same scenarios run through different models and prompt versions, and the results are compared.
Regression checks
Before a model switch, a scenario set shows what got worse.
A/B experiments
Variants are compared on real usage scenarios rather than synthetic examples.
07Workflows
A visual pipeline editor
The platform includes a canvas editor: a user assembles a pipeline from nodes, and server-side workers execute the steps. It is an applied tool shaped around that platform’s scenarios, not a general-purpose or distributed workflow engine.
Each node can use its own model. Every step exposes status, timing and output, so when something fails you can see where processing stopped. Long steps run in the background and do not block the interface.
08Billing
The money has to add up
A prepaid balance demands care because a model call costs money before its outcome is known.
Redis holds
The amount is reserved before the call, settled by actual cost after a successful answer, and released on failure. Concurrent requests cannot push a balance below zero.
A stream that breaks midway
If a provider cuts an answer off, the balance stays correct: actual usage is charged and the surplus reservation is returned.
Payment reconcile worker
On a schedule it compares provider payments with internal records and raises any mismatch.
Webhook signature checks
Notifications from the payment provider are accepted only after their signature is verified.
09Guards
Checks before and after generation
Requests and responses pass through a guard pipeline: jailbreak attempts and drift away from the assigned persona are caught before the model call and again after it. It protects the product from abuse and from answers that fall outside the service’s character.
10Observability
Seeing what is going on
Prometheus and Grafana
Several dashboards cover provider errors, latency, spend and queue state.
Structured logs
Machine-readable records make it possible to follow one specific request.
Versioned migrations
The PostgreSQL schema changes through versioned migrations, so database state is reproducible.
Docker delivery
Services run the same way locally and in deployed environments.
11Result
What came out of it
The result is a platform where the model is a replaceable dependency, while routing, memory, billing and quality evaluation are separate system loops.
A single provider can be swapped without rewriting the product, spend is counted explicitly, and a model change can be checked before it ships. That is what separates an AI product from a wrapper around one API.
12Stack
Technologies
Server
Interface
Data
Operations
Building an AI product?
Tell us what it should do and where it hurts today. We will walk through the architecture and say what is realistic.
Discuss an AI product14Contact
Tell us what
has to work.
Describe the problem, your current system and its constraints. If it’s a good fit, we’ll suggest a next step — usually a short call about data, integrations and timelines.
- Telegram@missuk2003
- Emailceo@closeflow.ru
- GitHubHusqvarnalox