Discuss a project

Case · under NDA

AI Platform Infrastructure Beyond the Model API

Calling a model takes one line of code. Routing across providers, memory, quality evaluation, limits and money take the rest, and that is the platform. This page describes how we built that layer for several consumer AI services. Details are generalised because of NDA.

  • Go · Next.js
  • PostgreSQL · pgvector · Redis
  • multi-LLM · eval · billing
  • under NDA

02Problem

The model call is the easy part

The work spans several consumer AI services. Because of the NDA we do not name the products or describe their domain, but the engineering problems of such systems are similar, and those are what this page is about.

Providers fail and slow down. Request cost and response latency vary significantly across models and providers, which makes model selection part of the product logic. Answer quality drifts after a model swap or a prompt edit. And a user’s money has to add up exactly, even when a call breaks halfway through a stream.

So a separate layer is built around the model: a gateway, memory, evaluation, billing and observability. In that layer the model is just one of the dependencies.

03Architecture

The platform around the LLM

A generalised view. The real topology is more detailed, but the roles of the nodes are preserved.

Diagram: Client, Go API, Billing, Memory, LLM gateway, Workflows, Providers, EvalsClientweb · WebSocket/SSEGo APIsessions · accessBillingRedis holdsMemorypgvector · RedisLLM gatewayrouting · breakerWorkflowscanvas · workersProvidersfallback chainEvalsLLM judge
fig. The gateway picks provider and model, billing reserves the balance before the call, and memory and evaluation run as their own loops.

04Routing

The gateway: routing and fault tolerance

The gateway is the single entry point to every model provider. Product code does not know whose model answered.

  • A circuit breaker per provider

    When a provider starts erroring or slowing down, traffic stops going there, so timeouts do not pile up behind it.

  • Fallback chain

    If a provider is unavailable, the request goes to another model of a comparable class. The user sees a short delay instead of an error.

  • Tier and intent routing

    Simple requests go to cheaper models and harder ones to stronger models. The decision rests on task type and cost.

  • Graceful degradation

    If the preferred model is down, the system offers an available replacement instead of failing as a whole.

  • Streaming

    The answer reaches the client as it is generated, over WebSocket or SSE, so nobody waits for the full text.

05Memory

Memory and context

This is about storing and finding information. It does not involve training or fine-tuning models, only retrieving the right data.

  • Semantic cache

    For requests close in meaning, an earlier answer found in PostgreSQL and pgvector is reused instead of calling a provider again.

  • Three-layer memory

    Short-term conversation context sits in Redis, structured facts in PostgreSQL, and semantic retrieval over embeddings in pgvector.

  • Session summarisation

    A long history is compressed into a summary, and what matters is saved to long-term memory.

  • Prompt under a token budget

    Context is composed from prioritised layers to fit the budget, rather than being cut off arbitrarily.

  • Retrieval instead of tuning

    New knowledge reaches the answer by retrieving relevant passages, while the model itself stays unchanged.

06Evaluation

Quality is measured, not guessed

Swapping a model or editing a prompt without a check is an experiment on users. That is why evaluation lives inside the server.

  • LLM-as-judge

    A judge model scores answers against defined criteria. It is a fast check, best complemented by human review.

  • Model and prompt comparison

    The same scenarios run through different models and prompt versions, and the results are compared.

  • Regression checks

    Before a model switch, a scenario set shows what got worse.

  • A/B experiments

    Variants are compared on real usage scenarios rather than synthetic examples.

07Workflows

A visual pipeline editor

The platform includes a canvas editor: a user assembles a pipeline from nodes, and server-side workers execute the steps. It is an applied tool shaped around that platform’s scenarios, not a general-purpose or distributed workflow engine.

Each node can use its own model. Every step exposes status, timing and output, so when something fails you can see where processing stopped. Long steps run in the background and do not block the interface.

08Billing

The money has to add up

A prepaid balance demands care because a model call costs money before its outcome is known.

  • Redis holds

    The amount is reserved before the call, settled by actual cost after a successful answer, and released on failure. Concurrent requests cannot push a balance below zero.

  • A stream that breaks midway

    If a provider cuts an answer off, the balance stays correct: actual usage is charged and the surplus reservation is returned.

  • Payment reconcile worker

    On a schedule it compares provider payments with internal records and raises any mismatch.

  • Webhook signature checks

    Notifications from the payment provider are accepted only after their signature is verified.

09Guards

Checks before and after generation

Requests and responses pass through a guard pipeline: jailbreak attempts and drift away from the assigned persona are caught before the model call and again after it. It protects the product from abuse and from answers that fall outside the service’s character.

10Observability

Seeing what is going on

  • Prometheus and Grafana

    Several dashboards cover provider errors, latency, spend and queue state.

  • Structured logs

    Machine-readable records make it possible to follow one specific request.

  • Versioned migrations

    The PostgreSQL schema changes through versioned migrations, so database state is reproducible.

  • Docker delivery

    Services run the same way locally and in deployed environments.

11Result

What came out of it

The result is a platform where the model is a replaceable dependency, while routing, memory, billing and quality evaluation are separate system loops.

A single provider can be swapped without rewriting the product, spend is counted explicitly, and a model change can be checked before it ships. That is what separates an AI product from a wrapper around one API.

12Stack

Technologies

Server

  • Go
  • WebSocket
  • SSE

Interface

  • Next.js
  • React

Data

  • PostgreSQL
  • pgvector
  • Redis

Operations

  • Prometheus
  • Grafana
  • Docker

Building an AI product?

Tell us what it should do and where it hurts today. We will walk through the architecture and say what is realistic.

Discuss an AI product

14Contact

Tell us what
has to work.

Describe the problem, your current system and its constraints. If it’s a good fit, we’ll suggest a next step — usually a short call about data, integrations and timelines.

brief.form4 fields · 2 minutes

Budget optional