Problem
An AI product in a demo and an AI product under load are different systems. Model providers go down and change their prices, conversations lose context, image generation needs GPUs, and every extra thousand tokens is a direct cost.
Architecture
Not a wrapper around one LLM API, but a platform. Go backend, Next.js and React front end, streaming chat over WebSocket and SSE. A single gateway across several LLM providers: smart routing by model and cost, circuit breakers, fallback routes, token counting and billing. On top of it — memory and a semantic cache on pgvector, a node-based workflow editor with its own execution engine, and eval harnesses. A separate pipeline handles LoRA training for SDXL and FLUX and generation through ComfyUI on rented GPUs.
Engineering decisions
- Routing between providers: model choice by task and cost, circuit breakers that cut off a failing provider, retries and fallback routes — one model going down doesn’t take the product with it.
- Three-layer memory: short-term context in Redis, structured facts in PostgreSQL, semantic retrieval in pgvector — with confidence weighting. Prompts are assembled against a token budget.
- A semantic response cache and an eval harness with an LLM judge: quality and cost are measured, not guessed; A/B experiments on real scenarios.
- Billing integrity: balance holds in Redis and a background worker that reconciles payments. Safety and persona guards before and after generation.
- A node-based workflow engine: a graph editor in the UI and server-side step execution, with observability per node.
- A separate GPU pipeline: LoRA training, ComfyUI generation and character consistency across generations, on capacity rented from RunPod and Vast.ai.
- Observability and load testing: Prometheus and Grafana metrics, k6 load testing, zero-downtime migrations, backups and CI / CD.
Outcome
Systems built from architecture through release and the operational layer, with billing, monitoring and fallback paths. Client details are under NDA — what’s described here is the engineering, not a specific implementation.