Services / AI · LLM
AI/LLM Product Engineering
An LLM API call is one request. An AI product is the routing, memory, evaluation, limits, cost control, failure handling and business logic around it. Closeflow designs and builds that layer. We are not a chatbot agency wrapping a single API.
02What we build
What we build
LLM gateways and multi-provider routing
One entry point to several providers, with model choice driven by task, cost and availability.
RAG and semantic search
Retrieval over your own data using embeddings and pgvector, and context assembly for the answer.
Embeddings and long-term memory
What the system remembers about a user or session, how it is retrieved and how it reaches the prompt.
Agent and workflow pipelines
Multi-step processing across models, background jobs and checkpoints at each step.
AI features inside existing products
Wired into your backend, permissions and billing rather than a parallel system next to them.
AI-assisted internal tools
For example an analyst that answers questions through a read-only database session and prepares reports for an operations team.
Evaluation
Model and prompt comparison on real scenarios, with regression checks.
Billing and quotas
Token accounting, prepaid balances, per-plan and per-user limits.
Moderation and guard pipelines
Checks on the request and the response: jailbreak attempts, drifting out of the assigned persona.
03Gateway
The LLM gateway: how a request travels
A generalised view of the gateway we built in AI-platform projects under NDA.
A circuit breaker per provider: when one starts failing or slowing down, traffic stops going there instead of piling up timeouts behind it.
Fallback: when a provider is unavailable, the request goes to another model of a comparable class. The user sees a short delay, not an error.
Tier and intent routing: simple requests go to cheaper models, harder ones to stronger models. The decision is made on task type and cost.
Graceful degradation: if a model is unavailable, the system offers an available replacement rather than failing as a whole. Responses stream over WebSocket or SSE.
04Memory
Memory, retrieval and context
This is about how a system stores and finds information. It does not involve training models.
Three-layer memory
Short-term conversation context in Redis, structured facts in PostgreSQL, semantic retrieval over embeddings in pgvector.
Semantic cache
Similar requests reuse a previously found answer from Postgres and pgvector instead of calling a provider again.
Session summarisation
Long history is compressed into a summary, and what matters is saved to long-term memory.
Prompt assembly under a token budget
Context is composed from prioritised layers to fit the budget, rather than being cut off arbitrarily.
RAG and retrieval
Finding relevant passages by meaning and inserting them into the request together with their source.
05Evaluation
Behaviour is measured, not guessed
Switching a model or editing a prompt without a check is an experiment run on your users.
Model comparison
The same scenarios run through different models and are compared on quality and price.
LLM-as-judge
A judge model ranks answers against defined criteria. It is a fast check, best complemented by human review.
Regression checks
Before a model switch or a prompt change, a scenario set shows what got worse.
A/B experiments
Variants are compared on real usage scenarios, not synthetic examples.
06Workflows
Node-based pipelines
In one AI platform we built a visual node editor: users assemble a pipeline on a canvas, and server-side workers execute the steps. It is not a general-purpose workflow engine. It is an applied tool shaped around that platform’s scenarios.
Each node can use its own provider and model. Every step exposes status, timing and output, so when something fails you can see where processing stopped. Long steps run in the background and do not block the interface.
07Cost and reliability
Cost, limits and reliability
Prepaid balance with Redis holds
The amount is reserved before the model call and settled afterwards, so concurrent requests cannot overdraw a balance.
Payment reconcile worker
Provider payments and internal records are compared on a schedule, and signatures are verified at intake.
Provider fallback
One provider being down does not stop the product.
Token accounting
The cost of every request is recorded against user, model and plan.
Observability
Prometheus and Grafana dashboards for errors, latency and spend surface problems before users report them.
08Stack
What we use
Server
Python where it is needed.
Interface
Data
Operations
09Boundaries
What this is not
Closeflow is not a foundation-model research lab and does not train large models from scratch. The work is production AI product engineering: the infrastructure around models and their dependable integration into applications.
10When it fits
When to get in touch
A prototype must become a dependable feature
The demo works, but there are no limits, retries, monitoring or defined behaviour on error.
Cost and quality are unpredictable
Spend is growing and it is unclear which model suits which task.
You need to add several providers
You want to avoid dependence on a single API and pick models per task.
AI has to live inside an existing backend
The feature must use your data, permissions and billing.
Have an AI feature in mind?
Tell us what it should do and where it gets stuck today. We will walk through the architecture and say what is realistic.
Discuss an AI feature12Contact
Tell us what
has to work.
Describe the problem, your current system and its constraints. If it’s a good fit, we’ll suggest a next step — usually a short call about data, integrations and timelines.
- Telegram@missuk2003
- Emailceo@closeflow.ru
- GitHubHusqvarnalox