AI integration · Sample

AI assistant over company documentation

RAG with hybrid search and re-ranking over thousands of documents. RAGAS evaluation and cost-aware model routing.

Client
B2B services company (sample)
Year
2026
Scope
PoCRAGEvaluationFinOps
Stack
PythonFastAPIpgvectorClaudeOllama
AI integration
0.87faithfulness
-62%LLM cost
<2 sp95 response

Reference architecture

ORCHESTRATIONMODELS & KNOWLEDGEUserchat · API · agentAPIauth · rate limitLLM routercost ↔ qualityRetrieverhybrid + rerankModelsClaude · OllamaSemantic cacheRedispgvectorembeddingsKnowledgePDF · Confluence · DBDWG · AI-01 · nightdev
  1. Userchat · API · agent
  2. APIauth · rate limit
  3. LLM routercost ↔ qualityRetrieverhybrid + rerank
  4. ModelsClaude · OllamaSemantic cacheRedispgvectorembeddingsKnowledgePDF · Confluence · DB
Reference architecture · AI assistant over company documentation

This is a sample case study showing the format. Replace it with a real project.

Problem

The support team searched ~4k documents (PDFs, Confluence, procedures). Finding the right answer took well over ten minutes on average, and new hires needed months to find their way around.

Decisions

  • Hybrid search (BM25 + pgvector embeddings) with re-ranking — pure vector search missed proper names and procedure numbers.
  • Every answer cites its sources with a link to the document fragment.
  • Model router: simple questions go to a cheaper model, complex ones to a stronger one. Semantic cache for repeated questions.
  • A 200-question test set with RAGAS evaluation run on every prompt or index change.

Outcome

  • Faithfulness 0.87 on the test set.
  • LLM cost -62% compared to one large model for everything.
  • p95 response time under 2 seconds.

Got a project that has to be fast and work at scale?

Describe it in 2 minutes. I'll reply within 24 hours with first insights and a proposed next step.