PROPELOO

RAG / RETRIEVAL-AUGMENTED GENERATION

Build AI that answers from your data, not from hallucination.

PROPELOO engineers RAG systems — from document ingestion pipelines and chunking strategy through vector database design, hybrid search, reranking and the evaluation framework that tells you when the system is actually working. RAG is not an API call. It is a retrieval engineering problem dressed up as an AI problem.

A RAG system with poor retrieval produces confident, wrong answers. The LLM is not the problem — the retrieval pipeline is.

Most RAG implementations fail not because of the language model but because of the retrieval layer. Chunks that are too large or too small produce irrelevant context. Semantic-only search misses exact keyword matches that matter for technical content. Embedding models that perform well on general benchmarks perform poorly on domain-specific vocabulary. No reranking means the most relevant chunk might be retrieved third instead of first — and the LLM uses the first one most. PROPELOO treats RAG as a retrieval engineering problem: we evaluate retrieval quality independently from generation quality, tune chunking and embedding on your specific content, and measure context precision before measuring answer quality.

The full RAG engineering stack.

A production RAG system has six distinct engineering layers, each of which can fail independently.

System Layers

  • Ingestion Layer: Document loading, parsing, cleaning, chunking, embedding generation, vector store indexing
  • Retrieval Layer: Query embedding, vector search, keyword search, hybrid fusion, metadata filtering
  • Reranking Layer: Cross-encoder reranking, MMR diversity, contextual compression, relevance scoring
  • Generation Layer: Prompt construction, context injection, LLM call, output validation, source citation
  • Evaluation Layer: Retrieval precision, answer faithfulness, answer relevance, hallucination detection

Core Technical Capabilities

  • Document Ingestion Pipeline

    Load from PDFs, Word, HTML, Notion, Confluence, SharePoint, databases and APIs. Parse with layout-aware extractors (Unstructured, LlamaParse). Clean noise (headers, footers, page numbers). Chunk with overlap. Generate embeddings. Upsert to vector store with metadata.

  • Chunking Strategy Optimisation

    Fixed-size with overlap, sentence-level, semantic (topic boundary detection), recursive character splitting and hierarchical (parent-child chunks for context expansion). Evaluated against retrieval precision on your specific corpus.

  • Hybrid Search

    Combine dense vector search (semantic similarity) with sparse BM25 keyword search via Reciprocal Rank Fusion (RRF). Captures both conceptual similarity and exact term matching. Consistently outperforms semantic-only retrieval for technical and domain-specific content.

  • Reranking Pipeline

    Cross-encoder reranking (Cohere Rerank, BGE Reranker) of top-K retrieved chunks. Reranker sees the query and full chunk text together — more accurate relevance than embedding similarity alone. MMR (Maximal Marginal Relevance) for result diversity.

  • Advanced RAG Patterns

    HyDE (Hypothetical Document Embeddings) for query expansion, multi-query retrieval, parent-child chunk retrieval (retrieve child for precision, expand to parent for context), RAG-fusion for multiple search strategies and self-query metadata filtering.

  • RAG Evaluation Framework

    RAGAS metrics: context precision (are retrieved chunks relevant?), context recall (are all relevant chunks retrieved?), answer faithfulness (is the answer grounded in retrieved context?), answer relevance (does the answer address the question?). LangSmith for production monitoring.

How we think about RAG engineering.

Retrieval quality determines answer quality. Improve the LLM from GPT-3.5 to GPT-4 and you get better language. Improve retrieval precision from 60% to 90% and you get better answers.

  • Measure retrieval before measuring generation

    Most RAG debugging starts at the wrong end — evaluating answer quality without first measuring whether the right context was retrieved. If retrieval precision is 60% (4 out of 10 retrieved chunks are relevant), no LLM can consistently produce correct answers. Measure context precision and context recall first. If retrieval is poor, fix the chunking, embedding model and search strategy before changing the prompt.

    Axiom:

  • Chunking is the most impactful parameter

    In our experience across many RAG systems, chunking strategy has more impact on retrieval quality than LLM model selection. A 512-token fixed-size chunk that cuts a sentence in half destroys the semantic coherence of that passage. A semantically chunked document that keeps related content together retrieves correctly more often. Evaluate chunking strategy on your specific content before finalising — the optimal approach varies by document type.

    Axiom:

  • Hybrid search outperforms semantic-only for most corpora

    Semantic search excels at conceptual matching ("what is the refund policy?" matches "how do I get my money back?"). BM25 keyword search excels at exact term matching (product codes, API names, error messages, person names). Most real-world corpora benefit from both. Reciprocal Rank Fusion of BM25 + semantic results consistently outperforms either alone on benchmarks and in production.

    Axiom:

  • RAG needs a test set, not just an architecture

    The only way to know if a RAG system is working correctly is to measure it against a test set of question-answer pairs with ground truth. Generate 50-100 representative questions with known correct answers from your corpus. Measure retrieval precision (did we retrieve the right chunks?) and answer faithfulness (is the answer grounded in retrieved content?) before launch. Without this, you are deploying a system of unknown quality.

    Axiom:

The RAG system decisions that determine quality.

Each parameter significantly affects retrieval precision and answer quality.

  • Chunking strategy?

    Impact: Start with recursive character splitting (LangChain default). Evaluate context precision. If below 70%, try semantic chunking. For Q&A over technical docs, hierarchical chunking (small chunks for retrieval, larger chunks for context) often gives the best results.

    • Fixed-size (512 tokens, 10% overlap) — simple baseline, works for homogeneous docs
    • Recursive character splitting — respects sentence/paragraph boundaries, better than fixed-size
    • Semantic chunking (split on topic change) — best quality, more complex, slower ingestion
    • Hierarchical (parent-child) — retrieve child chunks, expand to parent for context
  • Embedding model?

    Impact: Evaluate on your specific domain — general benchmark performance does not always translate. text-embedding-3-large is the safe default. BGE-M3 self-hosted for cost-sensitive high-volume pipelines.

    • text-embedding-3-large (OpenAI) — top benchmark, API dependency, $0.13/1M tokens
    • Cohere embed-v3 — strong multilingual, good for mixed-language corpora
    • BGE-M3 (self-hosted) — competitive with commercial, free, requires GPU for speed
    • Voyage AI — strong for domain-specific content
  • Vector database?

    Impact: pgvector for systems already on PostgreSQL with <2M chunks. Pinecone for managed production deployment with minimal ops. Qdrant for complex metadata filtering requirements. Do not over-engineer the vector DB choice — they are more interchangeable than the ecosystem suggests.

    • pgvector — PostgreSQL extension, consolidates stack, good to 5M vectors
    • Pinecone — fully managed, production-ready, cost scales with index size
    • Qdrant — self-hosted or managed, fast, strong metadata filtering
    • Weaviate — multi-modal support, good for mixed content types
  • Semantic-only vs hybrid search?

    Impact: Hybrid search is the production standard. The marginal engineering effort to add BM25 alongside semantic search is small. The retrieval quality improvement on any corpus with technical terms, product names or exact phrases is significant.

    • Semantic only — simpler, works well for conceptual Q&A
    • BM25 only — exact matching, misses conceptual queries
    • Hybrid (semantic + BM25, RRF fusion) — best of both, correct for most production systems
    • Graph RAG — knowledge graph retrieval, best for highly interconnected knowledge
  • Reranking?

    Impact: Cohere Rerank or BGE Reranker should be standard in production RAG systems. The precision improvement from reranking top-20 chunks to top-3 for context injection is measurable. The latency cost (200-400ms) is acceptable for most applications.

    • No reranking — top-K by embedding similarity, fast, lower precision
    • Cohere Rerank API — managed, accurate, adds ~200ms latency
    • BGE Reranker (self-hosted) — competitive quality, no API cost, requires inference server
    • LLM-based reranking — most accurate, high latency and cost
  • Evaluation frequency?

    Impact: RAGAS in CI for every configuration change. Production sampling with LLM-as-judge for ongoing quality monitoring. Without automated evaluation, you will break the system without knowing it.

    • Manual spot-check — not scalable, catches obvious failures only
    • CI evaluation on test set — automated, catches regressions on every prompt/config change
    • Online evaluation (production sampling) — measures real query distribution, requires LLM-as-judge at scale
    • RAGAS automated metrics — retrieval + generation evaluation, no ground truth required for some metrics

What PROPELOO builds.

  • Enterprise Knowledge Base

    RAG over internal documentation — Confluence, Notion, SharePoint, PDFs — with access-control-aware retrieval (users only see docs they are authorised to access) and source citation.

  • Customer Support RAG

    Support bot grounded in product documentation, FAQs and historical ticket resolutions — with confidence-based response gating and escalation to human agents for low-confidence queries.

  • Legal & Compliance Q&A

    RAG over contracts, regulations and policy documents — with exact citation of the retrieved clause, audit trail of every query and response, and guardrails preventing legal advice output.

  • Financial Research Assistant

    RAG over financial reports, filings and market data — with structured data tool use for live price queries, multi-document synthesis and source attribution for every factual claim.

  • Technical Documentation Search

    Developer documentation RAG with code-aware chunking, API reference retrieval, code example extraction and hybrid search optimised for exact function/method name matching.

  • Multi-tenant RAG Platform

    RAG platform supporting multiple tenants with isolated vector indices, per-tenant document ingestion, access control at retrieval time and tenant-specific evaluation dashboards.

The RAG stack.

Ingestion, retrieval, reranking and evaluation each require specific tooling.

  • Ingestion

    Stack: LangChain (document loaders), LlamaIndex, Unstructured (PDF/Doc parsing), LlamaParse (layout-aware), Apache Tika

  • Embedding

    Stack: text-embedding-3-large, Cohere embed-v3, BGE-M3 (self-hosted), Voyage AI, Jina Embeddings

  • Vector Stores

    Stack: pgvector, Pinecone, Qdrant, Weaviate, Chroma (dev)

  • Search & Reranking

    Stack: BM25 (Elasticsearch/Typesense), Cohere Rerank, BGE Reranker, Reciprocal Rank Fusion

  • Orchestration

    Stack: LangChain, LlamaIndex, Haystack, Custom Python pipeline

  • Evaluation

    Stack: RAGAS, LangSmith, Arize Phoenix, PromptFoo, Custom eval suite

RAG system security extends beyond the LLM layer.

Multi-tenant retrieval and document access control are the primary security challenges.

  • Document Access Control

    In multi-tenant or multi-user RAG systems, retrieved chunks must only come from documents the querying user is authorised to access. Row-level security in the vector database, user-specific metadata filtering at query time and validation that retrieved chunks belong to authorised documents are required.

  • Prompt Injection via Documents

    Malicious content in indexed documents can contain prompt injection payloads — text designed to override system instructions when included in the LLM context. Input sanitisation of indexed content and output monitoring for instruction override patterns are required mitigations.

  • Data Privacy in Embeddings

    Embedding models convert text to vectors. When using third-party embedding APIs (OpenAI, Cohere), document content is sent to the provider. For sensitive content (PII, proprietary data), self-hosted embedding models (BGE-M3, Instructor) keep data within your infrastructure.

  • API Key Exposure

    Embedding API keys and vector database credentials must be server-side only. Client-side RAG (embedding queries in the browser) exposes API keys that will be extracted and abused.

  • Source Attribution Accuracy

    RAG systems that cite sources must accurately cite the specific chunk the answer was derived from — not fabricate citations. Verification that claimed sources actually contain the cited information is required for high-stakes applications.

  • Retrieval Poisoning

    In systems where users can submit documents for indexing, malicious users can inject documents designed to surface in retrieval for specific queries, biasing the RAG system's answers. Moderation of user-submitted content before indexing is required.

From documents to production RAG.

  1. 01. Data Audit

    Inventory source documents, assess format diversity, identify parsing challenges and define the question types the system must answer.

  2. 02. Test Set Construction

    Build 50-100 ground truth question-answer pairs from the corpus. This is the evaluation baseline for every subsequent decision.

  3. 03. Ingestion Pipeline

    Document loading, parsing, chunking and embedding. Evaluate chunking strategies against context precision on the test set.

  4. 04. Retrieval Optimisation

    Implement hybrid search, tune top-K, add reranking. Measure context precision improvement at each step.

  5. 05. Generation Layer

    Prompt engineering, context injection, source citation, guardrails and output validation.

  6. 06. Evaluation & Baselines

    RAGAS evaluation on test set. Set quality thresholds. Configure LangSmith for production monitoring.

  7. 07. Production Deployment

    API deployment, streaming responses, caching for repeated queries, rate limiting and quality monitoring dashboard.

Frequently Asked Questions

What is the difference between RAG and fine-tuning?

RAG retrieves relevant documents at query time and grounds the LLM response in verified content — the knowledge is external and instantly updatable. Fine-tuning trains the model weights on your data — the knowledge is internal and cannot be updated without retraining. RAG prevents hallucination by constraining responses to retrieved content. Fine-tuning does not — a fine-tuned model still hallucinates. For knowledge base Q&A, RAG is almost always the correct choice. Fine-tuning is appropriate for response style, format or domain-specific reasoning patterns that cannot be achieved through prompting.

What is context precision and why does it matter?

Context precision measures what percentage of the chunks retrieved for a query are actually relevant to that query. If you retrieve 5 chunks and only 2 are relevant, context precision is 40%. The LLM receives all 5 chunks as context — the 3 irrelevant ones add noise that increases hallucination risk and dilutes the relevant signal. Higher context precision means the LLM has less noise to ignore and produces more accurate, focused answers. Improving context precision (via better chunking, hybrid search, reranking) often improves answer quality more than upgrading the LLM.

How do you handle documents in multiple languages?

Multilingual RAG requires embedding models with strong cross-lingual performance. Cohere embed-v3 and BGE-M3 both perform well across 50+ languages. For hybrid search, language-specific BM25 tokenisation (using language-aware tokenisers) is important for non-English content. The LLM should handle multilingual context well — GPT-4o and Claude 3.5 both have strong multilingual capability. Test retrieval quality in each supported language with a representative question set.

What is HyDE and when should we use it?

HyDE (Hypothetical Document Embeddings) addresses the query-document embedding mismatch: a short question has a different embedding distribution than a long document passage. HyDE generates a hypothetical answer to the question using the LLM, then embeds the hypothetical answer instead of the original query. The hypothetical answer embedding is more similar to relevant document embeddings. HyDE helps when queries are very short or different in style from the indexed documents — technical Q&A systems often see significant retrieval improvement.

How many chunks should we retrieve (top-K)?

Start with top-5. Evaluate context precision — if it is below 70%, the LLM is receiving too much irrelevant context. Add reranking to improve precision before reducing K. With reranking, top-20 retrieve then rerank to top-3 often outperforms top-5 without reranking. Context window size also constrains K — at 8K tokens of context budget, you can fit roughly 8-16 chunks of 512 tokens. Match K to your context budget and retrieval precision target.