TL;DR
A RAG pipeline is the chain of stages that turns your documents into searchable chunks and then turns each question into a grounded answer. The offline half ingests, cleans, chunks, embeds and indexes your content. The online half rewrites the query, retrieves candidates, reranks them, builds a prompt and generates the answer, with caching in front. Build it one stage at a time and put a measurable check after each stage. Wrong answers usually start early, in parsing, chunking or a stale index, rather than in the model. Track recall, faithfulness, p95 latency and cost per successful answer, and rerun the same test set after every change.
Key Takeaways A RAG pipeline has an offline half that prepares content and an online half that answers each query. Errors travel downstream, so a better model cannot recover a passage that ingestion or retrieval lost. Every stage needs its own checkpoint, logged against one trace ID per query. Hybrid retrieval plus a reranker is the safer default for enterprise text full of codes and names. User-facing latency lives in the online stages, while most build cost sits in parsing and embedding. Cost per successful answer is a better tuning target than model price or average latency. Watch on YouTube
LLM Hallucination in Enterprise Data: The Semantic Layer Fix
Why a model gives confident wrong answers on company data, and how better context upstream of the model prevents them.
Where Did That Wrong Answer Come From? An HR assistant tells a manager that new parents get 12 weeks of paid leave. The policy moved to 16 weeks in March, and the revised PDF replaced the old one in the same SharePoint library. The answer even carries a citation, which makes it worse.
Work backward from the answer. The model quoted its context faithfully. The reranker put the old policy first because both versions read almost alike. Retrieval returned both because the index still held the superseded file.
Follow the trail one step further and the real cause appears. The ingestion job never captured an effective date, and deletes never reached the index. A problem born in stage one surfaced in stage ten, wearing a citation.
That is the case for treating a RAG pipeline as a set of stages you can test one at a time. Kanerika’s guide to retrieval-augmented generation covers what RAG is and how its architecture fits together. Here, each stage gets its failure signs, its checkpoint, and its latency and cost. That way, the next bad answer takes minutes to trace instead of a week of guessing.
What a RAG Pipeline Is and Where Each Stage Runs A RAG pipeline is the sequence of processing steps that lets a language model answer from your own content instead of from training memory alone. Like any data pipeline , it moves data through ordered steps, and each step can fail on its own. If you are still choosing between retrieval and training a model on your data, read RAG vs fine-tuning first.
This pipeline splits into two halves that run on different clocks. The offline half runs when documents change, preparing content once and refreshing it on a schedule or an event. Every question, on the other hand, triggers the online half, so every millisecond and token there is paid per query.
The Eleven Stages at a Glance Here are the eleven stages this guide walks through, in order.
Ingestion and parsing (offline). Pull files and records from source systems and extract clean text plus metadata.Cleaning and normalization (offline). Strip boilerplate, remove true duplicates and redact what should never be searchable.Chunking (offline). Split documents into retrieval units that still make sense on their own.Embeddings (offline and online). Convert chunks, and later each query, into vectors.Vector store and indexing (offline). Store vectors with metadata so search is fast and filterable.Query rewriting (online). Turn a conversational question into a searchable one.Retrieval (online). Find candidate chunks with dense, sparse or hybrid search.Reranking (online). Re-score candidates and keep the few that deserve prompt space.Prompt assembly (online). Pack instructions, evidence and the question into a token budget.Generation (online). Produce a cited answer, or decline when the evidence is thin.Caching (online). Reuse answers, embeddings and prompt prefixes when it is safe to do so.Evaluation runs between these stages as a set of checkpoints, so each section below ends with one. Turning the finished pipeline into a product is a separate job. Sign-in, a citation view and user feedback are covered in our guide to RAG applications .
Before You Build: A Test Set, Targets and a Trace ID Three things decide whether you can debug a RAG pipeline later. All three come before the first line of pipeline code.
Start with a test set of 50 to 200 real questions, each with a known answer and the passage that supports it. Include multi-document questions, exact lookups such as a contract number, and questions your content cannot answer. Those last ones test whether the pipeline knows when to say no.
Next, write down targets before anyone tunes anything. A workable starter set is Recall@10 for retrieval, a faithfulness score for answers, meaning how well each claim is supported by retrieved text, p95 latency for the online path, and cost per successful answer. The table below shows what each target tells you and where it gets measured.
Table 1: Starter targets for a RAG pipeline test set
Target What it tells you Where you measure it Recall@10 Whether the right passage reaches the candidate set After retrieval, on the test set Faithfulness Whether each claim in the answer is supported by retrieved text After generation Decline rate on unanswerable questions Whether the pipeline says no when it should After generation p95 latency How slow the slowest 5% of queries are Online path, per stage and end to end Cost per successful answer What a good answer costs once failures count Offline plus online spend, divided by passing answers
Cost per successful answer divides total spend, including amortized indexing, by the number of answers that passed your checks. It punishes a cheap pipeline that answers badly, which an average cost per query hides.
Give Every Query a Trace ID Give every query an ID that travels with it through each stage. Log the rewritten query, the retrieved chunk IDs with scores, the reranked list, the prompt size and the answer. Without that trail, the leave-policy hunt above turns into guesswork.
Stage 1: Ingestion and Parsing Ingestion pulls content from where it lives and turns it into text a machine can split and search. It is the least glamorous stage. Still, it sets the ceiling for every stage after it.
Extract Text Without Losing Tables and Structure PDFs cause most of the pain. A native PDF has a text layer, while a scanned one needs OCR. Either kind can scramble multi-column layouts or flatten a table into a single line of numbers.
So keep structure while you extract. Headings, list items and table rows should come out as headings, lists and rows, often as Markdown or HTML. Our guide to intelligent document processing covers the extraction side in more depth.
Kanerika ran into this, for instance, on a vendor-agreement project for a real estate developer. The contracts sat as unstructured PDFs in a content management system, and reliable extraction had to work before any question answering could.
Case Study
90% Faster Vendor Selection with LLM Agreement Processing
A real estate developer cut manual processing time by 82% once unstructured vendor agreement PDFs were reliably extracted and made searchable.
Read the Case Study → Carry Metadata Forward From the First Step Capture metadata at ingestion, because later stages cannot invent it. Each document needs a stable ID, source URL, version, last-modified time, effective date where relevant, owner, and the access groups allowed to read it.
Together, those fields drive the filters in retrieval, the citations in generation and the deletes in indexing. Access groups matter most, since permission checks belong inside retrieval rather than after it. Kanerika’s piece on data access governance explains the RBAC and ABAC models these fields usually map to.
What Breaks at Ingestion and How to Catch It Ingestion failures are quiet, unlike most software errors. A parser that drops page 14 or a connector that skips locked files raises no error. The gap shows up weeks later as a question nobody can answer.
The checkpoint, then, is a reconciliation report after every run. Compare documents found at the source with documents parsed, and pages expected with pages extracted. Flag pages with almost no text, since they usually mean a scan that skipped OCR. These are the same volume and freshness checks used in data observability .
Latency at this stage means throughput, counted in documents per hour. Cost is compute for parsing and OCR, plus connector API limits. Incremental loads that only touch changed files cut both, and our data ingestion guide compares batch and streaming patterns.
Stage 2: Cleaning and Normalization Clean text before you chunk it, or the noise gets copied into thousands of vectors. Repeated headers, footers, legal disclaimers and navigation menus are the usual offenders.
A disclaimer on every page, for example, matches many queries a little, which crowds out passages that match well. Strip it once, here, and keep a copy in metadata if compliance needs it.
Deduplication comes second, and it needs more care than it looks. Exact copies should go, but two versions of the same policy are different documents. Keep both, mark the older one as superseded, and let retrieval filter on the effective date.
Redaction comes third, and it is the one with legal weight. Personal data that no user of the assistant should see is safer removed before embedding than filtered later. In Kanerika’s RAG deployments, its PII redaction agent, Susan , handles this step before documents are ingested.
Our unstructured data governance guide covers classification for this kind of content. For the controls around the whole system, see our LLM security guide .
The checkpoint here is a sample review. Pull 30 random cleaned documents per run, compare them with the raw text, and confirm nothing meaningful was stripped. Compute cost is small here. The real risk is an over-eager rule that deletes useful content along with the noise. Kanerika’s AI data quality guide treats that problem in more detail.
On-Demand Webinar
LLM Security Risks: Financial Impact and How to Reduce Them
An on-demand session on the data exposure and access risks that appear when language models read enterprise content.
Watch the Webinar → Stage 3: Chunking for Retrieval, Not for Storage Chunking decides the unit of evidence your pipeline can retrieve. Make chunks too large and each one mixes three topics and matches none of them well. Make them too small, though, and the answer gets split across pieces that never show up together.
Pick a Starting Size, Then Measure It First, pick a documented default and treat it as a hypothesis. Microsoft’s Azure AI Search chunking guidance suggests 512 tokens with 25% overlap, about 128 tokens, as a starting point. The same page says the right overlap varies with content type.
Then measure answer-span retention against your test set. Specifically, for each question, check whether the supporting passage sits whole inside a single chunk. If that rate is low, change the boundaries before you touch the size.
Structure-aware splitting usually beats a fixed character count on enterprise documents. Split on headings and paragraph breaks first, and fall back to a token limit only when a section runs long. Keep tables whole, with their header row, even if that makes one chunk bigger than the rest.
Give Every Chunk Its Context Back A chunk that says “the period increases to 16 weeks” is useless without knowing which policy it came from. The usual fix, then, is to attach the document title and section path to each chunk before embedding. Azure’s guide suggests appending the title to chunks from the middle of long documents for this reason.
Anthropic’s contextual retrieval method takes the idea further. A model writes a short context line for each chunk before it is embedded. In its tests, that cut top-20 retrieval failures by 35%, and by 49% when paired with contextual BM25 keyword search. The trade-off is an extra model call per chunk at indexing time.
What Breaks at Chunking The common failures are split answers, orphaned table rows, and overlap so heavy that near-identical chunks fill the top results. Each one shows up later as a retrieval miss, so teams tend to blame the retriever.
After every run, check the token-length spread, empty chunks, duplicate chunks and the share of chunks that trace back to a source page. Chunk size also drives cost further down the line. Smaller chunks mean more vectors to store, while larger ones mean more prompt tokens per answer.
Stage 4: Embeddings Embedding turns each chunk into a vector, a list of numbers that places similar meanings near each other. The same model later embeds every incoming query, so the two sides have to match exactly.
One Model for Documents and Queries Vectors from different embedding models cannot be compared, even at the same dimension. Switching models means re-embedding the whole corpus, so record the model name and version on every vector.
Also watch the input limit. OpenAI’s embeddings guide lists an 8,192-token maximum input for text-embedding-3-small. Depending on the client, longer text is rejected or cut short, and a silent cut drops the end of a chunk without any error.
In addition, some models expect a different prefix or instruction for documents and for queries. Apply it the same way every time, or retrieval quality drifts for no visible reason.
What Embeddings Cost and When You Pay Again The first full embedding run is a one-time bill. After that comes a smaller recurring charge for changed documents and a tiny charge per query. Prices differ a lot by model, and OpenAI quotes roughly 62,500 pages per dollar for text-embedding-3-small against about 9,615 for text-embedding-3-large.
Dimension size also affects storage and search speed. The large model defaults to 3,072 dimensions against 1,536 for the small one, and the API can shorten vectors through a dimensions parameter. Test the smaller option on your own questions before paying for the bigger one.
Finally, the checkpoint is plain bookkeeping. Count chunks in, vectors out, failed batches and the model version used, then alert when any count drifts from the previous run.
Stage 5: Vector Store and Indexing The index is where the pipeline’s memory lives, and it is usually the last place anyone looks. Our vector database comparison covers product choices, so this section sticks to how the index behaves inside a pipeline.
Design the Index Schema Around Filters First, store more than the vector. Each record needs the chunk ID, parent document ID, the chunk text or a pointer to it, source version, effective date and access groups.
Then decide which fields are filterable before you load data. Filtering by department, region or effective date at query time is cheap. Discovering you need a new filter after loading millions of vectors means a rebuild.
Exact vs Approximate Search Small collections can use exact search, which compares the query with every vector. The pgvector documentation notes that exact search gives perfect recall, and that adding an index trades some recall for speed.
Beyond that, the same README compares two index types. HNSW gives a better speed-recall trade-off than IVFFlat, but it builds more slowly and uses more memory. Results can also change once an approximate index is added, so rerun your retrieval tests after any index change.
Filters also need extra care on approximate indexes. In pgvector, filtering happens after the index scan, so a narrow filter can return far fewer rows than you asked for. Teams on Databricks can see how similar choices play out in Databricks Vector Search .
Keep the Index in Step With the Source Index drift, in fact, is how the leave-policy answer happened. New files get added, but deletes and replacements often never arrive, so old chunks keep answering questions.
To prevent it, make updates idempotent by keying chunks on document ID plus version. When a document changes, delete its old chunks in the same job that writes the new ones. Then run a nightly check that compares document IDs at the source with document IDs in the index.
Index cost, in turn, grows with vector count, dimensions and replicas. Query latency depends on the index type, its search settings and how selective your filters are.
Stage 6: Query Rewriting People type the way they talk, so their questions rarely read like search queries. A follow-up like “what about contractors?” means nothing to a search index without the conversation before it.
Query rewriting therefore turns that into a standalone search query. The usual moves are folding in conversation history, expanding internal acronyms, fixing obvious typos and splitting a two-part question into two searches. More involved methods, such as searching with a drafted hypothetical answer, belong to our advanced RAG guide. Retrieval loops run by an agent are covered under agentic RAG instead.
The main failure mode here, however, is drift. A rewrite that adds a product name the user never mentioned sends retrieval somewhere else entirely. Log the original and rewritten query side by side, and test the pipeline with rewriting on and off.
An LLM-based rewrite adds a model call to every query, often the slowest step before retrieval. Rules or a small, fast model can handle acronyms and typos. Save the larger model for multi-turn conversations, where context really matters.
Stage 7: Retrieval With Dense, Sparse and Hybrid Search Retrieval finds the candidate chunks. Every later stage can only reorder or use what it returns, so a passage missed here is missed for good.
Why Hybrid Search Is the Safer Default Dense search compares embeddings and handles paraphrase well, matching “ending the contract early” to a termination clause. Sparse search, on the other hand, matches exact words, usually with BM25. It catches part numbers, error codes, clause IDs and names that embeddings tend to blur.
Enterprise questions need both. Hybrid search runs the two in parallel and merges the ranked lists. Azure AI Search, for example, uses Reciprocal Rank Fusion . It adds up 1/(rank + k) from each list a result appears in, with k set to a small value such as 60.
Apply access-control and freshness filters inside the retrieval call rather than on the results afterwards. Filtering afterwards can leave a user with three results out of ten, and it puts restricted titles into logs. Our AI access control guide covers the authorization model behind this.
Tune Top-K With Recall, Not Instinct Top-K is the number of candidates retrieval passes on. A larger K raises recall but feeds more text to the reranker, or straight into the prompt if there is no reranker.
Set K from data. Plot Recall@K for values from 5 to 100 on the test set and pick the point where the curve flattens. If recall stays low at every K, the problem sits upstream in parsing, chunking or embeddings, and a bigger K only hides it.
Stage 8: Reranking Retrieval is built for speed across millions of chunks, so its ranking is rough. A reranker then re-scores a short list with a slower, more accurate model. That model is usually a cross-encoder that reads the query and the chunk together.
The Sentence-Transformers documentation describes the classic pattern. A fast retriever returns around 100 candidates, and a cross-encoder re-ranks them. Scoring every pair in a large collection that way would be far too slow, which is why the two steps stay separate.
The gain can be large. In Anthropic’s contextual retrieval tests, contextual embeddings plus BM25 cut top-20 retrieval failures by 49%. Adding a reranker on top raised that cut to 67%, from a 5.7% failure rate to 1.9%. Your documents will behave differently, which is why you test.
Two failure modes belong to this stage. A reranker cannot add a passage that retrieval missed. A reranker fed truncated chunks scores the truncation rather than the content. Check both by comparing MRR and NDCG with reranking on and off.
Reranking latency grows with the pool size, so reranking 100 candidates costs far more than reranking 30. Find the smallest pool that keeps the quality gain.
Stage 9: Prompt Assembly Prompt assembly packs the system instructions, the selected chunks, the question and room for the answer into one token budget. Small choices here change answers more than most teams expect.
Order also matters. The Lost in the Middle study found that models use information best when it sits near the start or end of a long context. Accuracy drops for material buried in the middle, so put the strongest evidence first and keep the context short.
Kanerika Service
RAG Development Services
Kanerika designs and builds RAG pipelines with stage-level checkpoints, hybrid retrieval and access controls that fit your enterprise data.
Explore RAG Development Label every chunk with a source ID so the model can cite it. Then tell the model what to do when the evidence does not answer the question. A short template like the one below makes those rules explicit and testable.
SYSTEM
Answer using only the sources below.
Cite each claim with its source ID in square brackets, for example [S2].
If the sources do not contain the answer, reply exactly:
"I could not find this in the approved documents."
SOURCES
[S1] Leave Policy v4 | effective 2026-03-01 | section 2.1
<chunk text>
[S2] Benefits FAQ | updated 2026-04-12
<chunk text>
QUESTION
<rewritten user question>After that, budget tokens on purpose. Reserve a fixed slice for instructions, cap evidence at a set number of chunks, and hold back enough output tokens for a complete answer. Kanerika’s guide to context engineering vs prompt engineering goes further into how teams manage that budget, and our prompt engineering best practices cover template testing.
Finally, test the assembled prompt itself. In particular, confirm the expected chunks made it in, the token count stays under the cap, and no chunk was cut mid-sentence by a truncation step.
Stage 10: Generation and Grounding Checks Generation is where the answer appears, so it collects the blame for every earlier stage. Treat it as the last check in a chain, with a short list of failures that belong to it alone.
Typically, those failures are unsupported claims, missing or wrong citations, ignoring the best passage, and blending two conflicting sources into one confident answer. Our explainer on LLM hallucination covers why models fill gaps this way.
Case Study
80% Ticket Auto-Resolution with LLM-Driven AI
A B2B SaaS company built a knowledge base and prepared past tickets for an LLM, then answered 80% of tickets automatically and halved resolution time.
Read the Case Study → Then check grounding once the answer exists. A lightweight step can confirm that each cited source ID exists in the prompt and that the cited passage supports the sentence. When support is weak, return the decline message instead of the answer.
Track time to first token separately from total response time. Users feel the first, and budgets feel the second. Streaming helps perceived speed, while output length drives cost, so cap it for question types that need short answers.
Where data cannot leave your environment, the generation model can run as one of the private LLMs , and the grounding check works the same way. In production, faithfulness and decline rate become live signals, and our roundup of AI observability tools compares ways to watch them.
Stage 11: Caching Without Serving Stale Answers Caching saves money and time on repeat work. It can also freeze a wrong answer in place. A typical RAG pipeline has four caches worth considering.
Response cache. Stores final answers for exact repeat queries.Semantic cache. Returns a stored answer when a new query embeds close enough to an earlier one.Embedding cache. Skips re-embedding chunks whose text has not changed, keyed on a content hash.Prompt cache. Lets the model provider reuse a long, unchanged prompt prefix such as system instructions.Provider prompt caching is usually the easiest win, since it needs almost no design work. Anthropic’s prompt caching documentation, for instance, bills cache reads at 0.1 times the base input price on most models. The default cache lifetime is five minutes, refreshed each time the prefix is reused.
The stale-answer risk sits in the response and semantic caches. Key every cached answer to the index version, and clear entries whenever a document they cite changes. Set the semantic similarity threshold high, because “contractors” and “employees” can differ by one word and still need opposite answers. An LLM gateway is a common place to enforce these caching rules across teams.
RAG Pipeline Latency and Cost, Stage by Stage Latency and cost behave differently on the two halves of the pipeline. Offline stages are measured in throughput and total spend per refresh, whereas online stages are measured per query. Their delays also add up in sequence.
The table below names what drives each figure rather than benchmarks, because your models, index and hardware set the actual numbers.
A Stage-by-Stage Reference for Debugging Keep this table open while debugging. Read across a row when one stage is suspect, and down the last two columns when p95 or the bill moves.
Table 2: RAG pipeline stages compared by failure, checkpoint, latency and cost
Stage Runs Common failure Checkpoint Latency driver Cost driver Ingestion and parsing Offline Missing pages, flattened tables, no version metadata Source vs parsed document and page counts Documents per hour, OCR Parsing and OCR compute, connector limits Cleaning Offline Boilerplate kept, or real content stripped Sample diff of raw vs cleaned text Minor Minor compute Chunking Offline Answers split across chunks, orphaned table rows Answer-span retention, token-length spread Minor Chunk count drives vectors and prompt tokens Embeddings Offline and online Silent truncation, mixed model versions Chunks in vs vectors out, model version Batch throughput, one query call Tokens embedded, re-embedding on model change Indexing Offline Stale or duplicate entries, broken filters Source vs index ID reconciliation Index build time, query search time Storage, memory, replicas Query rewriting Online Rewrite changes the question’s meaning Original vs rewritten query, on/off test One extra model call Rewrite model tokens Retrieval Online Right chunk never in the candidate set Recall@K, Hit Rate@K Search plus filters Search compute Reranking Online Good chunk pushed down, truncated inputs MRR and NDCG with reranking on and off Grows with pool size Reranker calls per query Prompt assembly Online Evidence buried mid-context or cut off Expected chunks present, token count under cap Negligible Sets input token volume Generation Online Unsupported claims, wrong citations Faithfulness, citation accuracy, decline rate Time to first token, output length Input and output tokens Caching Online Stale answers served after a source changes Hit rate, invalidations per index version Saves time on hits Cache storage, saves model spend
Two Habits That Keep the Numbers Honest First, log latency per stage under the trace ID, not only end to end, so a slow reranker cannot hide inside a fast total. Report the p95 rather than the average, because the slow tail is what users remember. Second, judge every change by cost per successful answer, as defined earlier, rather than by cost per query.
Cutting cost usually starts with the cheapest levers. Shrink the rerank pool, trim the number of chunks in the prompt, cap answer length and turn on prompt caching. Only then consider a smaller generation model.
Evaluation Checkpoints Between Stages An end-to-end score tells you something is wrong. Stage-level checkpoints, however, tell you where. Kanerika’s LLM evaluation framework guide covers the metrics in depth, so this section focuses on where each check belongs.
Each stage section above ends with its own check, and the Checkpoint column in Table 2 collects them in one place. What matters here is cadence. Run the cheap offline checks on every refresh: reconciliation, answer-span retention, and vector and ID counts. Run the query-level checks on the full test set after every change: Recall@K and MRR after retrieval and reranking, then faithfulness, citation accuracy and decline rate after generation.
In practice, open-source tools can automate much of this. Ragas , for example, ships metrics such as context precision, context recall and faithfulness. Automated judges still need a human-reviewed test set, or they will score confidently against the wrong answers.
So run the full suite after every change to chunk size, embedding model, top-K, reranker or prompt template. A change that lifts one metric often lowers another, and the regression run is where you see it before users do.
Checklist
Generative AI Checklist: Secure Adoption and Governance
Work through data readiness, security, evaluation and governance checks before a RAG assistant moves past its pilot.
Get the Checklist → An Example RAG Pipeline Definition A pipeline definition puts every stage choice in one versioned file. That way, changes become reviewable, each regression score ties to a configuration, and the trace ID has a version to point at.
The YAML below is illustrative and not tied to any one framework. Swap the component names for whatever your stack uses, and see our RAG tools roundup if you are still picking those components.
pipeline: hr-policy-assistant
version: 2026.10.3
trace: {id_header: x-trace-id, log_stages: all}
offline:
ingest:
sources: [sharepoint://hr-policies, confluence://people-ops]
parser: layout_aware_pdf
ocr: on_low_text_pages
metadata: [doc_id, version, effective_date, owner, acl_groups]
mode: incremental # deletes and supersedes propagate
clean:
strip: [headers, footers, disclaimers]
dedupe: exact_only # keep versions, mark superseded
redact: [national_id, bank_account]
chunk:
strategy: by_heading
max_tokens: 512
overlap_tokens: 128
prepend: [doc_title, section_path]
embed:
model: text-embedding-3-small
dimensions: 1536
batch_size: 256
index:
store: pgvector
index_type: hnsw
filterable: [acl_groups, effective_date, department]
reconcile: nightly
online:
rewrite: {model: small-fast, history_turns: 3}
retrieve:
mode: hybrid
fusion: rrf
rrf_k: 60
top_k: 40
filters: [acl_groups, effective_date <= today]
rerank: {model: cross-encoder, pool: 40, keep: 6}
assemble:
max_context_tokens: 3000
order: best_first
cite_format: "[S{n}]"
generate:
max_output_tokens: 400
on_low_support: decline
cache:
response: {key: [query_hash, index_version], ttl: 24h}
semantic: {threshold: 0.97, key: index_version}
prompt_prefix: provider
checkpoints: # placeholders: set your own targets
recall_at_10: ">= 0.85"
faithfulness: ">= 0.90"
decline_rate_unanswerable: ">= 0.90"
p95_latency_ms: "<= 4000"Three lines in that file would have stopped the leave-policy answer. Incremental mode carries the delete through to the index, and the effective-date filter keeps any surviving old version out of the results. A response cache keyed to the index version then makes sure users get the fix, not a cached copy of the old answer.
The thresholds under checkpoints, however, are placeholders. Set them from your own baseline run, then tighten them as the pipeline matures.
How to Trace a Bad Answer Back Through the Pipeline When a wrong answer is reported, work backward from the output, because each step rules out one stage cheaply. Pull the trace ID first, then check the stages in this order.
Cache. Was the answer served from the response or semantic cache, and under which index version? If so, clear the entry before checking anything else.Generation. Do the cited sources say what the answer claims? If they do, the model behaved and the problem is earlier.Prompt assembly. Was the correct passage in the prompt, and where did it sit?Reranking. Was the correct chunk in the candidate pool but dropped from the final set?Retrieval. Was the correct chunk retrieved at all? Test dense and sparse search separately and check the filters.Query rewriting. Did the rewrite change what the user asked?Index. Does the correct chunk exist, and does an outdated one sit beside it?Chunking and cleaning. Is the answer split across chunks, or stripped as boilerplate?Ingestion. Did the source document parse fully, with the right version and metadata?Once you find the cause, add the question to the test set so the same failure cannot return unnoticed. Over a few months, that habit builds the most useful regression suite your team will own.
Tracing by hand works for a pilot. In production, though, you want spans, stage-level metrics and alerts, which Kanerika’s guide to LLM observability explains step by step.
Common RAG Pipeline Mistakes These turn up in almost every first build. Each one stays hidden until real users and real documents arrive.
Indexing files before checking extraction quality. Picking a chunk size from a blog post instead of your own test set. Updating documents without deleting their old chunks. Raising top-K to cover for weak retrieval. Adding a reranker without measuring its gain and its latency. Treating a fluent, cited answer as proof that the evidence was right. Reporting average latency while ignoring p95 and cost per successful answer. How Kanerika Builds RAG Pipelines That Survive Production Kanerika’s RAG development services team builds pipelines stage by stage, with retrieval and answer quality measured from the start. As a rule, the work runs in five steps.
Assess the content. Profile formats, scan quality, versioning and access rules before choosing any component.Build the test set with the business. Collect real questions, expected answers and supporting passages, including questions the assistant should decline.Build the offline half. Tune parsing, cleaning, chunking and indexing against answer-span retention and reconciliation reports.Build the online half. Tune hybrid retrieval, reranking, prompt assembly and caching against recall, faithfulness, p95 latency and cost per successful answer.Operate and govern. Keep trace IDs, nightly reconciliation and regression runs going after launch.The pattern also shows up in Kanerika’s delivery work. For a B2B SaaS company serving customers in more than 40 countries, Kanerika first built a knowledge base and prepared historical support tickets. Then an LLM-based resolution system went on top. The client saw 80% of tickets answered automatically, a 50% drop in resolution time and 70% lower staffing cost.
For a Middle East real estate developer, by contrast, the hard part sat at ingestion. Vendor agreements lived as unstructured PDFs, so the team built reliable extraction before adding a chat interface for vendor search. Manual processing time fell by 82%, and vendor selection became 90% faster.
Lessons That Repeat Across Projects The same lessons repeat across these projects. Version metadata decides whether answers stay current, and access groups have to travel with every chunk. The first regression suite should exist before the first demo, since demos are where unmeasured changes creep in. Teams that also need a custom model alongside retrieval can pair this with Kanerika’s LLM development services .
Talk to Kanerika
Get a RAG Pipeline Review
Walk through your pipeline with Kanerika’s engineers and find which stage is costing you accuracy, latency or money.
Book a Meeting → Wrapping Up A RAG pipeline is easier to build, and far easier to debug, when you treat it as eleven stages with a check after each. Ingest and clean with care, chunk for retrieval, and keep the index in step with the source. Use hybrid retrieval with a reranker, assemble tight prompts, ground every answer and tie caches to the index version. Then measure recall, faithfulness, p95 latency and cost per successful answer on the same test set after every change. The next time an answer comes back wrong, the trace will show you which stage to fix.
Frequently Asked Questions
What is a RAG pipeline? A RAG pipeline is the chain of processing stages that lets a language model answer from your own content. An offline half ingests, cleans, chunks, embeds and indexes documents. An online half rewrites each question, retrieves and reranks passages, builds a prompt and generates a cited answer. Each stage can fail separately, so each needs its own check.
What are the stages of a RAG pipeline? A typical RAG pipeline has eleven stages. The offline stages are ingestion and parsing, cleaning, chunking, embedding and indexing. The online stages are query rewriting, retrieval, reranking, prompt assembly, generation and caching. Evaluation runs as checkpoints between those stages rather than as one final step, so each failure can be traced to its source.
How do you build a RAG pipeline step by step? Start with a test set of real questions and expected answers, plus targets for recall, faithfulness, latency and cost. Then build the offline half and check extraction, chunk quality and index counts. Add hybrid retrieval, a reranker, a prompt template and grounded generation. Finally add caching, trace IDs and a regression run after every change.
What is the best chunk size for a RAG pipeline? No single size works for every collection. Microsoft’s Azure AI Search guidance suggests starting near 512 tokens with about 25 percent overlap, then testing. Measure how often the passage that answers each test question sits whole inside one chunk. Split on headings and paragraphs first, and keep tables intact with their header rows.
Do you need a vector database for a RAG pipeline? You need somewhere to store vectors with their metadata and search them quickly. That can be a dedicated vector database, a search engine with vector support, or a relational database with an extension such as pgvector. Small collections can even use exact search. Choose based on scale, filtering needs and what your team already operates.
What is hybrid search in a RAG pipeline? Hybrid search runs dense vector search and sparse keyword search such as BM25 in parallel, then merges the two ranked lists. Dense search handles paraphrased questions, while keyword search catches exact codes, IDs and names. A common merge method is Reciprocal Rank Fusion, which rewards passages that rank well in both lists.
Why add a reranker to a RAG pipeline? First-stage retrieval is tuned for speed, so its ranking is rough. A reranker, usually a cross-encoder, reads the question and each candidate together and scores them more accurately. That puts the best evidence at the top of a small prompt. It cannot recover passages retrieval missed, so measure recall before adding one.
Where do RAG pipelines most often fail? Most failures start early. Parsers drop pages or flatten tables, chunking splits answers, and indexes keep outdated versions after a document changes. Those problems surface later as wrong or unsupported answers, so the model gets blamed. Logging each stage’s output under one trace ID lets you work backward from a bad answer to the stage that caused it.
How do you reduce RAG pipeline latency? Measure p95 latency per online stage before changing anything. Common fixes are a smaller or rule-based query rewriter, a smaller rerank pool, fewer chunks in the prompt and shorter answers. Streaming improves how fast the answer feels. Caching repeat answers and long prompt prefixes removes work entirely when it is safe to reuse results.
How do you calculate the cost of a RAG pipeline? Add offline costs such as parsing, embedding and index storage, spread across the queries they serve. Then add per-query costs for rewriting, retrieval, reranking and generation tokens. Divide the total by answers that passed your quality checks. That cost per successful answer shows whether a cheaper setting is really cheaper once bad answers count.
Where should evaluation checks sit in a RAG pipeline? Evaluate each stage with its own check rather than one overall score. Reconcile document counts after ingestion, measure answer-span retention after chunking, and track Recall@K after retrieval. After generation, score faithfulness, citation accuracy and how often the system declines unanswerable questions. Rerun the same human-reviewed test set after every pipeline change.
How often should a RAG pipeline re-index documents? Re-index on change rather than on a fixed calendar when your sources allow it. Incremental jobs should add new versions and delete superseded chunks in the same run. A nightly reconciliation between source and index catches anything missed. Fast-moving content such as policies or tickets may need near real-time updates, while archives can refresh weekly.