TL;DR
A RAG application is a complete software product built around a retrieval-augmented generation pipeline, so people get answers grounded in company data. The pipeline finds the right content and writes the answer. The application decides who may ask, what each person may see, and how answers and sources appear. Common types are knowledge assistants, support copilots, document review tools, enterprise search and assistants built into business software. Every production version needs sign-in, document-level permissions, checkable citations, guardrails, feedback and monitoring, added in staged releases with pass tests. Buy a packaged assistant for general questions over common content, and build when the workflow, rules or permissions are your own.
Key Takeaways A RAG application wraps the retrieval pipeline in a product, so it adds an interface, identity, permissions, citations and feedback. Five application types cover most enterprise demand, but each one has its own must-have feature and its own costly failure. Permissions must be enforced at retrieval time, per document, before any text reaches the model. A prototype becomes production in stages, and each stage closes only when quality, security, cost and latency tests pass. Evaluate the app on four layers, from answer quality and permission tests to user outcomes and cost per answer. Most enterprises land on a hybrid, managed retrieval underneath and their own application layer on top. Watch on YouTube
AI Agent Access Control: From Excessive Agency to Least Privilege
Kanerika explains how to scope what AI systems can reach, from least-privilege permissions to stopping assistants that see more than their users should.
The Answer the Bot Should Never Have Given A sales analyst types a quick question into the company’s new internal assistant. What did the leadership team receive in bonuses last year? Four seconds later the answer arrives, neatly summarized and cited to an HR compensation file the analyst has never been allowed to open.
Nothing in that exchange looks broken. Retrieval found the most relevant passage, the model stayed faithful to it, and the citation was accurate. Still, nobody checked whether this person could read that file before its text reached the model.
OWASP lists sensitive information disclosure as LLM02 in its 2025 Top 10 for LLM applications, and the bonus answer is a textbook case. The pipeline did its job. The application around it, though, had not been built yet.
What Is a RAG Application, and Where Does the Product Begin? A RAG application is software that answers plain-language questions from an organization’s own documents and data, with sources attached. Underneath sits a retrieval-augmented generation pipeline, the technique Patrick Lewis and colleagues named in a 2020 research paper . If you want the concept and the reference architecture first, our guide to retrieval-augmented generation covers both.
This article starts where that guide stops. A pipeline is an engine, while an application is the car around it, with seats, locks, a dashboard and someone responsible for servicing it.
RAG Pipeline vs RAG Application The pipeline ingests content, splits it into chunks, embeds it and retrieves the best matches for a question. Then it asks a model to answer from those matches. Each stage has its own failure modes and tuning, so our step-by-step RAG pipeline guide walks through them one at a time.
The application, by contrast, is everything a user, a security reviewer and a support team touch. The quickest way to see the line is to ask who owns each question once the system is live.
Table 1: What the pipeline answers vs what the application answers
Question Pipeline layer Application layer Which passages match this question? Yes, retrieval and ranking No Is this user allowed to see those passages? No Yes, filtered before ranking Is the answer supported by the sources? Partly, through prompting Yes, citation checks and evaluation What happens when the answer is wrong? No Yes, feedback, escalation and fixes Who sees usage, cost and incidents? No Yes, logging, audit and monitoring Which topics are out of bounds? No Yes, guardrails and business rules
Why a Working Demo Is Not a Finished Product Demos run on a tidy folder of fifty documents, one friendly user and questions the builder already knows. Production, on the other hand, brings thousands of changing files and people with different access rights. It also brings a compliance team that wants to know who saw what.
That shift is why a team can be proud of a prototype on Friday and blocked by security on Monday. Retrieval quality was rarely the problem. Instead, the missing pieces were the ones a product manager would have listed in week one.
Gaps like these also explain why so many promising pilots stall. Our AI proof of concept framework shows how to design a pilot so the path to production is visible from the start.
Five Types of RAG Applications, With an Example of Each Five types cover the bulk of enterprise demand, and they differ more in who uses them than in how retrieval works. The same pipeline can sit under all five. What changes, then, is the user, the sources, the feature that cannot be skipped and the mistake that hurts most.
1. Internal Knowledge Assistant An internal knowledge assistant answers employee questions about policies, processes and internal guidance, such as HR rules, IT how-tos or brand standards. Its sources are usually SharePoint, Confluence or a wiki. Because an HR folder and an engineering wiki rarely share an audience, the must-have feature is permission-aware answers.
For example, Kanerika built an assistant of this type for a payment technology provider whose teams kept waiting on brand experts for routine logo questions. A conversational assistant over the approved guidelines cut approval turnaround by 60% and manual effort per query by 70%. Our piece on AI knowledge management also covers the content side of this problem.
2. Customer Support Copilot and Self-Service Support applications answer customer or member questions from manuals, help articles and resolved tickets. Some face the customer directly, while others sit beside a human agent and draft the reply. Either way, the must-have features are citations to approved content and a clean handoff to a person when the system is unsure.
Getting this wrong in public costs money. In February 2024, British Columbia’s Civil Resolution Tribunal held Air Canada liable for wrong bereavement-fare advice from its website chatbot. The tribunal’s Moffatt decision says the airline is responsible for all the information on its website, chatbot included.
Getting it right shows up in support numbers. For instance, a Kanerika support agent for an expert network resolved 65% of member queries on its own, routing low-confidence cases to live staff. As a result, ticket volume fell 42% and cost per ticket fell 31%. For the wider operating model, see our guide to AI for customer service .
Case Study
65% Self-Service Resolution with an AI Support Agent
An AI agent connected to the knowledge base and Zendesk resolved routine member questions, cut ticket volume by 42% and lowered cost per ticket by 31%.
Read the Case Study → 3. Document Q&A and Review Workspace Document applications help specialists question long files, such as contracts, RFPs, regulatory filings or due-diligence packs. For example, a legal reviewer might ask which clauses depart from the standard playbook, then jump to each passage. The must-have feature is source-linked answers that show the exact text, so the expert verifies rather than trusts.
An investment bank Kanerika worked with needed answers from long RFI documents and databases without reading each file, under strict access rules. Its chat-based retrieval layer enforced role-based access at the point of retrieval. Retrieval became 43% faster while role-based compliance held at 100%.
Pairing this type with intelligent document processing also helps when the source files start life as scans.
4. Enterprise Search With Generated Answers Enterprise search applications replace a list of blue links with a direct answer plus the ranked sources behind it. Because they span many repositories, the hardest part is syncing permissions across systems that each manage access differently. The must-have feature is that the answer and every source respect the searcher’s rights in the original system.
Packaged tools already work this way. Microsoft states that Microsoft 365 Copilot only surfaces data a user has at least view permission for . So the real risk becomes oversharing in SharePoint, a pattern our article on Microsoft Copilot security concerns unpacks.
5. Analyst Assistant Embedded in Business Software Embedded assistants live inside a tool people already use, such as a CRM record, a service ticket or an operations dashboard. They answer questions about the record on screen and use its context to narrow retrieval. The must-have feature is workflow context, so the answer is about this customer or this ticket rather than a generic one.
Kanerika’s closest example starts from a record rather than a typed question. For an expert network, an AI agent reads each survey request and then finds specialists through semantic search across skills and domains. Next, it checks them against past participation and shows each match with context and source links in one dashboard.
Overall, mapping accuracy rose 40% and mismatch tickets fell 80%.
How to Choose the Right RAG Application Type Start from the user and the decision they make, not from the documents you happen to have. Then check how sensitive the sources are and what happens after someone reads the answer. A wrong answer read by an employee is a nuisance, but a wrong answer sent to a customer is a liability.
Table 2: Requirements by RAG application type
Application type Primary user Must-have feature Main risk Success metric Internal knowledge assistant Employees Permission-aware answers Restricted content exposed Time saved per question Support copilot or self-service Customers and agents Citations and human handoff Wrong guidance given as policy Verified resolution rate Document Q&A and review Domain specialists Passage-level source links Unsupported conclusions Review time per document Enterprise search with answers Everyone Cross-system permission sync Stale or overshared access Search success rate Embedded analyst assistant Operational teams Record and workflow context Answer about the wrong record Task completion time
Industry examples, such as banking or retail, sit in our roundup of generative AI use cases . The table above, though, is about something narrower, which is what each product must include before real users touch it.
The Components Around the Pipeline That Make a RAG Application a Product Once a pipeline works, seven components turn it into a product. None of them improves retrieval on its own, yet each one decides whether people trust the system and keep using it. For where they sit in the wider request flow, see our generative AI architecture guide.
Interface and API Layer The interface is where people ask and read. A chat window is common, but a search box, a side panel or a Teams bot often fits the work better. Good interfaces show progress while the answer streams, keep conversation history and allow follow-up questions.
Behind the interface sits an API that other systems call. Treat it like any production service, with versioning, rate limits, timeouts and a clear message when retrieval returns nothing. That way, teams that serve several front ends avoid building the same permission logic three times.
Identity and Document-Level Permissions Every request must carry a verified identity from single sign-on, never a name typed into the prompt. The application then looks up that person’s current groups and passes them to retrieval as a filter. As a result, documents the user cannot open in the source system drop out before ranking, so the model never sees them.
This is the control that would have stopped the bonus question. OWASP’s entry on vector and embedding weaknesses (LLM08) recommends permission-aware vector stores and strict partitioning between user groups. Managed search services also support the pattern, and Microsoft documents how Azure AI Search trims results by user or group at query time .
# Permission-aware retrieval: filter BEFORE the model sees any text
user = auth.verify(request.token) # identity from SSO, not from the prompt
groups = directory.current_groups(user.id) # read live, so revoked access drops out
hits = search_index.query(
text=request.question,
top_k=8,
filter={"allowed_groups": {"any_of": groups}}, # trimmed at query time
)
if not hits:
return reply("I could not find that in sources you can access.")
answer = llm.generate(question=request.question, context=hits)
return reply(answer, citations=[h.source_url for h in hits])Two details catch teams out. First, access lists change daily, so permission metadata in the index must sync as fast as access is revoked. Second, filtering after the model writes its answer is too late, because the restricted text has already shaped the reply.
The policy side matters just as much as the code. Our guides to AI access control and data access governance cover how to define who may see what before you encode it.
Citations Users Can Check A citation is only useful when it opens the right passage in a version the user can read. So show the document title, the section and a link that lands on the quoted text rather than the whole file. For long documents, marking the supporting sentence also saves the reader a search.
Some citations are decorative, since they point at a real source that does not support the claim. Production apps therefore check that each cited passage carries the substance of its sentence. When support is weak, the safer move is to say so, which also limits the hallucination on enterprise data that erodes trust fastest.
Application Guardrails Guardrails are the business rules wrapped around each request and response. They keep the assistant on approved topics and redact personal data in answers. They also decline legal or medical advice when that is out of scope, and they refuse politely when sources are thin.
Indirect prompt injection needs its own defense, because hostile instructions can hide inside a retrieved document. Many teams therefore centralize these checks in an LLM gateway , so every application inherits the same policy. Our LLM security guide lists the attacks to test, while the wider controls sit in our generative AI security guide.
Feedback Loop and Human Escalation Every answer should carry a thumbs up, a thumbs down and a short reason field. That signal only matters when someone reviews it weekly and turns it into fixes. Typical fixes include a missing document, a stale page, a retrieval miss or a new guardrail.
Escalation is the other half. When confidence is low, or the user asks for a person, the conversation should move to a human. The question, the answer and the sources should travel with it, so nobody starts over.
On-Demand Webinar
LLM Security Risks: Financial Impact and How to Reduce Them
An on-demand Kanerika session on the security risks LLM applications carry, what they cost when they go wrong, and the controls that reduce them.
Watch the Webinar → Logging, Audit, and Observability Log each request with the user, the documents retrieved and the answer, as well as the latency and token cost. Security teams need that trail to answer who saw what, while engineers need it to debug a bad answer days later. Keep personal data out of logs, or mask it, so the audit trail does not become a new leak.
Traces, spans and alerts for model calls are a discipline of their own. Our sibling guide on LLM observability shows how to instrument them step by step.
Content Ownership and Freshness An assistant is only as current as its sources. That is why each source collection needs a named owner who approves new content, retires outdated pages and handles deletion requests. Show the document date beside each citation, so users can spot an answer built on last year’s policy.
This is governance work more than engineering. Our article on unstructured data governance explains how to keep document estates clean enough to feed an assistant.
From Prototype to Production: A Six-Stage Release Plan for a RAG Application Getting to production takes six stages, and each ends with a test that must pass before the next begins. The stages focus on the product and the people around it. Meanwhile, ingestion and retrieval tuning continue inside the pipeline workstream.
Stage 1: Scope the Product and Its Success Measure First, pick one repeated task with identifiable users, approved sources and a measurable cost today. Then write down the questions it must answer, the ones it must refuse and the business number it should move. The exit test is a one-page scope that the business owner, security and the build team all sign.
Stage 2: Prototype Against Real User Tasks Build a thin end-to-end version with real documents and 50 to 100 questions collected from future users. Include questions that should fail, too. The exit test is that target users can finish their actual tasks with it, judged by them rather than by the builders.
Stage 3: Wire In Identity, Permissions, and Feedback Connect single sign-on, permission filtering at retrieval, feedback capture, error handling and audit logging. This comes before wider testing because permissions change which documents are retrieved, and so they change answer quality. The exit test is a permission test suite that passes for every role you intend to serve.
Watch on YouTube
From AI Pilot to Production | How to Scale AI Successfully
What it takes to move an AI pilot into production, from scoping and ownership to the checks that keep a working demo from stalling.
Stage 4: Test Quality, Security, Cost, and Latency Run the full question set, the refusal set, prompt injection attempts and a load test at expected peak usage. Set targets before testing for grounded answers, p95 response time and cost per answered question. The exit test is meeting those targets, with any gaps signed off as accepted risks.
Stage 5: Launch to a Controlled Group Release to one team or region with training, a support channel and a named product owner. Then watch adoption, thumbs-down reasons and escalations daily for the first few weeks. The exit test is stable quality and steady usage rather than a launch-day spike.
Stage 6: Run It as a Product Schedule regular reviews of failures, content changes, costs and feedback, then ship fixes on a release cadence. Expand to new users or sources only when the numbers from the controlled group hold. In other words, treat the assistant as a product with a roadmap, not a project with an end date.
Table 3: Prototype vs production checklist for a RAG application
Requirement Typical prototype Production pass condition Identity Shared test login SSO on every request, so no anonymous access Permissions Everyone sees everything Retrieval filtered by current groups, then tested per role Citations File names listed Links open the supporting passage, and support is checked Quality Builder’s spot checks Fixed question set meets the agreed target before release Refusals Not tested Declines out-of-scope questions when sources are thin Feedback None Ratings captured, then reviewed weekly by an owner Operations Runs on a laptop p95 latency, uptime and cost per answer monitored Ownership The builder Named product, content and support owners
Teams that skip the checklist still rediscover it later, one incident at a time.
How to Evaluate a RAG Application Once Real Users Arrive Evaluation at this stage means testing the product, not only the pipeline. Retrieval and generation scores tell you whether the engine works. However, they do not tell you whether the right people got safe, useful answers at an acceptable cost.
Answer Quality and Citation Accuracy Keep a fixed test set of real questions with expected sources, and rerun it on every change. The open-source Ragas library defines faithfulness as the share of claims in a response that the retrieved context supports. The app’s test set, however, also carries the refusal questions from stage two, so a release cannot score higher by answering less.
Then add a citation check on top. Sample answers each week and confirm that every cited passage supports the sentence it is attached to. Our LLM evaluation framework guide compares the tooling for this.
Permission Tests First, create test users for each role. Then, as each of them, ask questions whose answers live in documents they must not see. A pass is a refusal or an answer drawn only from permitted sources, and the suite reruns whenever the permission model, the index or the source systems change.
User Outcomes Track weekly active users, repeat usage, thumbs-down rate, escalation rate and questions answered without a follow-up ticket. Also tie at least one number to the business measure chosen in stage one. Usage that drops after week three usually means answers stopped being useful, even when quality scores look fine.
Latency, Availability, and Cost per Answer Watch p50 and p95 response time, error rate and uptime, just as you would for any customer-facing service. Then divide total model, search and hosting spend by questions answered successfully to get cost per answer. That figure makes return on investment talks concrete, as our article on generative AI ROI explains.
Table 4: Four evaluation layers for a RAG application
Layer Example metric Typical owner Cadence Answer quality Faithfulness and citation accuracy AI engineering Every release, as well as weekly samples Permissions Restricted-content leak rate (target zero) Security Whenever the index or access changes User outcomes Resolution rate and thumbs-down rate Product owner Weekly Operations p95 latency and cost per answer Platform team Continuous, with alerts when targets slip
Build vs Buy a RAG Application: How to Decide The build vs buy choice comes down to how much of the application layer you need to control. Retrieval engines are increasingly a commodity. Permission logic, workflow fit and business rules, however, are not.
When Buying Makes Sense Buy when the need is general question answering over content in one suite, and when the access model already lives there. A packaged assistant such as Microsoft 365 Copilot brings administration, support and permission handling on day one. The trade-off is limited control over prompts, refusals, citations and evaluation.
When Building Makes Sense Build when the workflow is specific or the sources span several systems with different access models. The same applies when answers must follow domain rules that a generic product cannot encode. You then own the quality bar and the roadmap, but you also own the maintenance.
The Hybrid Most Enterprises Land On Most enterprise teams settle on a hybrid. Managed retrieval services like Azure AI Search or Amazon Bedrock Knowledge Bases handle indexing and search. Meanwhile, the team builds the interface, permission checks, guardrails and evaluation, and our list of RAG tools for enterprise AI reviews the frameworks and vendors.
Kanerika Service
RAG Development Services
Kanerika designs and builds RAG applications with permission-aware retrieval, checkable citations, evaluation and production support.
Explore RAG Development Table 5: Build vs buy decision factors
Factor Lean buy Lean build or hybrid Content location One suite, one access model Many systems with different permissions Workflow General questions and answers Specific tasks, records or approvals Audience Employees Customers, partners or regulated users Control needed Default refusals and citations are fine Custom rules, tone and citation format Launch date Weeks A quarter or more is acceptable Team capacity No engineers to maintain it A team that can own it after launch
What to Count in the Cost Comparison Licenses are the visible cost, but rarely the largest. Count integration work, security review, permission sync, evaluation effort, user support and the engineers who keep it running. On the build side, model and search usage grows with adoption, while on the buy side switching costs rise if the vendor’s roadmap drifts from yours.
Common Mistakes That Sink RAG Applications The same six mistakes appear in most stalled projects, but each has a simple fix.
Treating a demo as production-ready. Run the checklist in Table 3 before anyone outside the build team gets access.Checking permissions only at login. Instead, filter every retrieval by the user’s current groups, per document.Showing citations nobody verifies. Sample answers weekly and confirm each source supports its sentence.Launching without a feedback owner. Name the person who reviews thumbs-down reasons and ships fixes.Measuring only model accuracy. Track adoption, resolution, latency and cost per answer as well.Letting content rot. Give every source collection an owner, and show document dates in answers.Some teams then add agents that plan multi-step lookups, which our explainer on agentic RAG covers. Those agents inherit every gap above, so fix the basics first.
Checklist
Generative AI Checklist for Secure AI Adoption and Governance
A practical checklist for rolling out generative AI applications safely, covering data rules, access, evaluation and governance.
Get the Checklist → How Kanerika Builds Production RAG Applications Kanerika treats a RAG application as a product, not a pipeline demo. In our RAG development services , role-based access, audit trails and data masking are part of the architecture from day one. So they are never retrofitted after a security review.
Access control sits inside retrieval, so users only retrieve what they are authorized to see. In addition, every retrieval event is logged for audit, while evaluation measures answer quality continuously and policy guardrails cover sensitive topics.
Timelines follow scope. For example, a focused use case with one document type typically takes 8 to 12 weeks from architecture to production. Multi-source, cross-system or regulated deployments, by contrast, run 3 to 6 months.
The two engagements above show both halves in practice. At the investment bank, role-based access ran inside the retrieval layer. Similarly, the member-support agent handed low-confidence cases to people with a ticket summary and suggested next steps, and member satisfaction rose 25%.
Our view is blunt. Model choice rarely decides whether a RAG application ships, but the access model often does, and the same pitfalls repeat across projects. Permission metadata drifts from the source system, test sets go stale, and nobody owns content after launch. We also see assistants answer beyond their brief because refusals were never tested.
Case Study
43% Faster Document Retrieval for an Investment Bank
Chat-based retrieval across documents and databases with role-based access enforced at the retrieval layer, delivering 100% role-based compliance.
Read the Case Study → For broader generative AI or custom model work, see our generative AI services and LLM development services .
Wrapping Up A RAG application is the product people actually use, built on top of a retrieval pipeline. The pipeline finds and grounds the answer. The application, though, decides who can ask, what they can see, how sources appear and what happens when an answer is wrong.
So pick the application type from the user and the decision, then build in identity, document-level permissions, citations, guardrails and feedback before launch. Move through staged releases with clear exit tests, and measure outcomes as well as accuracy. Buy where the work is generic, and build or go hybrid where the rules are yours.
Frequently Asked Questions
What is a RAG application? A RAG application is a software product that answers questions in plain language using an organization’s own documents and data, with the sources shown. It runs on a retrieval-augmented generation pipeline, then adds the parts users and security teams need, such as sign-in, document-level permissions, citations, guardrails, feedback and monitoring.
How is a RAG application different from a RAG pipeline? The pipeline ingests content, retrieves the passages that match a question and asks a model to answer from them. The application wraps that engine in a product. It verifies who is asking, filters what they may see, presents answers with checkable sources, collects feedback and logs activity for audit and cost tracking.
What are the main types of RAG applications? Five types cover most enterprise demand. They are internal knowledge assistants, customer support copilots and self-service tools, document Q&A and review workspaces, enterprise search with generated answers, and assistants embedded in business software such as CRM or ticketing. Each one has a different primary user, must-have feature and main risk to manage.
What is an example of a RAG application? A common example is an HR or IT assistant that answers employee questions from approved policy documents and links each answer to the source page. Another is a support copilot that drafts replies from product manuals and past tickets, then hands the conversation to a human agent when it is unsure.
How do you keep a RAG application from showing restricted documents? Pass the verified user identity from single sign-on to retrieval and filter results by the user’s current groups before the model sees any text. Store permission metadata with each document, sync it from the source system on a tight schedule, and run permission tests for every role after each index or access change.
How do you test a RAG application before launch? Use a fixed set of real user questions with expected sources, plus questions that should be refused. Score grounded answers and citation accuracy, run permission tests as each role, try prompt injection, and load test at expected peak usage. Agree on pass targets for quality, latency and cost before testing starts.
How do you measure a RAG application after launch? Track four layers. Answer quality covers faithfulness and citation accuracy. Permission tests should show zero restricted-content leaks. User outcomes include resolution rate, repeat usage and thumbs-down reasons. Operations covers p95 latency, uptime and cost per answered question. Tie at least one metric to the business goal set at the start.
How long does it take to move a RAG application from prototype to production? There is no universal timeline. It depends mostly on how many source systems are involved, how complex their permission models are, how much testing the use case demands, and how many users the first release serves. A narrow internal assistant usually moves faster than a customer-facing or regulated application with many sources.
Should you build or buy a RAG application? Buy a packaged assistant when the need is general question answering over content in one suite with its own access model. Build, or use a hybrid of managed retrieval plus your own application layer, when the workflow is specific, sources span several systems, or answers must follow your own rules and citation standards.
What does it cost to run a RAG application? Costs include model usage, search and storage, hosting, integration work, permission sync, security reviews, evaluation, user support and the engineers who maintain it. Model and search spend grows with adoption, so track cost per successfully answered question. Packaged products add license fees, plus possible switching costs later if the vendor roadmap changes.