What is RAG (retrieval-augmented generation)? How it works on your own documents
RAG, short for retrieval-augmented generation, is a way to make an AI model answer from your own documents. The system searches your files for the passages that match a question, hands those passages to the model with the question, and the model writes its answer from them and cites where each fact came from. The model is not retrained, so a corrected document changes the next answer once the index is updated. This guide explains how RAG works in plain words, where it helps a plant or distributor, where it fails, how it compares with fine-tuning, and how to test it on your own data.

What is RAG?
AWS defines retrieval-augmented generation as “the process of optimizing the output of a large language model, so it references an authoritative knowledge base outside of its training data sources before generating a response.” A large language model (LLM) is the kind of model behind ChatGPT and Claude. A knowledge base, in this sense, is the set of documents the system is allowed to search: your drawings, procedures, quality records or a shared mailbox.
A model on its own answers from what it learned in training, which is fixed at training time and does not include your private files. RAG adds a search step before the answer, so the model reads the relevant passages from your documents first. AWS notes that RAG does this “all without the need to retrain the model.” A RAG chatbot is the same idea behind a chat window: each question triggers a search, and the reply is written from what the search found.
Where the term comes from
The name comes from a 2020 paper by Patrick Lewis and eleven colleagues at Facebook AI Research, University College London and New York University, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, accepted at the NeurIPS 2020 conference. The authors paired a language model with “a dense vector index of Wikipedia, accessed with a pre-trained neural retriever,” built from 21 million passages of 100 words each. They found that RAG models “generate more specific, diverse and factual language” than a comparable model that answered from its training alone.
The paper also shows why RAG suits documents that change. “Parametric-only models like T5 or BART need further training to update their behavior as the world changes,” the authors write, where T5 and BART are models that answer from training alone. They swapped in an older Wikipedia index to show that the index “can be hot-swapped to update the model without requiring any retraining.”
How RAG works, step by step
Anthropic’s guide to contextual retrieval describes the usual setup. Before anyone asks a question, the system prepares the documents:
- Split each document into chunks. Anthropic says chunks are “usually no more than a few hundred tokens.” A token is a small piece of text, often part of a word, so a chunk is a few paragraphs.
- Turn each chunk into an embedding. OpenAI’s embeddings guide says “An embedding is a vector (list) of floating point numbers. The distance between two vectors measures their relatedness.” Passages with similar meaning end up close together, even when they use different words.
- Store the embeddings in an index. This is often a vector database, a database built to find the nearest embeddings to a new one. Many systems also keep a keyword index. Anthropic says the keyword method BM25 is “particularly effective for queries that include unique identifiers or technical terms,” which in a plant means part numbers, drawing numbers and material grades.
Then, for each question:
- Retrieve. The system turns the question into an embedding, finds the closest chunks, runs the keyword search, and keeps the best matches.
- Generate. The model receives the question, the retrieved passages and instructions. Anthropic’s guide to reducing hallucinations recommends telling the model to “only use information from provided documents and not its general knowledge,” and giving it permission to say it does not know.
- Cite. The answer points to the passages it used. Claude’s citations feature returns the exact cited text, and Anthropic says these citations “are guaranteed to contain valid pointers to the provided documents.” OpenAI’s file search tool returns file citations with its answers.
A small document set may not need retrieval at all. Anthropic says that if a knowledge base is smaller than 200,000 tokens, about 500 pages, it can go into the prompt whole, with no need for RAG.
Where RAG helps a plant or distributor
RAG fits questions whose answer sits in a document someone would otherwise search for by hand. The table shows common document sets in a plant or distribution business and what to watch in each. An NCR, in the third row, is a nonconformance report.
| Documents | A question RAG can answer | What to watch |
|---|---|---|
| Drawings and revisions | What tolerance does revision C of this part give for the bore? | Scanned drawings need their text extracted first, and Claude’s citations do not cite images inside PDFs |
| Specifications and work instructions | What torque does our work instruction give for this fastener? | Superseded revisions must leave the index or be marked as superseded |
| Quality records | Have we logged an NCR for this defect on this part before, and what fixed it? | Records can name customers and employees, so limit who can search them |
| Price files and past quotes | What did we quote this customer for a similar part last year? | Prices change often, so read the current price from the ERP through a live connection |
| What lead time did the supplier confirm for this order? | Mailboxes belong to people, so each person should search only what they can already open |
For scanned documents, text extraction from images covers turning pages into searchable text. For live values such as stock and prices, an MCP server gives the model a controlled connection to the ERP, and many useful assistants combine both. AI for manufacturing and AI for supply chain cover more uses on the floor and in purchasing.
Where RAG fails
Bad retrieval
If the search step misses the right passage, the model cannot use it. Anthropic writes that “traditional RAG solutions remove context when encoding information, which often results in the system failing to retrieve the relevant information from the knowledge base.” Its example is a chunk that reads “The company’s revenue grew by 3% over the previous quarter” and names neither the company nor the quarter. In a plant, the same problem looks like a chunk that gives a torque value without the part it applies to.
Anthropic’s fix has a model write a sentence or two of context for each chunk from the whole document, such as the company and the quarter in its example, and adds it to the chunk before indexing. Across its tests on code, fiction and research papers, that change plus keyword search cut the share of relevant passages missing from the top 20 results by 49%, from 5.7% to 2.9%. Adding a reranking step, a second pass that reorders the results, cut it by 67%, to 1.9%. Those tests did not use plant documents, so measure retrieval on your own.
Stale documents
The index is a copy of your documents, and it only knows what was copied. Amazon’s Bedrock documentation says: “Each time you add, modify, or remove files from your data source, you must sync the data source so that it is re-indexed to the knowledge base.” An old revision left in the index can still be retrieved. The OWASP Gen AI Security Project warns of “knowledge conflict errors” when data from several sources contradicts itself. Sync on a schedule, remove superseded revisions, and show the document date next to each citation.
Access control
A RAG system can show a person a passage they could never open in the original folder. OWASP’s LLM08:2025 entry, on weaknesses in vectors and embeddings, warns that “Inadequate or misaligned access controls can lead to unauthorized access to embeddings containing sensitive information,” and recommends “permission-aware vector and embedding stores.” It adds that attackers can “invert embeddings and recover significant amounts of source information,” so the index needs the same protection as the documents.
Search services support this. Microsoft’s Azure AI Search enforces permissions “from data ingestion through query execution,” and Microsoft notes that permission changes in the source reach search results only after they are synchronized to the index. Microsoft’s Copilot “only surfaces organizational data to which individual users have at least view permissions,” which also means an overshared folder is overshared to the AI.
Answers can still be wrong
Grounding a model in retrieved passages reduces made-up answers without ending them. Anthropic says its techniques “significantly reduce hallucinations” but “don’t eliminate them entirely,” and advises: “Always validate critical information, especially for high-stakes decisions.” For a quote, a tolerance or a reply to a customer, a named person should check the answer against its citation. Human in the loop covers how to set up that check.
The index holds a copy of your documents’ text, so store it where the documents are allowed to be. AWS says data at rest in its Canada (Central) Region, including Bedrock knowledge bases, stays in that region while the model may process requests in another one (AWS). Private AI for business and sovereign AI cover both questions.
Test RAG on your own documents
Tell Derik which documents your team searches most often. He will suggest how to test retrieval on them before anything is built.
Start a conversationRAG vs fine-tuning
Fine-tuning further trains a model on your own examples so it behaves differently. RAG changes what the model reads for each question. Microsoft’s guide to customizing LLMs calls prompt engineering, RAG and fine-tuning “complementary methods that in combination can be applicable to a specific use case.” It says RAG is advantageous with “an organization’s private data” or when a model’s public training data “might have become outdated,” and that good cases for fine-tuning include “steering the model to output content in a specific and customized style, tone, or format.”
| Method | What it changes | Best for | Updating facts | Shows its sources |
|---|---|---|---|---|
| Better prompts | The instructions and examples sent with each request | Format, tone and simple rules | Edit the prompt | Only for sources pasted into the prompt |
| RAG | The passages the model reads for each question | Answers drawn from private or changing documents | Sync the index | Yes, through citations |
| Fine-tuning | The model’s weights, the numbers it learned in training | A consistent format, style or narrow task | Train again | No |
Microsoft also notes that fine-tuning “requires the use of high-quality training data” and has upfront training costs plus “additional hourly costs for hosting the custom model once it’s deployed.” Fine-tuning a language model covers when it pays off and what providers charge.
How to test RAG on your own data
Anthropic’s guide to evaluations says to “Design evals that mirror your real-world task distribution,” and that more test questions with automated grading beat fewer questions graded by hand. For a RAG system, that becomes six checks:
- Collect real questions. Ask the people who will use the system for 50 to 100 questions they asked this year, each with the correct answer and the document that holds it.
- Grade retrieval on its own. For each question, check whether the right passage is among the passages retrieved. If it is missing, no prompt will fix the answer.
- Grade answers and citations. Check that each answer is right and that the cited passage supports it.
- Include questions with no answer. The system should say it does not know. Anthropic lists “Irrelevant or nonexistent input data” among the edge cases to test.
- Test permissions and freshness. Ask as a person who should not see a document, then change a document, sync, and ask again.
- Re-run the same set after every change. A new chunk size, embedding model or prompt can fix one question and break another.
AI training covers how to teach a team to write these questions, and how to build an AI agent shows where a test set fits in a project.
Questions people ask
What is RAG?
What does RAG stand for in AI?
What is a RAG chatbot?
Is RAG better than fine-tuning?
Does RAG stop AI hallucinations?
Is RAG safe for confidential company documents?
How ThriveAI helps
ThriveAI is an AI engineering company in Ottawa. It builds private AI systems on the client’s own data for manufacturers and distributors in Ontario and Quebec, including assistants that answer from drawings, procedures, quality records and email, with a citation for every answer. Derik Lawlis, the founder, leads every project and stays close to the build.
The platform is designed to keep each client’s data, and the index built from it, on its own server in Canada. You choose the model that reads it: one on that server, or a hosted model under a written zero data retention agreement, under which the provider keeps no copy of a request or its answer. A hosted model may process requests outside Canada, so the contract names the model. Retrieval follows each person’s permissions, and a named person approves every action before anything is sent or saved. Enterprise AI platform shows how the pieces fit together, and About ThriveAI covers the company.