What is RAG (retrieval-augmented generation)? How it works on your own documents

RAG, short for retrieval-augmented generation, is a way to make an AI model answer from your own documents. The system searches your files for the passages that match a question, hands those passages to the model with the question, and the model writes its answer from them and cites where each fact came from. The model is not retrained, so a corrected document changes the next answer once the index is updated. This guide explains how RAG works in plain words, where it helps a plant or distributor, where it fails, how it compares with fine-tuning, and how to test it on your own data.

A CNC spindle milling an aluminium gear blank in a bright mist of coolant

What is RAG?

AWS defines retrieval-augmented generation as “the process of optimizing the output of a large language model, so it references an authoritative knowledge base outside of its training data sources before generating a response.” A large language model (LLM) is the kind of model behind ChatGPT and Claude. A knowledge base, in this sense, is the set of documents the system is allowed to search: your drawings, procedures, quality records or a shared mailbox.

A model on its own answers from what it learned in training, which is fixed at training time and does not include your private files. RAG adds a search step before the answer, so the model reads the relevant passages from your documents first. AWS notes that RAG does this “all without the need to retrain the model.” A RAG chatbot is the same idea behind a chat window: each question triggers a search, and the reply is written from what the search found.

Where the term comes from

The name comes from a 2020 paper by Patrick Lewis and eleven colleagues at Facebook AI Research, University College London and New York University, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, accepted at the NeurIPS 2020 conference. The authors paired a language model with “a dense vector index of Wikipedia, accessed with a pre-trained neural retriever,” built from 21 million passages of 100 words each. They found that RAG models “generate more specific, diverse and factual language” than a comparable model that answered from its training alone.

The paper also shows why RAG suits documents that change. “Parametric-only models like T5 or BART need further training to update their behavior as the world changes,” the authors write, where T5 and BART are models that answer from training alone. They swapped in an older Wikipedia index to show that the index “can be hot-swapped to update the model without requiring any retraining.”

How RAG works, step by step

Anthropic’s guide to contextual retrieval describes the usual setup. Before anyone asks a question, the system prepares the documents:

  1. Split each document into chunks. Anthropic says chunks are “usually no more than a few hundred tokens.” A token is a small piece of text, often part of a word, so a chunk is a few paragraphs.
  2. Turn each chunk into an embedding. OpenAI’s embeddings guide says “An embedding is a vector (list) of floating point numbers. The distance between two vectors measures their relatedness.” Passages with similar meaning end up close together, even when they use different words.
  3. Store the embeddings in an index. This is often a vector database, a database built to find the nearest embeddings to a new one. Many systems also keep a keyword index. Anthropic says the keyword method BM25 is “particularly effective for queries that include unique identifiers or technical terms,” which in a plant means part numbers, drawing numbers and material grades.

Then, for each question:

  1. Retrieve. The system turns the question into an embedding, finds the closest chunks, runs the keyword search, and keeps the best matches.
  2. Generate. The model receives the question, the retrieved passages and instructions. Anthropic’s guide to reducing hallucinations recommends telling the model to “only use information from provided documents and not its general knowledge,” and giving it permission to say it does not know.
  3. Cite. The answer points to the passages it used. Claude’s citations feature returns the exact cited text, and Anthropic says these citations “are guaranteed to contain valid pointers to the provided documents.” OpenAI’s file search tool returns file citations with its answers.

A small document set may not need retrieval at all. Anthropic says that if a knowledge base is smaller than 200,000 tokens, about 500 pages, it can go into the prompt whole, with no need for RAG.

Where RAG helps a plant or distributor

RAG fits questions whose answer sits in a document someone would otherwise search for by hand. The table shows common document sets in a plant or distribution business and what to watch in each. An NCR, in the third row, is a nonconformance report.

DocumentsA question RAG can answerWhat to watch
Drawings and revisionsWhat tolerance does revision C of this part give for the bore?Scanned drawings need their text extracted first, and Claude’s citations do not cite images inside PDFs
Specifications and work instructionsWhat torque does our work instruction give for this fastener?Superseded revisions must leave the index or be marked as superseded
Quality recordsHave we logged an NCR for this defect on this part before, and what fixed it?Records can name customers and employees, so limit who can search them
Price files and past quotesWhat did we quote this customer for a similar part last year?Prices change often, so read the current price from the ERP through a live connection
EmailWhat lead time did the supplier confirm for this order?Mailboxes belong to people, so each person should search only what they can already open

For scanned documents, text extraction from images covers turning pages into searchable text. For live values such as stock and prices, an MCP server gives the model a controlled connection to the ERP, and many useful assistants combine both. AI for manufacturing and AI for supply chain cover more uses on the floor and in purchasing.

Where RAG fails

Bad retrieval

If the search step misses the right passage, the model cannot use it. Anthropic writes that “traditional RAG solutions remove context when encoding information, which often results in the system failing to retrieve the relevant information from the knowledge base.” Its example is a chunk that reads “The company’s revenue grew by 3% over the previous quarter” and names neither the company nor the quarter. In a plant, the same problem looks like a chunk that gives a torque value without the part it applies to.

Anthropic’s fix has a model write a sentence or two of context for each chunk from the whole document, such as the company and the quarter in its example, and adds it to the chunk before indexing. Across its tests on code, fiction and research papers, that change plus keyword search cut the share of relevant passages missing from the top 20 results by 49%, from 5.7% to 2.9%. Adding a reranking step, a second pass that reorders the results, cut it by 67%, to 1.9%. Those tests did not use plant documents, so measure retrieval on your own.

Stale documents

The index is a copy of your documents, and it only knows what was copied. Amazon’s Bedrock documentation says: “Each time you add, modify, or remove files from your data source, you must sync the data source so that it is re-indexed to the knowledge base.” An old revision left in the index can still be retrieved. The OWASP Gen AI Security Project warns of “knowledge conflict errors” when data from several sources contradicts itself. Sync on a schedule, remove superseded revisions, and show the document date next to each citation.

Access control

A RAG system can show a person a passage they could never open in the original folder. OWASP’s LLM08:2025 entry, on weaknesses in vectors and embeddings, warns that “Inadequate or misaligned access controls can lead to unauthorized access to embeddings containing sensitive information,” and recommends “permission-aware vector and embedding stores.” It adds that attackers can “invert embeddings and recover significant amounts of source information,” so the index needs the same protection as the documents.

Search services support this. Microsoft’s Azure AI Search enforces permissions “from data ingestion through query execution,” and Microsoft notes that permission changes in the source reach search results only after they are synchronized to the index. Microsoft’s Copilot “only surfaces organizational data to which individual users have at least view permissions,” which also means an overshared folder is overshared to the AI.

Answers can still be wrong

Grounding a model in retrieved passages reduces made-up answers without ending them. Anthropic says its techniques “significantly reduce hallucinations” but “don’t eliminate them entirely,” and advises: “Always validate critical information, especially for high-stakes decisions.” For a quote, a tolerance or a reply to a customer, a named person should check the answer against its citation. Human in the loop covers how to set up that check.

Where the index lives

The index holds a copy of your documents’ text, so store it where the documents are allowed to be. AWS says data at rest in its Canada (Central) Region, including Bedrock knowledge bases, stays in that region while the model may process requests in another one (AWS). Private AI for business and sovereign AI cover both questions.

Test RAG on your own documents

Tell Derik which documents your team searches most often. He will suggest how to test retrieval on them before anything is built.

Start a conversation

RAG vs fine-tuning

Fine-tuning further trains a model on your own examples so it behaves differently. RAG changes what the model reads for each question. Microsoft’s guide to customizing LLMs calls prompt engineering, RAG and fine-tuning “complementary methods that in combination can be applicable to a specific use case.” It says RAG is advantageous with “an organization’s private data” or when a model’s public training data “might have become outdated,” and that good cases for fine-tuning include “steering the model to output content in a specific and customized style, tone, or format.”

MethodWhat it changesBest forUpdating factsShows its sources
Better promptsThe instructions and examples sent with each requestFormat, tone and simple rulesEdit the promptOnly for sources pasted into the prompt
RAGThe passages the model reads for each questionAnswers drawn from private or changing documentsSync the indexYes, through citations
Fine-tuningThe model’s weights, the numbers it learned in trainingA consistent format, style or narrow taskTrain againNo

Microsoft also notes that fine-tuning “requires the use of high-quality training data” and has upfront training costs plus “additional hourly costs for hosting the custom model once it’s deployed.” Fine-tuning a language model covers when it pays off and what providers charge.

How to test RAG on your own data

Anthropic’s guide to evaluations says to “Design evals that mirror your real-world task distribution,” and that more test questions with automated grading beat fewer questions graded by hand. For a RAG system, that becomes six checks:

  1. Collect real questions. Ask the people who will use the system for 50 to 100 questions they asked this year, each with the correct answer and the document that holds it.
  2. Grade retrieval on its own. For each question, check whether the right passage is among the passages retrieved. If it is missing, no prompt will fix the answer.
  3. Grade answers and citations. Check that each answer is right and that the cited passage supports it.
  4. Include questions with no answer. The system should say it does not know. Anthropic lists “Irrelevant or nonexistent input data” among the edge cases to test.
  5. Test permissions and freshness. Ask as a person who should not see a document, then change a document, sync, and ask again.
  6. Re-run the same set after every change. A new chunk size, embedding model or prompt can fix one question and break another.

AI training covers how to teach a team to write these questions, and how to build an AI agent shows where a test set fits in a project.

Questions people ask

What is RAG?
RAG, or retrieval-augmented generation, is a way to make an AI model answer from your own documents. The system searches an index of your files for passages that match the question, gives them to the model with the question, and the model answers from them and cites them. The model itself is not retrained.
What does RAG stand for in AI?
RAG stands for retrieval-augmented generation. The term comes from a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research, University College London and New York University, which paired a language model with a searchable index of Wikipedia.
What is a RAG chatbot?
A RAG chatbot is a chat assistant that answers from a set of documents through retrieval-augmented generation. For each question, it searches the documents for relevant passages and passes them to the model, so the answer can cite its sources. Its quality depends on the search step and on how current the documents are.
Is RAG better than fine-tuning?
They solve different problems. RAG suits answers drawn from private or changing documents, and it can cite sources. Fine-tuning further trains a model on examples, which suits a consistent format, style or narrow task. Microsoft describes prompt engineering, RAG and fine-tuning as complementary methods that can be combined.
Does RAG stop AI hallucinations?
RAG reduces made-up answers without stopping them. Retrieved passages, permission to say I don't know and required citations all help, but Anthropic says such techniques don't eliminate them entirely. Check critical answers against the cited source before acting on them.
Is RAG safe for confidential company documents?
It can be, if retrieval respects each person's permissions. OWASP warns that weak access controls in a RAG system can expose sensitive information, and recommends permission-aware vector stores. Treat the index as a copy of your documents: store it where the documents are allowed to live, and sync it when files or permissions change.

How ThriveAI helps

ThriveAI is an AI engineering company in Ottawa. It builds private AI systems on the client’s own data for manufacturers and distributors in Ontario and Quebec, including assistants that answer from drawings, procedures, quality records and email, with a citation for every answer. Derik Lawlis, the founder, leads every project and stays close to the build.

The platform is designed to keep each client’s data, and the index built from it, on its own server in Canada. You choose the model that reads it: one on that server, or a hosted model under a written zero data retention agreement, under which the provider keeps no copy of a request or its answer. A hosted model may process requests outside Canada, so the contract names the model. Retrieval follows each person’s permissions, and a named person approves every action before anything is sent or saved. Enterprise AI platform shows how the pieces fit together, and About ThriveAI covers the company.

Contact

Get answers from your own files, with the source attached

Tell Derik which questions your team answers by digging through files. He will tell you whether RAG can answer them from your documents and how to test it first.

Prefer to talk? Book a meeting.

Your message goes to Derik Lawlis, the founder.