What is Retrieval-augmented generation?
Also called: RAGDefinition
Retrieval-augmented generation, or RAG, is a technique that improves the answers of a large language model by first retrieving relevant information from a chosen source, such as company documents or a database, and adding it to the prompt. The model then generates its response grounded in that retrieved material, ideally citing it. RAG lets an LLM use current, private or specialised knowledge without retraining, and reduces, though does not eliminate, hallucination. Its quality depends heavily on how well the retrieval step finds the right passages.
A typical RAG pipeline has two stages. First, documents are split into chunks, converted into embeddings and stored in a vector index, often alongside a keyword index. Second, when a user asks a question, the system retrieves the most relevant chunks, may re-rank them, and passes them to the LLM with instructions to answer only from that material and cite sources.
RAG is one of the most common patterns for enterprise generative AI because it keeps knowledge in documents the organisation already controls. Updating the answer is as simple as updating the document, access controls can be enforced at retrieval time, and citations let users check the source.
Most RAG failures are retrieval failures: poor chunking, missing metadata, no keyword search for exact terms such as product codes, or retrieving too many loosely related passages. Teams also skip evaluation, so they cannot tell whether a change improved or worsened answers. Measuring retrieval quality and answer faithfulness separately is essential.
Key points
- Retrieve relevant content first, then generate the answer from it.
- Uses embeddings and a vector index, often with keyword search.
- Adds current or private knowledge without retraining.
- Reduces but does not remove hallucination.
- Evaluate retrieval and answer faithfulness separately.
An example at work
A Chennai-based NBFC builds a RAG assistant over its credit policy manuals so branch officers can ask eligibility questions and receive answers that quote and link the exact policy clause.
Where this is used at Bodhih
Related terms
Embeddings
Embeddings are numerical vectors that represent the meaning of text, images or other data, so that similar items sit close together mathematically.
Large language model
A large language model (LLM) is an AI model trained on vast amounts of text to understand and generate language by predicting the next token.
Fine-tuning
Fine-tuning is further training of a pre-trained AI model on a smaller, task-specific dataset to adapt its behaviour, style or performance.
LLM evaluation
LLM evaluation is the systematic testing of a language model or LLM application against defined criteria to measure quality, accuracy and safety.
AI hallucination
An AI hallucination is output from a generative AI model that sounds confident and plausible but is false, unsupported or made up.
AI agent
An AI agent is a system in which an LLM plans and takes actions, using tools such as search, code or business software, to complete a goal over several steps.
More about Retrieval-augmented generation
What is the difference between RAG and fine-tuning?
RAG gives the model relevant information at question time, so it suits knowledge that changes or must be cited. Fine-tuning changes the model’s weights by training on examples, so it suits teaching a consistent style, format or narrow task. They are complementary: many systems use RAG for facts and fine-tuning, if needed, for behaviour.
Does RAG stop AI hallucinations?
No, it reduces them. If retrieval returns the wrong passage, or the model goes beyond the retrieved text, the answer can still be wrong. Good RAG systems instruct the model to answer only from sources, show citations, say when the information is not available, and are evaluated regularly on a test set of real questions.
Do I need a vector database for RAG?
You need some way to search your content. A vector index enables semantic search using embeddings, and many teams combine it with keyword search for exact matches. For a small document set, a simple in-memory index can be enough; dedicated vector databases help at larger scale or with complex filtering.