What is Embeddings?
Also called: Vector embeddings · Text embeddings · Embedding vectorsDefinition
Embeddings are lists of numbers, called vectors, that represent the meaning of a piece of data such as a sentence, document, image or product. An embedding model is trained so that items with similar meaning produce vectors that are close together, typically measured by cosine similarity. This lets software search by meaning rather than exact words, cluster similar items and find related content. Embeddings are the foundation of semantic search and of the retrieval step in retrieval-augmented generation systems.
To use embeddings, each item is passed through an embedding model, which returns a vector with hundreds or thousands of dimensions. These vectors are stored in an index. At query time, the query is embedded in the same way, and the system finds stored vectors closest to it. Because similarity is about meaning, a search for "leave policy for new parents" can find a document titled "maternity and paternity benefits".
In business applications, embeddings power semantic search over internal documents, RAG assistants, recommendation of related courses or products, duplicate detection, and grouping of customer feedback by theme. They are usually cheap to compute compared with generating text.
Common mistakes include mixing vectors from different embedding models in one index, embedding very long or poorly split chunks, and relying on semantic search alone when users search for exact codes or names, where keyword search works better. Embeddings must also be regenerated if you change the embedding model.
Key points
- Vectors that represent meaning numerically.
- Similar meanings produce nearby vectors.
- Enable semantic search, clustering and recommendations.
- Core to the retrieval step in RAG.
- Use one embedding model per index; combine with keyword search.
An example at work
An edtech company in Bengaluru embeds every lesson description so that when a learner finishes a module on Excel pivot tables, the platform recommends related lessons on data summarisation, even when titles share no words.
Where this is used at Bodhih
Related terms
Retrieval-augmented generation
Retrieval-augmented generation (RAG) is a technique where an AI system retrieves relevant documents first and gives them to an LLM to ground its answer.
Large language model
A large language model (LLM) is an AI model trained on vast amounts of text to understand and generate language by predicting the next token.
Machine learning
Machine learning is a branch of AI in which computers learn patterns from data to make predictions or decisions without being explicitly programmed for each rule.
Fine-tuning
Fine-tuning is further training of a pre-trained AI model on a smaller, task-specific dataset to adapt its behaviour, style or performance.
LLM evaluation
LLM evaluation is the systematic testing of a language model or LLM application against defined criteria to measure quality, accuracy and safety.
More about Embeddings
What is the difference between embeddings and a vector database?
Embeddings are the vectors themselves, produced by an embedding model. A vector database, or vector index, is where those vectors are stored and searched efficiently for the nearest matches. You create embeddings first, then store them in a vector index to retrieve similar items quickly.
Why are embeddings important for RAG?
In RAG, the system must find the passages most relevant to a question before the LLM answers. Embeddings let it match by meaning, so relevant passages are found even when they use different wording from the question. Poor embeddings or chunking lead to poor retrieval, and therefore poor answers.