HomeAI CoursesGuidesStorePlatformPricingContact
Sign inFree diagnostic
Home/Glossary/Embeddings
Glossary · Artificial intelligence

What is Embeddings?

Also called: Vector embeddings · Text embeddings · Embedding vectors

Definition

Embeddings are lists of numbers, called vectors, that represent the meaning of a piece of data such as a sentence, document, image or product. An embedding model is trained so that items with similar meaning produce vectors that are close together, typically measured by cosine similarity. This lets software search by meaning rather than exact words, cluster similar items and find related content. Embeddings are the foundation of semantic search and of the retrieval step in retrieval-augmented generation systems.

To use embeddings, each item is passed through an embedding model, which returns a vector with hundreds or thousands of dimensions. These vectors are stored in an index. At query time, the query is embedded in the same way, and the system finds stored vectors closest to it. Because similarity is about meaning, a search for "leave policy for new parents" can find a document titled "maternity and paternity benefits".

In business applications, embeddings power semantic search over internal documents, RAG assistants, recommendation of related courses or products, duplicate detection, and grouping of customer feedback by theme. They are usually cheap to compute compared with generating text.

Common mistakes include mixing vectors from different embedding models in one index, embedding very long or poorly split chunks, and relying on semantic search alone when users search for exact codes or names, where keyword search works better. Embeddings must also be regenerated if you change the embedding model.

Key points

  • Vectors that represent meaning numerically.
  • Similar meanings produce nearby vectors.
  • Enable semantic search, clustering and recommendations.
  • Core to the retrieval step in RAG.
  • Use one embedding model per index; combine with keyword search.

An example at work

An edtech company in Bengaluru embeds every lesson description so that when a learner finishes a module on Excel pivot tables, the platform recommends related lessons on data summarisation, even when titles share no words.

Where this is used at Bodhih

LLM EngineeringLevel 3: embeddings and retrieval in practice.Guide: RAG, agents and evaluation

Related terms

Retrieval-augmented generation

Retrieval-augmented generation (RAG) is a technique where an AI system retrieves relevant documents first and gives them to an LLM to ground its answer.

Large language model

A large language model (LLM) is an AI model trained on vast amounts of text to understand and generate language by predicting the next token.

Machine learning

Machine learning is a branch of AI in which computers learn patterns from data to make predictions or decisions without being explicitly programmed for each rule.

Fine-tuning

Fine-tuning is further training of a pre-trained AI model on a smaller, task-specific dataset to adapt its behaviour, style or performance.

LLM evaluation

LLM evaluation is the systematic testing of a language model or LLM application against defined criteria to measure quality, accuracy and safety.

Bodhih Training Solutions, Bengaluru · Updated 2 October 2026 · All 51 terms
Common questions

More about Embeddings

What is the difference between embeddings and a vector database?

Embeddings are the vectors themselves, produced by an embedding model. A vector database, or vector index, is where those vectors are stored and searched efficiently for the nearest matches. You create embeddings first, then store them in a vector index to retrieve similar items quickly.

Why are embeddings important for RAG?

In RAG, the system must find the passages most relevant to a question before the LLM answers. Embeddings let it match by meaning, so relevant passages are found even when they use different wording from the question. Poor embeddings or chunking lead to poor retrieval, and therefore poor answers.

Corporate training since 2008, now measured. Part of a family with AssessAll, Jobulary and Pewple — one shared record of a person.

963, 2nd Floor, 3rd Cross, 1st Block,
HRBR Layout, Bengaluru 560043, India
solutions@bodhih.com

Product

AI coursesGuidesStorePlatformPricingAll courses

AI courses

AI for Sales courseAI for Marketing courseAI for HR courseAI for Finance courseAI for Managers courseAI/ML Foundations courseApplied AI/ML Practitioner courseLLM Engineering course

English at work

Workplace English: business English course

Resources

AnswersGlossaryCompetency frameworkSolutionsIndustriesLocationsTrainer Toolkits

Family

AssessAll — measurementJobulary — the SpinePewple — coachingBodhih Training — classroom

Company

About BodhihContactBook a diagnosticSign inTerms of usePrivacy policyRefunds
© 2026 Bodhih Training Solutions Private Limited · BengaluruMon–Fri, 9 AM – 6 PM IST