What is Large language model?
Also called: LLM · Foundation modelDefinition
A large language model, or LLM, is an artificial intelligence model trained on very large amounts of text to predict the next token, a word or part of a word, in a sequence. Through this training, and further tuning on instructions and human feedback, it learns to answer questions, summarise, translate, write and reason over text. Most LLMs use the transformer architecture. They power assistants such as ChatGPT, Claude and Gemini, but can produce confident errors, so their output needs checking.
An LLM is first pre-trained on a broad text corpus, learning patterns of language and a great deal of general knowledge. It is then refined, typically with instruction tuning and reinforcement learning from human feedback, to follow requests helpfully and safely. At run time, the model reads the prompt within its context window and generates a response one token at a time.
For organisations, LLMs are general-purpose language engines. They can be used directly through chat tools, or built into applications via an API, often combined with retrieval over company documents, tools the model can call, and evaluation to check quality. This is the basis of most enterprise generative AI use today.
Common misunderstandings are that an LLM looks facts up (it does not, unless connected to retrieval or search), that it remembers past conversations by default, and that bigger is always better. The right model depends on the task, cost, speed, data-handling terms and measured quality on your own examples.
Key points
- Trained to predict the next token on very large text datasets.
- Most use the transformer architecture.
- Refined with instruction tuning and human feedback.
- Has a limited context window per request.
- Can produce fluent but incorrect output.
An example at work
A Hyderabad GCC builds an internal assistant on an LLM accessed via API, connecting it to the company’s HR policy documents so employees get answers that cite the relevant policy section.
Where this is used at Bodhih
Related terms
Generative AI
Generative AI is a type of artificial intelligence that creates new content, such as text, images, audio, video or code, from patterns learnt from data.
Prompt engineering
Prompt engineering is the practice of writing and refining instructions to a generative AI model so it produces accurate, useful output for a task.
Retrieval-augmented generation
Retrieval-augmented generation (RAG) is a technique where an AI system retrieves relevant documents first and gives them to an LLM to ground its answer.
Fine-tuning
Fine-tuning is further training of a pre-trained AI model on a smaller, task-specific dataset to adapt its behaviour, style or performance.
Embeddings
Embeddings are numerical vectors that represent the meaning of text, images or other data, so that similar items sit close together mathematically.
AI hallucination
An AI hallucination is output from a generative AI model that sounds confident and plausible but is false, unsupported or made up.
More about Large language model
How does a large language model work?
It converts text into tokens, processes them through many transformer layers that weigh how each token relates to the others, and predicts the most likely next token. Repeating this produces a response. Its abilities come from patterns learnt during training on large text datasets and later tuning on instructions and feedback.
What is the difference between an LLM and ChatGPT?
An LLM is the underlying model. ChatGPT is a product, a chat application built on top of LLMs developed by OpenAI, with an interface, safety systems and extra features. Claude and Gemini are similarly products built on their developers’ own models. The same LLM can also be used in other applications through an API.