What is LLM evaluation?
Also called: LLM evals · Evals · AI evaluation · Model evaluationDefinition
LLM evaluation is the practice of systematically measuring how well a large language model or an application built on one performs on the tasks it is meant to do. Teams build a test set of realistic inputs with expected outcomes or scoring criteria, run the system, and score outputs using exact checks, human reviewers, or another model acting as a judge. Evaluation covers accuracy, faithfulness to sources, format, safety, cost and latency, and is repeated whenever prompts, models or data change.
An evaluation set should reflect real usage, including difficult and edge cases. Some criteria can be checked automatically, such as valid JSON or a correct classification label. Open-ended answers need a rubric scored by people or by an LLM judge, which must itself be checked against human ratings. For RAG systems, retrieval quality and answer faithfulness are measured separately.
Evaluation turns LLM development from guesswork into engineering. Without it, a team cannot tell whether a new prompt, model or chunking strategy is better or worse, and regressions reach users unnoticed. With it, changes can be compared on numbers, and production monitoring can flag drift.
Common mistakes are relying on public benchmark scores instead of testing on your own tasks, using a handful of hand-picked examples, trusting an LLM judge without calibration, and evaluating only once before launch. Evaluation is a continuous part of operating an LLM system.
Key points
- Test on a realistic, representative dataset of your own tasks.
- Combine automatic checks, human review and calibrated LLM judges.
- For RAG, measure retrieval and faithfulness separately.
- Rerun on every prompt, model or data change.
- Public benchmarks do not replace task-specific evaluation.
An example at work
A Bengaluru legal-tech start-up keeps 300 real, anonymised contract questions with lawyer-approved answers and reruns them whenever it changes its prompt or model, blocking release if faithfulness scores drop.
Where this is used at Bodhih
Related terms
Large language model
A large language model (LLM) is an AI model trained on vast amounts of text to understand and generate language by predicting the next token.
Retrieval-augmented generation
Retrieval-augmented generation (RAG) is a technique where an AI system retrieves relevant documents first and gives them to an LLM to ground its answer.
AI hallucination
An AI hallucination is output from a generative AI model that sounds confident and plausible but is false, unsupported or made up.
AI agent
An AI agent is a system in which an LLM plans and takes actions, using tools such as search, code or business software, to complete a goal over several steps.
Responsible AI
Responsible AI is the practice of designing, deploying and using AI so that it is fair, safe, transparent, accountable and respects privacy.
Fine-tuning
Fine-tuning is further training of a pre-trained AI model on a smaller, task-specific dataset to adapt its behaviour, style or performance.
More about LLM evaluation
How do you evaluate an LLM application?
Collect realistic inputs, define what a good output looks like for each, and score the application’s outputs using automatic checks where possible and human or calibrated LLM-judge scoring where not. Track metrics such as accuracy, faithfulness, safety, cost and latency, and rerun the set after every change.
What is LLM-as-a-judge?
LLM-as-a-judge means using a language model to score another model’s outputs against a rubric. It scales evaluation of open-ended answers, but judges can be biased or inconsistent, so their scores should be checked against a sample of human ratings before being trusted.