LLM engineering: what RAG, agents and evaluation actually require in production
By Bodhih Training Solutions · UpdatedThe short answer
In production, LLM engineering requires four things a demo does not. Retrieval choices such as chunking and re-ranking must be made by measurement on a fixed question set. Agents need guardrails that have been attacked, not assumed. Evaluation needs a golden set, separate retrieval and answer metrics, and a gate that fails a bad release. The system also needs a latency budget, a cost model and an audit trail.
What does LLM engineering involve once the demo works?
A demo proves that a model can answer a question when the right text is put in front of it. A production system has to answer thousands of questions it was never shown, at a cost someone has approved, within a response time users will tolerate, and with a record of what it did.
The gap between the two is engineering discipline rather than a cleverer prompt. Most of the work sits in four places: how context is retrieved, how actions are constrained, how quality is measured, and how the system is operated and accounted for. The table below sets out what each one asks of you.
| Part of the system | What production requires | What a reviewer should ask to see |
|---|---|---|
| Retrieval (RAG) | Chunking, embedding and re-ranking choices compared on one question set | The comparison table and the reason for the choice |
| Agents and tools | Validated arguments, an allowed-tools list, a step budget, defined termination | The results of an injection test that was actually run |
| Evaluation | A golden set, retrieval and answer metrics kept apart, a regression gate | A release the gate blocked, and why |
| Operations | A latency budget per step, a cost model at real volume, traces, rollback | The measured p95 under a stated load |
| Governance | Residency, PII handling, an audit trail, an incident runbook | One worked incident, start to finish |
What does RAG require beyond a vector database?
Retrieval augmented generation fails most often at retrieval, not generation. If the right passage is not in the context, no prompt will rescue the answer. So the first job is to measure whether the right passage arrives, using a metric such as recall@k on a set of questions with known source documents.
Chunk size, overlap and structure-aware splitting should be treated as hypotheses. Try several configurations on the same questions and record both recall and the context cost each one carries. Sometimes the configurations land within noise of each other, and the correct decision is to take the cheapest. That is still a measured decision.
Embedding search is not always better than lexical search. Exact identifiers, product codes and names often retrieve better by keyword, which is why hybrid search and a re-ranking step earn their place. Near-duplicate documents also deserve a rule, because they can fill the context with the same passage several times.
What do AI agents need before they touch real tools?
An agent is a loop: the model proposes an action, the system runs a tool, and the result goes back to the model. Every turn of that loop is a place where something can go wrong, so the design question is what the system does when a tool fails, when arguments are malformed, and when the loop should stop.
- Validate every tool argument against a schema before the tool runs.
- Keep an explicit list of allowed tools, and refuse anything outside the task’s scope.
- Set a step budget so a confused agent stops instead of looping.
- Scrub personal data before it reaches a tool or a log.
- Run tools with side effects in a sandbox or behind a human approval.
How does prompt injection reach a system that only reads documents?
Prompt injection does not need a malicious user. An instruction can arrive inside a retrieved document, a web page or a tool result, and the model may treat it as part of its task. Any system that feeds retrieved text to a model with tools attached is exposed.
The useful response is to attack your own system and write down what happened. A short table of attempts, outcomes and the changes you made is worth more than a claim that guardrails exist. An attack that succeeded and was fixed tells you more than one that was never tried.
How do you evaluate an LLM system at scale?
Start with a golden set: questions with expected answers and expected sources, grown from real usage and kept out of anything the system is tuned on. Report retrieval metrics and answer metrics separately, because a single blended score cannot tell you which half broke.
Using a model as a judge is practical, but its validity has to be shown. Label a sample by hand and check that the judge agrees with you before trusting it on the rest. Where several people label, check that they agree with each other too.
The failure to design for is the silent regression: headline metrics hold steady while the system begins giving wrong answers with valid citations. Catching it needs checks on whether the cited passage supports the claim, and a gate in the release pipeline that can fail.
What do cost, latency and caching look like at real volume?
Break the response time into its steps: retrieval, re-ranking, the model call, any tool calls and post-processing. Give each a budget and measure the median and the tail under a stated load. Streaming improves perceived latency without changing the total, which matters for how you report it.
Model cost per thousand questions at the volumes you expect, in the currency your finance team uses. Caching lowers both cost and latency, but each kind carries a correctness risk. An exact-match cache is safe and rarely hits. A semantic cache hits often and can return an answer to a question that only looked similar.
Version prompts and models together, trace every request, and release changes to a small share of traffic first with a rollback that has been rehearsed.
What does governance mean for an engineer?
Governance is the set of questions someone outside the team will eventually ask. Where do the documents, the embeddings and the prompts physically reside? What personal data enters the system, and what does redaction cost in wrongly removed text? Can you reconstruct what the system answered on a given day, and from which sources?
A model card written for the whole system, an incident runbook and a statement of where a human reviews the output answer most of them. This is general engineering guidance, not legal advice; your organisation’s counsel decides what the law requires.
How do you build these skills deliberately?
Reading helps, but each of these skills is learned by doing it on a system you own and having someone qualified question the result. Pick one system, set numeric success criteria before building, and work through retrieval, guardrails, evaluation and operations in that order.
Bodhih’s LLM Engineering pathway is structured this way: five modules around one reference system, then a six-week capstone with reviewed checkpoints and a recorded viva. Whether or not you take a course, the checklist in the table above is a fair test of any LLM system you are asked to sign off.
Build one of these systems and have it examined
The LLM Engineering certification from Bodhih takes about 45 hours: five modules, a six-week capstone marked by two people, a recorded viva with two evaluators and a certification exam under exam conditions. It assumes you already write Python and have shipped something. It is ₹64,999 plus GST for an individual, with a verifiable credential at the end.
Questions people ask next
What is the difference between prompt engineering and LLM engineering?
Prompt engineering is the craft of writing instructions that get a useful response from a model. LLM engineering is the wider discipline of building a dependable system around the model: retrieval, tool use, guardrails, evaluation, cost and latency control, and governance. A good prompt is one component of that system and is rarely the reason a production system fails.
What is a golden set in LLM evaluation?
A golden set is a curated collection of questions with expected answers and expected source documents, used to test an LLM system repeatedly. A golden set should be grown from real usage, kept separate from anything the system is tuned on, and run before every release so that a change that lowers quality is caught before users see it.
Is RAG better than fine-tuning?
RAG and fine-tuning solve different problems, so neither is better in general. RAG supplies facts the model does not hold, such as your own documents, and lets you cite sources and update content without retraining. Fine-tuning changes how a model behaves or writes. Many production systems begin with RAG because the knowledge changes often and answers need to be traceable.
Can an LLM reliably judge another LLM’s answers?
An LLM can judge another LLM’s answers usefully, but only after its judgements have been checked against human labels. Hand-label a sample, compare the judge’s verdicts with yours, and look for biases such as favouring longer or more confident answers. An unvalidated judge gives a number that looks like a measurement without being one.
How do you stop an AI agent from taking a harmful action?
You stop an AI agent from taking a harmful action by constraining it in code rather than in the prompt. Validate tool arguments against a schema, keep an allowed-tools list, set a step budget, run side-effecting tools in a sandbox or behind human approval, and test the whole arrangement with injection attempts through retrieved text and tool results.