Which parts of your eval framework you can stop maintaining

Anyone who has put an LLM feature into production has probably written evaluation code, and after a few of them you are maintaining a small private framework: golden cases, a judge prompt, thresholds, a scheduled job. I came across Google's managed Gen AI evaluation service while looking into something else, and the question it raised was which parts of that framework I could stop maintaining. Working through it loop by loop produced what I think is a fairly clean rule.

What a managed judge is actually for

Buy the managed service for the loops where you would otherwise be writing and calibrating a judge, which in practice means subjective quality scoring of generated text. That is where it gives you something you cannot easily build, namely adaptive rubrics, which generate pass or fail tests per prompt rather than applying one fixed scale to everything. For example, a question about a fee schedule and one about an account balance get different checks rather than a shared helpfulness ladder. There is a second argument that has nothing to do with effort, which is that a judge you wrote yourself is one you can gradually tune until it flatters your own system.

What does not move

Three kinds of loop should stay yours, for three different reasons.

The first is any check that does not need a model. A gate replaying deterministic routing or business logic against expected results is a contract test, so putting a hosted judge behind it buys nondeterminism, latency and a per-run bill in exchange for nothing.

The second is drift detection, because comparing a run against a stored baseline with tolerance bands is what makes a canary a canary, and managed services score a run in isolation rather than against your history.

The third is anything needing provenance for the judge itself. If you record which model version served each scored row, so a quality drop can be attributed rather than merely noticed, a managed autorater cannot help, because its version belongs to the provider.

Share the inputs, keep the runners apart

Hooking a managed loop into a framework you already have is mostly a question of what to share. Share the inputs, meaning the same golden cases the rest of your evaluation uses, so the numbers describe one dataset rather than two that quietly drift apart. For example, my drift canary and the managed loop replay the same seven cases, which is what makes their scores comparable. Derive any policy rubric from whatever machine-readable source your system is already instructed from, so it cannot fall behind the rules.

Keep the runner, dependencies, schedule and gate separate. The cost profiles differ enough that the paid loop should not be able to run in the same job as the free one by accident, and a contract gate that quietly acquires a large service SDK has stopped being cheap.

A banking assistant answering an account balance question, grounded in the output of the tool it called

What adoption costs

Adoption is not free, and the failures do not announce themselves. Testing these against a reference banking assistant of mine produced four faults in a session, each returning a plausible number rather than an error.

The first run scored the hallucination metric at 1.0, the best possible, and I nearly recorded that as evidence the assistant was well grounded. Instead I fed it a fabricated answer, a balance of 98,700 dollars and a pending wire to Belize, against tool output saying 1,234.56, and it scored that one 1.0 as well. The metric reads its evidence from the prompt and nowhere else, and I had supplied it in a context column, which the SDK forwards but that metric ignores. With nothing to check against it treated the answer as its own source.

The same fabricated answer scored four ways: 1.0 via a context column, 1.0 via instruction, 1.0 via reference, and 0.0 when the evidence is placed in the prompt

The other three have the same shape. A custom metric's judge output is parsed as JSON while the SDK's prompt builder emits prose, so the judge reasons correctly and the service reports null anyway. The autorater is also noisy enough that the same fabrication scored 0.0, 0.0, 0.0 and then 1.0 over four runs, so anything you gate on wants majority voting.

Worse than any of those, a gate that skipped a null score as non-numeric reported "gates passed" for a policy it had never measured, and a missing control is at least visible where a green one is reassuring.

Reading the output, not the score

The step worth being deliberate about is building cases where you already know the answer, running them through the metric, and reading what the judge wrote rather than what it scored. The number in my case looked entirely healthy, and what gave the fault away was the per-case rationale, where the supporting excerpt cited for the fabricated balance was the fabricated balance itself.

So before reporting a grounding number, send the metric an answer you know is false against evidence you control, confirm it scores badly, then send a clean one and confirm it does not. That costs a handful of calls and a few minutes of reading, and it is the difference between a metric that works and one that returns a number.

0 Comments

Leave a Comment