Pixzest — Zest for Innovation
Start a project
Pixzest — Zest for Innovation
Start a project
Quality

Evals are the new unit tests

LLM features don't fail like normal code. Evaluation suites are how teams keep them improving instead of quietly regressing.

Traditional tests assume the same input gives the same output. Features built on large language models break that assumption: the same prompt can produce different wording, a model upgrade can shift behaviour overnight, and "correct" is often a judgement rather than an exact match. That doesn't make quality unmeasurable. It means the tests look different.

Build a golden set

Start with a set of real inputs — questions users actually ask, documents they actually upload — each paired with what a good answer must contain or must avoid. Keep it small enough to run on every change and representative enough to catch what matters. It will grow every time production surprises you.

Score what you can, judge what you can't

Some checks are mechanical: did the output parse as valid JSON, cite a source, stay under a length, avoid a forbidden phrase? Automate those first. For qualities like accuracy or tone, a second model can act as a grader, but only after you have checked its scores against human judgement on a sample — an unvalidated grader is just a confident opinion.

Treat prompts and models as code

  • Version prompts, model choices and settings alongside the code that uses them.
  • Run the eval suite in CI, and block releases that regress on the cases you care about.
  • Re-run it before adopting a new model version, however small the upgrade sounds.

Test for the adversarial

Users will try to talk a model out of its instructions, and documents can carry hidden instructions of their own. Include prompt-injection and jailbreak attempts in the suite, and design features so that a successful one can't reach anything sensitive.

Keep watching in production

Evals before release catch regressions you anticipated. Monitoring after release catches the ones you didn't: sample real outputs for review, track user corrections and thumbs-down, and feed the failures back into the golden set.

Our QA, Testing & Maintenance team builds these suites into delivery from the start, alongside the models our AI & Machine Learning practice puts into production.