← All posts

Evaluation · 8 min read

A practical eval stack for LLM apps you can set up in a day

You don't need a platform to start evaluating your LLM app. A spreadsheet of real cases, a few scoring methods and a script you run on every change will get you most of the way.

Subhash Nunna ·

Teams often delay evaluation because they think it needs a dedicated platform. It doesn’t. Here’s a stack you can build in a day that will catch most regressions.

Step 1: Build a golden set

Gather 50–100 real inputs your system should handle. Pull them from logs, support tickets or user interviews — not from your imagination. For each one, record:

  • the input
  • what a good answer must contain (key facts, required fields, tone)
  • anything it must not contain

Keep it in a spreadsheet or a JSONL file in your repo. Version it like code.

Step 2: Choose a scoring method per case

Different cases need different checks:

  • Exact or structural checks — does the JSON parse, are required fields present, is the number in range? Cheap, fast and reliable.
  • Reference checks — does the answer include the key facts from the reference? Simple string or semantic matching works for many cases.
  • Model-graded checks — a separate model scores the answer against a written rubric. Useful for tone, helpfulness and reasoning, but calibrate it against human judgement first.

Use the cheapest method that’s trustworthy for each case.

Step 3: Automate the run

Write a script that runs every case through your app, scores it and writes a report: overall pass rate, pass rate by category, and the failing examples side by side with the expected result. Run it in CI on every pull request that touches prompts, models or retrieval.

Step 4: Review failures weekly

The score is a signal; the failures are the insight. Every week, read the failing cases, fix what’s fixable, and add new cases from real-world problems your users hit. Your golden set should grow as your product does.

What this gets you

With this in place, you can change models, prompts or retrieval with confidence, because you’ll see the impact before your users do. That alone puts you ahead of most teams shipping AI today.