Case studyEvaluating an agent with no answer key

Evaluation for agents that don't fit the template.

We help teams find out whether their AI agents actually work: by customizing proven evaluation frameworks, and by bringing methods from the latest research into production before any product offers them.

trace · it-support-agent3 judges · 2-of-3 vote
user"I lost VPN access after my password reset."
01toolget_employee_profile(email)GEA
02toolsearch_tickets(user, "VPN")GEA
03toolsearch_tickets(user, "VPN")redundantGEA
04reply"Your ticket INC-4821 is resolved."ungroundedGEA
G groundednessE efficiencyA goal alignmentillustrative

Services

Two ways we help

Framework implementation

Make the evaluation stack you already use fit your agent

Off-the-shelf metrics rarely match what matters for a multi-step, tool-using agent. We tailor existing frameworks to your system instead of starting from scratch.

  • Custom metrics for your domain and workflows
  • Evaluation datasets built from production traces
  • Regression gates in CI before every release

Research to production

Methods from recent papers, working on your agent

The most useful evaluation ideas are in papers, not products. We adapt them to how your agent actually behaves, and fix the parts that break on real traffic.

  • Step-level trajectory evaluation, not just final answers
  • Reference-free scoring when there is no answer key
  • Judge ensembles with agreement reporting

The method

From raw traces to a decision you can defend.

  1. 01

    Traces

    Real conversations, tool calls, and results, exactly as your agent ran them.

  2. 02

    Evidence

    For every step, what the agent knew at that point: the user's goal and earlier tool results.

  3. 03

    Judges

    Independent LLM judges from different providers score each step on each metric.

  4. 04

    Agreement

    Majority vote, with agreement reported per metric, so you know which numbers to trust.

  5. 05

    Decision

    Findings tied to the choice you face: ship, switch models, or cut cost.

Case study

Evaluating a production IT support agent without an answer key

Adapting recent trajectory-evaluation research to a real tool-calling agent, then using it to compare model configurations on quality, cost, and latency.

4

step-level quality metrics, with no answer key

3

independent LLM judges from different providers, majority vote

~9×

cost gap between model configurations, weighed against quality

Read the full case study

Working together

How an engagement works

  1. Scope

    We learn what your agent does, what data you have, and which decision the evaluation needs to support: a release, a model switch, a cost cut.

  2. Evaluate

    We build or adapt the evaluation and run it on your real traffic, reporting how far each number can be trusted.

  3. Decide

    You get findings, a clear recommendation, and a pipeline your team can keep running after we leave.

Vendor-neutral by design. We don't sell a platform. Recommendations follow the evidence, including when the answer is that your current setup is good enough.

Have an agent that needs evaluating?

Tell us what it does and what decision you're facing. We'll reply with how we would approach it.

Get in touch