What ships with every agent

The evaluation report, before anything touches a customer.

An agent is done when it clears a pass mark you set on your real past cases, with the misses explained. This example uses a fictional client, Halcyon, with figures we would scope for.

Stackjoinery
Evaluation report · support-agent v3 · Halcyon
Overall accuracy 96.2% Bar set by Halcyon: 95% · Pass
Evaluation set
412 real past tickets with known right answers, Jan to Jun 2026, sampled from each category
Labeled by
Two Halcyon support leads, disagreements resolved jointly
Model and tools
Claude, lookups in the help center and changelog, read-only tools for Stripe and the account database
Recommendation
Go, with the promotion path in section 5

1 What "correct" means here

Each ticket in the set has a labeled ideal outcome: a correct answer, or a correct escalation to a person. The agent is scored correct when its draft matches the label in substance. It is scored escalated correctly when it hands off a ticket that should be handed off. It is scored wrong when it answers incorrectly or escalates something it should have handled. Wrong answers count the same as wrong escalations. Halcyon chose that deliberately.

2 Results by category

CategoryTicketsCorrectEscalated correctlyWrongLaunch mode
How-to questions 171 98.2% 0.6% 1.2% Autonomous
Billing questions 104 96.2% 2.9% 0.9% Autonomous
Account changes 58 91.4% 8.6% 0.0% Supervised
Refund requests 41 100.0% 0.0% 0.0% Always gated
Bugs and outages 38 94.7% 5.3% 0.0% Supervised

3 The misses, and what changed

Six tickets were scored wrong. Three were variants of the same cause and are shown once below. We applied each fix and re-ran the set before issuing this report. The figures above are from the re-run.

  • #17204How-to
    Answered how to export reports from the old dashboard; the customer was on the new one.

    The lookup returned both articles and the model picked the older one. Fixed by adding the dashboard version to the lookup filter.

  • #17388Billing
    Quoted the annual price without the customer's negotiated discount.

    Discounts live on the Stripe customer record, which the tool could not read. We added the field.

  • #17511How-to
    Told the customer a feature was unavailable; it had shipped two weeks earlier.

    The help center lagged the release. We added the changelog to the lookup index and a freshness note to the agent instructions.

4 Cost and speed

Median time to draft
6.1 s
Time for 95% of drafts
14.8 s
AI model cost per ticket
$0.011
Projected monthly, 5,600 tickets
$62 AI model · $140 tools

5 Promotion path

  1. Shadow, 1 week. Drafts only. Nobody sees them but the team. This confirms the numbers hold on live traffic.
  2. Supervised, 2 weeks. Every reply waits for a one-click approval. We track the edit rate per category.
  3. Autonomous for how-to and billing. Starts when the approval rate stays above 97% for two weeks. Account changes and bugs stay supervised. Refunds stay gated indefinitely.
IssuedJune 28, 2026 · Stackjoinery
AcceptedJuly 1, 2026 · Head of Support, Halcyon
Re-runMonthly on a retainer. Otherwise, after any change to the instructions, tools, or model

Want to know the number before you commit?

Agent builds include this report and the evaluation set behind it. You decide the pass mark. We show you whether it clears.

esc
PagesHomeAutomation and AI agents for growing companiesPagesServicesAutomations, workflows, and agentsServicesAutomationsRule-based handoffs that remove copy-paste work between tools.ServicesComplex workflowsProcesses that span several tools, with approvals, exceptions, and clean data.ServicesAI agentsAgents that read, decide, and act across your tools. People stay in the loop.PagesSolutions by industryLogistics, healthcare, professional services, e-commerce, SaaSPagesWorkCase studies with the numbersCase studiesLumen LogisticsQuote requests from inbox to booked load in nine minutesCase studiesHalcyonA first-line support agent that resolves 61% of tickets on its ownCase studiesMeridian Dental GroupRecall reminders that cut no-shows by a third in 90 daysPagesPricingFixed prices and the instant estimatorPagesSample scopeThe one-page scope every client receivesPagesSample evaluation reportWhat ships with every agent buildPagesExample dashboardThe reporting every client getsPagesAboutA small senior teamPagesInsightsField notes from the buildInsightsWhen not to use an AI agentMost automation value comes from rule-based work. A guide to choosing rules, workflows, or agents for each process.InsightsDesigning approval gates that people useHuman review fails when people are asked too often or too late. Here are the patterns we ship, with the numbers behind them.InsightsThe true monthly cost of an automationAI model usage, tool seats, hosting, and the upkeep nobody budgets for. A worked example from a real quote pipeline.PagesContactBook a discovery call
↑↓ navigate↵ open⌘K or / to open anywhere