- Evaluation set
- 412 real past tickets with known right answers, Jan to Jun 2026, sampled from each category
- Labeled by
- Two Halcyon support leads, disagreements resolved jointly
- Model and tools
- Claude, lookups in the help center and changelog, read-only tools for Stripe and the account database
- Recommendation
- Go, with the promotion path in section 5
1 What "correct" means here
Each ticket in the set has a labeled ideal outcome: a correct answer, or a correct escalation to a person. The agent is scored correct when its draft matches the label in substance. It is scored escalated correctly when it hands off a ticket that should be handed off. It is scored wrong when it answers incorrectly or escalates something it should have handled. Wrong answers count the same as wrong escalations. Halcyon chose that deliberately.
2 Results by category
| Category | Tickets | Correct | Escalated correctly | Wrong | Launch mode |
|---|---|---|---|---|---|
| How-to questions | 171 | 98.2% | 0.6% | 1.2% | Autonomous |
| Billing questions | 104 | 96.2% | 2.9% | 0.9% | Autonomous |
| Account changes | 58 | 91.4% | 8.6% | 0.0% | Supervised |
| Refund requests | 41 | 100.0% | 0.0% | 0.0% | Always gated |
| Bugs and outages | 38 | 94.7% | 5.3% | 0.0% | Supervised |
3 The misses, and what changed
Six tickets were scored wrong. Three were variants of the same cause and are shown once below. We applied each fix and re-ran the set before issuing this report. The figures above are from the re-run.
- #17204How-toAnswered how to export reports from the old dashboard; the customer was on the new one.
The lookup returned both articles and the model picked the older one. Fixed by adding the dashboard version to the lookup filter.
- #17388BillingQuoted the annual price without the customer's negotiated discount.
Discounts live on the Stripe customer record, which the tool could not read. We added the field.
- #17511How-toTold the customer a feature was unavailable; it had shipped two weeks earlier.
The help center lagged the release. We added the changelog to the lookup index and a freshness note to the agent instructions.
4 Cost and speed
- Median time to draft
- 6.1 s
- Time for 95% of drafts
- 14.8 s
- AI model cost per ticket
- $0.011
- Projected monthly, 5,600 tickets
- $62 AI model · $140 tools
5 Promotion path
- Shadow, 1 week. Drafts only. Nobody sees them but the team. This confirms the numbers hold on live traffic.
- Supervised, 2 weeks. Every reply waits for a one-click approval. We track the edit rate per category.
- Autonomous for how-to and billing. Starts when the approval rate stays above 97% for two weeks. Account changes and bugs stay supervised. Refunds stay gated indefinitely.