The situation
Halcyon sells scheduling and billing software to service businesses. After a Series B the customer count roughly doubled in a year, and the support team did not. Nine people were handling about 1,400 tickets a week in Zendesk. First response had drifted to six hours. The backlog sat above 900 tickets most Monday mornings, and the team was triaging by which customers were loudest.
When we read a sample of the queue, the shape was clear. Roughly half the tickets were how-to questions that the help center already answered. A quarter were billing questions that needed one look at Stripe to resolve. The rest were account changes, refund requests, and real bugs, and those were the tickets a person needed to handle. The problem was that the first two groups were burying the third.
Priya's team had tested two off-the-shelf support bots. Both answered confidently and were wrong often enough that agents stopped trusting them. The requirement for us was blunt: show the accuracy on our tickets before the agent talks to a customer, and never let it touch money.
What we built
A first-line agent that works inside Zendesk. It sorts each new ticket, looks up the relevant material, and drafts a reply with citations. Then it does one of two things, depending on what the reply would commit Halcyon to.
The agent looks things up in three sources indexed in pgvector: the help center, the last 18 months of resolved tickets, and a read-only view of the customer's billing state from Stripe. The draft must cite a passage for every factual claim. If the model cannot find support for an answer, it says so and routes the ticket to a person with what it did find. That is the behavior that made agents trust the queue again: a handoff arrives with the research already done.
Refunds, plan changes, seat changes, and data deletion never auto-send. The agent drafts the reply and the proposed Stripe or account action, and a support agent approves both with one click or edits them. Everything else is sent immediately with a citation link and a one-tap rating. Langfuse records every step, so when a customer disputes an answer, the team can see exactly which passage the agent relied on.
The rollout
Weeks one and two. Discovery, read-only access to Zendesk and Stripe, and the evaluation set: a collection of past tickets with known right answers. Two support leads pulled 400 resolved tickets across every category and graded each with the answer they would have wanted sent. That set became the pass mark for everything that followed.
Weeks three to five. Build and evaluate. On the first full run the agent matched or improved on the human resolution for 71% of the set. We fixed how help center articles are split up for search, added the past-tickets index, and tightened the citation rule. By the end of week five the number was 84%, and it correctly declined to answer on 97% of the tickets the leads had marked as needing a person.
Weeks six and seven. Shadow mode, where the agent drafts without sending. Support agents rated the drafts on every ticket, and customers saw none of them. We set confidence thresholds from those ratings.
Weeks eight and nine. Auto-send for how-to and billing questions, then the full policy. A written runbook covered the daily review, how to add a help center article to the index, and how to change the gate list.
Results
| Measure | Before | After 90 days |
|---|---|---|
| Tickets resolved without a person | 0% | 61% |
| First response time (median) | 6 hrs | 2.4 min |
| Customer satisfaction on agent-handled tickets | n/a | 4.8 / 5 |
| Customer satisfaction on human-handled tickets | 4.3 / 5 | 4.6 / 5 |
| Monday morning backlog | 900+ | under 120 |
| Escalation accuracy on held-back test set | n/a | 97% |
The human satisfaction score rose because people now spend their day on the tickets that need judgment, and reach them in minutes instead of hours. No one on the support team was let go. Two agents moved into a new customer onboarding role that Halcyon had been unable to staff.
What's next
Halcyon is on a monthly retainer. The gate list has been widened once, to let the agent apply a documented goodwill credit under a fixed amount, after the team watched 300 proposed credits and approved all but four. Next up is a proactive flow that opens a ticket when a card payment fails, drafts the outreach, and gives the customer a link to update their card before the account is paused.