Example engagement · B2B SaaS · Series B · AI agent

A first-line support agent that resolves 61% of tickets on its own

It answers from the help center, billing data, and past tickets. Refunds and account changes wait for a human. Every other answer cites its source.

61% of tickets resolved without a person
2.4 min first response, from 6 hours
4.8/5 customer satisfaction score on agent-handled tickets

The situation

Halcyon sells scheduling and billing software to service businesses. After a Series B the customer count roughly doubled in a year, and the support team did not. Nine people were handling about 1,400 tickets a week in Zendesk. First response had drifted to six hours. The backlog sat above 900 tickets most Monday mornings, and the team was triaging by which customers were loudest.

When we read a sample of the queue, the shape was clear. Roughly half the tickets were how-to questions that the help center already answered. A quarter were billing questions that needed one look at Stripe to resolve. The rest were account changes, refund requests, and real bugs, and those were the tickets a person needed to handle. The problem was that the first two groups were burying the third.

Priya's team had tested two off-the-shelf support bots. Both answered confidently and were wrong often enough that agents stopped trusting them. The requirement for us was blunt: show the accuracy on our tickets before the agent talks to a customer, and never let it touch money.

What we built

A first-line agent that works inside Zendesk. It sorts each new ticket, looks up the relevant material, and drafts a reply with citations. Then it does one of two things, depending on what the reply would commit Halcyon to.

Support agent workflow: ticket, classify, retrieve from help center and billing and past tickets, draft with citations, gate on irreversible actions, auto-send or human queue Help center Billing · Stripe Past tickets New ticketZendesk trigger Classifyintent · risk · language Retrievepgvector · top 8 chunks Draft replyClaude · every claim cited Irreversible?refund · plan · seats · data Reply sentwith citation · rated Human queueproposed action · one click yes no Every step, prompt, and retrieved chunk is traced in Langfuse and linked from the ticket.
The gate is a fixed policy list the support team owns. Anything that moves money, changes a plan, adds or removes seats, or deletes data goes to a person with the action pre-filled.

The agent looks things up in three sources indexed in pgvector: the help center, the last 18 months of resolved tickets, and a read-only view of the customer's billing state from Stripe. The draft must cite a passage for every factual claim. If the model cannot find support for an answer, it says so and routes the ticket to a person with what it did find. That is the behavior that made agents trust the queue again: a handoff arrives with the research already done.

Refunds, plan changes, seat changes, and data deletion never auto-send. The agent drafts the reply and the proposed Stripe or account action, and a support agent approves both with one click or edits them. Everything else is sent immediately with a citation link and a one-tap rating. Langfuse records every step, so when a customer disputes an answer, the team can see exactly which passage the agent relied on.

The rollout

Weeks one and two. Discovery, read-only access to Zendesk and Stripe, and the evaluation set: a collection of past tickets with known right answers. Two support leads pulled 400 resolved tickets across every category and graded each with the answer they would have wanted sent. That set became the pass mark for everything that followed.

Weeks three to five. Build and evaluate. On the first full run the agent matched or improved on the human resolution for 71% of the set. We fixed how help center articles are split up for search, added the past-tickets index, and tightened the citation rule. By the end of week five the number was 84%, and it correctly declined to answer on 97% of the tickets the leads had marked as needing a person.

Weeks six and seven. Shadow mode, where the agent drafts without sending. Support agents rated the drafts on every ticket, and customers saw none of them. We set confidence thresholds from those ratings.

Weeks eight and nine. Auto-send for how-to and billing questions, then the full policy. A written runbook covered the daily review, how to add a help center article to the index, and how to change the gate list.

Results

MeasureBeforeAfter 90 days
Tickets resolved without a person0%61%
First response time (median)6 hrs2.4 min
Customer satisfaction on agent-handled ticketsn/a4.8 / 5
Customer satisfaction on human-handled tickets4.3 / 54.6 / 5
Monday morning backlog900+under 120
Escalation accuracy on held-back test setn/a97%

The human satisfaction score rose because people now spend their day on the tickets that need judgment, and reach them in minutes instead of hours. No one on the support team was let go. Two agents moved into a new customer onboarding role that Halcyon had been unable to staff.

What's next

Halcyon is on a monthly retainer. The gate list has been widened once, to let the agent apply a documented goodwill credit under a fixed amount, after the team watched 300 proposed credits and approved all but four. Next up is a proactive flow that opens a ticket when a card payment fails, drafts the outreach, and gives the customer a link to update their card before the account is paused.

Tell us what your team still does by hand.

One paragraph is enough. We reply within a business day with questions or a time for a 45-minute call.

esc
PagesHomeAutomation and AI agents for growing companiesPagesServicesAutomations, workflows, and agentsServicesAutomationsRule-based handoffs that remove copy-paste work between tools.ServicesComplex workflowsProcesses that span several tools, with approvals, exceptions, and clean data.ServicesAI agentsAgents that read, decide, and act across your tools. People stay in the loop.PagesSolutions by industryLogistics, healthcare, professional services, e-commerce, SaaSPagesWorkCase studies with the numbersCase studiesLumen LogisticsQuote requests from inbox to booked load in nine minutesCase studiesHalcyonA first-line support agent that resolves 61% of tickets on its ownCase studiesMeridian Dental GroupRecall reminders that cut no-shows by a third in 90 daysPagesPricingFixed prices and the instant estimatorPagesSample scopeThe one-page scope every client receivesPagesSample evaluation reportWhat ships with every agent buildPagesExample dashboardThe reporting every client getsPagesAboutA small senior teamPagesInsightsField notes from the buildInsightsWhen not to use an AI agentMost automation value comes from rule-based work. A guide to choosing rules, workflows, or agents for each process.InsightsDesigning approval gates that people useHuman review fails when people are asked too often or too late. Here are the patterns we ship, with the numbers behind them.InsightsThe true monthly cost of an automationAI model usage, tool seats, hosting, and the upkeep nobody budgets for. A worked example from a real quote pipeline.PagesContactBook a discovery call
↑↓ navigate↵ open⌘K or / to open anywhere