Support Ticket Mining at Scale
The task
5,000 free-text support tickets. The task is simple to state but hard to do at scale. Classify each ticket. Extract the structured fields that matter: product, issue type, severity, whether a refund was requested. Aggregate into trends a human can act on.
The tickets are messy. “why is my card chrgd twice???” sits next to a three-paragraph saga about a failed migration. Neither has a neat label.
The setup: I give the agent the task. I give the yardstick: run journal, deterministic completeness check, critic score. Then I get out of the way. No workflow design. No tool list. The agent figures out what to build.
Here is the key insight: the order of the work is reversed from Parts 1 and 2. In those parts, the deterministic tool came first, then the LLM interpreted. Here, free text cannot be aggregated until it has been understood. The LLM has to classify and extract first. Only then does the deterministic tool have something to aggregate.
The agent is about to discover this constraint, and it will shape every round that follows.
Round 0: the agent discovers the baseline breaks at scale
No custom tools. No workflow. Just the agent and the tickets.
The agent reads them one at a time. Classify, note, move on. It works at first. Then the context fills. The agent scrolls back to re-check earlier labels. It re-reads the current ticket to remember where it was. The journal shows re-entries - the agent scrolling and re-reading.
By ticket 3,000, the agent has lost the running tally. It re-tallies from scratch. The new tally does not match the old one.
Worse: the classification drifts. The same kind of ticket, a customer charged twice asking for a refund, is ‘billing’ in one pass and ‘refund’ in another. By ticket 3,000 the context is full. The label it reaches for is whatever was closest to hand.
The journal shows all the re-entries and re-reads. The journal looks busy. The output is inconsistent.
The agent just discovered: hand-classifying 5,000 tickets at scale does not work. The classification drifts. The tallies disagree. The output is not trustworthy.
Round 1: the measurements reveal what needs to be built
The agent measures Round 0. Three signals, facts first.
- Tokens: 1.9M. The agent re-read the same tickets over and over. Every re-read paid the full token price.
- Re-entries: 214. The agent scrolling back to re-check its own earlier classifications. That is rework. The journal’s way of saying the agent did not trust its own first pass.
- Completeness: 4,612 of 5,000 tickets. The deterministic check finds 388 tickets that were never classified. They fell off the back of the context window.
- Consistency: 96 pairs of the same issue type classified two different ways. ‘billing’ vs ‘refund’ for duplicate charges. ‘bug’ vs ‘feature request’ for the same missing button.
- Critic score: thorough and well-organized. The critic thought the work was good.
The critic’s opinion and the deterministic facts are pointing opposite directions. The consistency check says 96 pairs are labeled differently. The critic says the work is thorough.
The variable part of the job is unconstrained, so it drifts. The volume part is done in context, so it does not scale. The LLM is doing two jobs at once: understanding text and keeping the books. It is bad at keeping the books.
The agent now knows what to build.
Round 2: the agent discovers it needs a schema and a tool
The problem is not the LLM being careless. The LLM’s output is prose - a classification note, a tally, a report. Prose is for humans. Nothing downstream can consume it. The agent has to hold the whole state in context, and that is where the drift and re-reading come from.
The solution: make the variable extraction produce structured output, constrained to a schema. Then a deterministic tool can aggregate it.
So the agent defines a schema. Small, on purpose:
{
"ticket_id": "string",
"product": "string",
"issue_type": "enum: billing | refund | bug | feature_request | account | other",
"severity": "enum: low | medium | high",
"refund_requested": "boolean",
"summary": "string"
}
Every ticket the LLM reads comes back as one record of exactly that shape. No free-form label. No ‘billing-ish’. The enum is the constraint: if the ticket is a duplicate-charge complaint, the issue type is either ‘billing’ or ‘refund’, and the schema forces the agent to pick one, the same way, every time. The variable part is still variable. The LLM still has to understand the messy prose. But its output is now a record a machine can hold, count, and diff.
With the schema in place, the agent builds the deterministic core: a ticket_agg tool. It takes the structured records and does the volume work: counts by issue type, by product, by severity; dedupes tickets that are the same complaint filed twice; computes the metrics. It operates on the schema, not on the raw text. The raw tickets never enter the tool’s context. They never had to.
Here is the architecture once both halves exist. The LLM is in its lane, doing the variable work on raw text. The tool is in its lane, doing the deterministic work on structured records. The boundary between them is the schema.
Notice what the tool cannot do. It can count ‘refund’ tickets. It cannot tell you that ticket 2,140 is a refund request buried in a rant about a failed migration. That judgment is still the LLM’s. The tool only computes; it does not interpret. That split, the deterministic core in its lane and the variable layer in its, is the same principle as Parts 1 and 2, just with the order reversed. It is the deterministic shell, probabilistic core split from the Agents in Production series, except here the shell is something the agent builds and tunes itself. The LLM produces the structured input, and the tool produces the metrics. Each half is done by the thing that is best at it.
Token cost drops immediately, because the context now holds the schema and the records, not the raw corpus. But the run is still not gated. The agent can still emit a malformed record, or a record that is missing, and nothing would catch it. That is Round 3.
Round 3: the variable layer and the gate
The schema and the tool are in place, but the run is still not gated. The agent can emit a malformed record, or a record that is missing, and nothing would catch it. So the agent builds the finite state machine. It is the workflow that puts each half in its lane and checks the boundary between them.
The FSM has three working states and two bookends. extract is the variable layer: the LLM reads a batch of tickets and produces structured records. verify is the deterministic gate: it checks that the records are schema-valid and complete for the batch. report is the deterministic core: it runs ticket_agg over all the records to produce the final metrics. done is the terminal state. escalate is the escalation state, for a batch that fails verification too many times.
The re-entrant loop is the quality gate. If verify rejects a batch, the work is sent back to extract for re-extraction. This is the state that makes the workflow able to fail. Without it, the journal is clean and the output is not. With it, a bad batch is caught and reworked, and the journal records the rework as re-entries.
Here is the FSM, written in the Aegis DSL. It is the final version, the one that converged in Round 4.
fsm:
name: ticket-mining
version: 4
initialState: extract
context:
batch: { type: string }
passes: { type: list, of: string }
drafts: { type: string }
verdict: { type: string }
report: { type: string }
agents:
worker:
canTransition: [extract]
guidance: full
states:
extract:
description: >-
Read the next batch of tickets and extract one structured record per
ticket to the schema. Set `batch` to the extracted records for this
pass and append a pass marker to `passes`.
requires:
- { name: batch, from: agent }
transitions:
- { toState: verify }
verify:
description: >-
Deterministic gate. Check the batch's records are schema-valid and
complete against the raw tickets. Set `verdict` to the tool's result
exactly: approved, or rejected with reasons.
requires:
- { name: verdict, from: { tool: ticket_verify } }
transitions:
- toState: extract # re-enter: batch rejected, re-extract
guards: [{ backend: cel, expression: "context.verdict.startsWith('rejected') && size(context.passes) < 10" }]
onTransition:
- { assign: { batch: "", verdict: "" } } # consume the trigger so extract re-collects a fresh batch
- toState: report
guards: [{ backend: cel, expression: "context.verdict == 'approved'" }]
- toState: escalate
guards: [{ backend: cel, expression: "context.verdict.startsWith('rejected') && size(context.passes) >= 10" }]
report:
description: "Run ticket_agg over all verified drafts to produce the final metrics report."
requires:
- { name: report, from: { tool: ticket_agg } }
transitions:
- { toState: done }
done: { description: "Report complete.", terminal: true }
escalate: { description: "A batch failed verification too many times. A human takes over with the pass ledger.", terminal: true }
policy:
maxExchangesPerStage: 20
maxExchangesPerAgentPerStage: 10
checkpointEvery: 2
A few things to notice about the YAML, because they are the DSL rules in action. The context fields are typed, and there is no default: key; a declared-but-unwritten field is simply absent. The re-entrant edge back into extract carries the assign, and the assign is literal-only: it writes batch and verdict to empty strings verbatim. That is the consume-the-trigger idiom. extract’s trigger is batch. A state’s requires is satisfied the moment the named field is present, so a re-entrant state whose requirement is already set will not re-collect. By blanking batch on the edge coming back in, the next pass through extract re-collects a fresh batch instead of routing on the stale one. The loop is bounded by a list field, passes, that appends one entry per pass. The guard size(context.passes) < 10 is the counter, and it is a comparison, not a presence test, which is exactly what the re-entrant trap requires. canTransition lists extract, the only agent-driven state, and it is the re-entrant one. Omit it and the run stalls permanently. And the loop has an exit: when a batch is rejected and passes has already hit the limit, the guard routes to escalate instead of re-entering extract. That is the declared escalation state, and it is reachable.
The verify state is where the gate lives. It is tool-driven, not agent-driven: its requires names { tool: ticket_verify } as the source, not agent. It runs ticket_verify, a deterministic tool that recomputes the expected output from the raw input and diffs it. It returns approved or rejected: <reasons>. That is a checkable fact, not an opinion. And it gates the FSM: a rejected batch is sent back, and the loop only fires when something is actually wrong.
That is the whole architecture. The LLM is in its lane, doing the variable work on raw text, constrained to a schema. The tool is in its lane, doing the deterministic work on structured records. The gate checks the boundary between them. And the loop is bounded, so it converges instead of spinning.
Round 4: the agent discovers the optimum by tuning
The workflow works. The agent now discovers how to make it better. One change at a time. Measure. If the metric moves, keep it. If not, do not keep tuning for no gain.
Change one: batch size. The agent measures and realizes the first version extracts one ticket per pass. That is 5,000 passes. The agent batches 500 tickets per pass instead. Ten passes, not 5,000. Re-run and measure: tokens drop from 1.9M to 410K. Re-entries drop from 214 to 9. One number moved: tokens down 78 percent.
Change two: schema fields. The agent measures and realizes the summary field is 140 characters on average - the LLM is hedging. The agent trims it to 40 and adds a confidence enum. The gate now rejects only when confidence is low and the issue type is ambiguous. Re-run and measure: re-entries drop from 9 to 3. The consistency check finds 0 pairs classified two different ways. The drift is gone. One number moved.
Change three: the verify recompute. The agent tightens the verify check to confirm every ticket has exactly one record. Re-run and measure: completeness jumps to 5,000. All tickets processed. One number moved.
Convergence. The agent runs three more times. The metrics stay flat. The agent discovered it hit the ceiling and reported it.
The optimum: what the agent discovered
Here is the baseline against the converged optimum:
| Metric | Round 0 (baseline) | Optimum | Discovery |
|---|---|---|---|
| Tokens | 1.9M | 410K | The agent discovered batch extraction reduces token cost |
| Re-entries | 214 | 3 | The agent discovered the gate works when the schema is right |
| Tickets processed | 4,612 of 5,000 | 5,000 of 5,000 | The agent discovered the verify check must be strict |
| Inconsistent pairs | 96 | 0 | The agent discovered the confidence enum eliminates drift |
| Critic score | ”thorough" | "thorough” | The agent discovered the critic was not the signal |
The critic said “thorough” both times. The critic was the weakest signal and it did not move. The numbers that mattered were the facts, and they moved a lot. Tokens down 78 percent. Re-entries down 98 percent. Inconsistency down to zero. Completeness up to 100 percent.
The signal that told the agent it had converged was the consistency check hitting zero. The inconsistent pairs stayed at zero for two consecutive runs. The metrics stopped moving.
And here is the headline:
No human designed the workflow, the tools, the schema, the batch size, or the tuning strategy. The agent did.
The human gave the task: classify 5,000 tickets. The human gave the yardstick: run journal, completeness, consistency. Then the human stepped back.
The agent built the schema because Round 1 revealed the LLM’s output had to be constrained to avoid drift. The agent built the ticket_agg tool because Round 2 revealed the structured records needed to be aggregated deterministically. The agent built the FSM and gate because Round 3 revealed the workflow needed to verify its own work. The agent tuned the batch size, the schema fields, and the verify logic because Round 4 revealed each one could be improved.
Every decision was driven by a measurement. Every measurement revealed what needed to be built. The agent built it. Measured. Stopped when the metrics stopped moving.
The optimum is what the agent discovered by iterating, not what a human would have designed upfront.
A closing note
I keep coming back to the Round 0 journal. It was clean: linear, no stalls, no rejections. If the only thing I had looked at was the journal, I would have shipped the 4,612-ticket report and called it a day. The difference between that run and the optimum is a gate. A workflow that cannot fail cannot tell you when it has failed, so the agent that builds its own workflow has to build the gate too.
That is the lesson this run keeps returning to. Hand the agent a problem and a yardstick, and it will build the tools and the state machine to solve it, measure itself against the yardstick, and revise until the numbers stop moving. The human’s part is the task, the yardstick, and the approval of the dangerous operations. Everything else is the agent’s.
Part 4 flips the problem again: instead of one corpus to classify, several overlapping record sets to dedupe and resolve, and the gate has to catch the false merge. The loop is the same. The order is not.