Log and Telemetry Triage at Scale
The 40MB log that breaks the baseline
Picture this: 2:47 AM on a Tuesday. Checkout is degraded. The only artifact you have is 40MB of JSONL error records covering the last six hours.
The task is simple: triage. What broke, how often, what’s the root cause, and what’s actually severe?
The setup: I give the agent the task. I give it a yardstick: every distinct error class in the file has to appear in the report, with counts, rates, and a defensible root-cause narrative. Then I get out of the way. No workflow design. No tool list. The agent figures out what to build.
So the agent does what seems obvious first. It reads the log by hand, in context.
For the first few thousand lines, the pattern is clear. Payment timeout, payment timeout, payment timeout. A retry storm from one flaky upstream. The narrative is obvious: one provider degrading, everything else is noise. The agent feels confident.
Then the context window fills. To keep reading, the agent summarizes. Summarize, read, summarize, read. With each pass, the earlier details compress. Details on early pages become summaries. Summaries become mentions. Mentions get dropped.
At line 300,000 there is an error class that appears 41 times. SignatureVerificationFailed from the webhook receiver. All 41 in a 90-second window. All on a single tenant. This is the actual root cause. The payment timeouts are downstream noise.
The agent read it, compressed it into “a handful of webhook errors,” and moved on.
The report came out confident, coherent, and wrong about what mattered.
Hand-reading a 40MB file is a lossy compression algorithm. The volume is bigger than the context window. Details get compressed into noise. Noise gets dropped. By the end, the agent has a confident narrative and no idea that it’s wrong.
The baseline is broken.
Round 0: the agent discovers the baseline is broken
No custom tools. No workflow. Just the agent and the raw data. The agent reads the log, reasons over it, writes a report.
The run journal is spotless. One long pass through the data. Every transition fires once. No re-entries. No stalls. The process ran perfectly.
And that’s the trap. A clean journal is not a good outcome. The journal records how the process ran. It says nothing about whether the output is right. A workflow with no quality gate cannot fail, so the journal stays clean while the output breaks silently.
Round 0 did produce an answer. The agent followed its steps: read, reason, report. The output missed the critical error at line 300,000. The journal shows no trace of that. The process and the output are measuring different things.
The agent is about to discover both problems.
Round 1: the measurements reveal what to build
The agent measures Round 0. Three signals. Numbers before opinions.
| Signal | What it reveals | Tool |
|---|---|---|
| Efficiency | How did the process run? | getFsmRun |
| Completeness | Is the output actually complete? | Deterministic recompute check |
| Quality | Was the output any good? | evalById (LLM critic) |
The measurements for Round 0:
- Tokens: 612,400. The raw log sat in context, repeatedly, as the agent re-read and re-summarized.
- Completeness: 60 percent. The recompute check found 20 distinct error classes in the file. The report only named 12. Eight classes are missing, including the
SignatureVerificationFailedat line 300,000. - Critic score: 8.6 out of 10. “Clear structure, confident tone, good use of examples.”
The critic’s score is high. The completeness number is not. Both are true. The critic is judging the narrative quality. The completeness check is saying “you missed 8 error classes.” The narrative sounds great and is incomplete.
The token number also names a gap: 612,400 tokens to hand-read a file. A tool could process the same volume for a fraction of the cost. Two problems: the output is incomplete, and the process is expensive.
The agent now figures out what to build.
Round 2: the agent discovers it needs a tool
The measurements revealed the gap. The agent now realizes what needs to be built: a tool that can process the 40MB log outside the context window.
The agent builds log_triage.
Here is the shape of the run:
The tool does four things, in order:
- Parse and normalize. Each JSONL line becomes a record with a timestamp, a service, an error class, and a message.
- Dedupe by error signature. A signature is the error class plus a normalized message template: variable parts (ids, numbers, timestamps, paths) are stripped. Forty-one
SignatureVerificationFailedlines with different tenant ids collapse into one signature with a count of 41. - Aggregate. Counts per signature, per service, and per time bucket. Rates per minute. A trend flag for anything that is rising or falling across the window.
- Summarize. The whole volume becomes a compact JSON summary: a few kilobytes, not forty megabytes.
The shape of the tool, in pseudocode:
def log_triage(path):
records = [parse(line) for line in open(path)]
sigs = {}
for r in records:
sig = signature(r) # class + normalized template
sigs[sig] = sigs.get(sig, 0) + 1
return summarize(sigs, records) # counts, rates, trends
That is the whole trick. The tool processes the volume outside the context window. The agent’s context now holds the aggregated summary, not the raw log. It can read the entire summary at once, every time, with no lossy compression.
The agent runs the tool and measures:
- Tokens: 24,800. Down from 612,400. The raw log never enters context.
- Completeness: 100 percent. All 20 error classes are in the summary.
- Critic score: 6.1 out of 10. “Accurate counts. But it is a list, not an explanation.”
The tool fixed the completeness gap. The critic score dropped. The tool only counts. It cannot explain. A table of 20 signatures with counts is complete and accurate. It is not a triage. It does not say that webhook failures are the root cause. It does not judge severity. It does not produce a narrative a human can act on.
The tool cannot do the judgment part. The agent needs something that puts the tool in its lane and an LLM in its lane. The tool crushes the volume. The LLM interprets. A gate between them checks that the interpretation is correct.
The agent has only solved half the problem.
Round 3: the agent discovers it needs a workflow with a gate
The agent now builds what Round 2 revealed was missing: a workflow that can interpret the data and check its own work.
The agent designs a finite state machine. The tool in one lane, the LLM in another, a gate between them.
The shape of it:
Five states, one re-entrant loop.
triageruns thelog_triagetool. Deterministic. Produces the compact summary.root_causeis where the LLM works. It reasons over the aggregated signatures, correlates across services, and writes a draft narrative: what broke, why, what is severe.verifyis the gate. A deterministic recompute tool. It recomputes the set of error classes from the raw file and diffs it against the draft. It returnsapprovedorrejected: <reasons>. It is a checkable fact, not an opinion.reportwrites the final narrative once the gate has passed.escalateis the ceiling. If the gate rejects three times, the run stops and reports the ceiling instead of looping.
The gate is what makes the workflow able to fail. In Round 0, the journal was clean because nothing could reject the output. Now, something can. If the gate rejects, the work goes back and the LLM re-reasons with the gap explicitly named. The re-entry is the workflow doing what it should.
Here is the FSM:
fsm:
name: log-triage
version: 1
initialState: triage
context:
summary: { type: string }
draft: { type: string }
verdict: { type: string }
rejections: { type: list, of: string }
report: { type: string }
agents:
worker:
canTransition: [triage, root_cause, verify, report]
guidance: full
states:
triage:
description: >-
Run the log_triage tool over the raw file. Set summary to the compact
JSON summary it returns.
requires:
- { name: summary, from: agent }
transitions:
- { toState: root_cause }
root_cause:
description: >-
Reason over summary and write a draft triage narrative in draft.
Account for every signature in summary, name the likely root cause,
and judge severity. If rejections is non-empty, address each named
gap before rewriting.
requires:
- { name: draft, from: agent }
transitions:
- { toState: verify }
verify:
description: >-
Run the verify tool. It recomputes the distinct error classes from
the raw file and diffs them against draft. Set verdict to its result
verbatim, approved or rejected with reasons.
requires:
- { name: verdict, from: agent }
transitions:
- { toState: report,
guards: [{ backend: cel, expression: "context.verdict == 'approved'" }] }
- { toState: root_cause,
guards: [{ backend: cel, expression: "size(context.rejections) < 3" }],
onTransition:
- { append: { rejections: "rejected" } }
- { assign: { draft: "", verdict: "" } } }
- { toState: escalate,
guards: [{ backend: cel, expression: "size(context.rejections) >= 3" }] }
report:
description: >-
Write the final triage report in report from the approved draft.
requires:
- { name: report, from: agent }
transitions:
- { toState: done }
done:
description: "Triage complete and verified."
terminal: true
escalate:
description: "Gate rejected three times. Report the ceiling."
terminal: true
policy:
maxIterations: 8
maxExchangesPerStage: 10
A few of the DSL rules are doing real work in that file, and they are not decoration.
Typed context, no defaults. Every field is declared with a type. There is no default: key anywhere. A declared field that has not been written is simply absent; a declared list binds as an empty collection in a guard. That is why rejections can be compared with size() from the very first pass.
Literal-only assign. The assign writes literals only. On the reject edge it blanks draft and verdict. It cannot increment a counter or copy one field into another. The loop-bound list grows with a separate append op, which also takes a literal. So the bound is not a number that gets incremented. It is the list itself, one entry per rejected pass.
Loop-bounding via the list. rejections appends one entry per rejected pass. The counter is the data. The guard size(context.rejections) < 3 bounds the loop, and escalate is the declared terminal for when the bound is hit. No hidden counter, no arithmetic.
Consume-the-trigger. root_cause is re-entrant, and its requirement is draft. When verify rejects, the run takes the verify to root_cause edge, and that edge’s onTransition blanks draft (and verdict). Without it, the stale draft from the previous pass would still be present, root_cause would treat its requirement as already satisfied, and it would fall straight through without anyone re-reasoning. Blanking the trigger forces a fresh draft on every re-entry.
The re-entrant trap. A state’s requires is satisfied when the named field is present. A re-entrant state whose requirement is already set will not re-collect. So the guard that decides the loop must compare (size(context.rejections) < 3), not test presence. This is the trap that silently stalls a naive FSM.
canTransition lists every agent-driven state. Under the worker agent it lists triage, root_cause, verify, and report, including the re-entrant root_cause. Omit one and the run stalls permanently, because the agent is not allowed to act in a state it is not listed in.
The agent measures Round 3:
- Tokens: 26,400. The context holds the summary and the draft.
- Completeness: 100 percent. The gate guarantees it. A run cannot reach
reportunless the gate approves. - Re-entries: 1. On the first pass, the draft dropped a low-count class. The gate caught it, named the class, sent the work back. The LLM fixed it on the second pass. The loop fired once because something was actually wrong. Re-entries mean the gate is working.
- Critic score: 8.9 out of 10. The narrative now explains, not just counts.
The gate caught a real gap and fixed it. The gate is not friction. It is the workflow finally being able to fail in a way that fixes itself.
Round 4: the agent discovers the optimum by tuning one change at a time
The workflow works. The agent now discovers how to make it better. One change at a time. Measure. If the metric moves, the agent keeps it. If it doesn’t, the agent does not keep tuning for no gain.
Change one: Tighten the dedupe signature. The agent measures and realizes the signature is too coarse. Two different errors collapsed into one. The tool’s output is complete, but the narrative is muddied. The agent refines the signature to keep the error class. Re-run and measure: signature count went from 18 to 20. The critic noted the two connection errors are now distinct. One number moved.
Change two: Add a severity threshold. The agent measures and realizes the narrative buries the lede. All 20 signatures are treated equally. The 41-count webhook failure gets the same space as a 2-count warning. The agent adds a severity flag to the tool and tells the LLM to lead with the severe ones. Re-run and measure: critic score went from 8.9 to 9.3. The narrative now opens with the webhook failure. One number moved.
Change three: Nothing. The agent looks at the reject loop bound (set to 3). The measurements show it only fired once. The agent checks: does this need tuning? No. It leaves the bound at 3. No change, no re-run. This is discovery through discipline: the measurements do not show a problem, so the agent does not build a solution.
Convergence. The agent runs one more time to confirm the numbers are stable. Tokens: 26,400. Completeness: 100 percent. Re-entries: 1. Critic: 9.3. Nothing moved from the last round. Three consecutive runs with no improvement. The agent discovers it has reached the ceiling and reports it. The optimum is not a target the agent was told to hit. It is the point where the agent’s own measurements stopped moving.
The optimum: what the agent discovered
Here is the arc from baseline to optimum:
| Metric | Round 0 (baseline) | Optimum | Discovery |
|---|---|---|---|
| Tokens | 612,400 | 26,400 | The agent discovered it needs a tool that reads the volume outside context |
| Completeness | 60 percent | 100 percent | The agent discovered it needs a gate that recomputes and diffs |
| Re-entries | 0 (no gate) | 1 (gate caught a real gap) | The agent discovered re-entries are the feature, not a bug |
| Critic score | 8.6 | 9.3 | The agent discovered severity thresholds guide the narrative |
The most important difference is in the re-entries row. In Round 0, zero re-entries looked good. Nothing could fail. In the optimum, one re-entry means the gate found a real gap, named it, sent the work back, and the LLM fixed it. The re-entry is the agent learning that it was wrong and correcting itself.
The signal that told the agent it had converged was the token count. Across three consecutive tuning rounds: 26,400, 26,400, 26,400. Flat. The cost of a complete, correct answer stopped improving.
No human designed the workflow, the tools, the gate, or the tuning strategy. The agent did.
The human gave the task: triage the log. The human gave the yardstick: every error class must appear with counts and a defensible narrative. Then the human stepped back.
The agent built log_triage because Round 1 revealed incompleteness. The agent built the FSM because Round 2 revealed the tool could not interpret. The agent added the gate because Round 3 revealed the workflow needed to be able to reject its own output and retry. The agent tuned the signature and added the severity threshold because Round 4 revealed the narrative could be clearer.
Every single decision was driven by a measurement. Every measurement revealed what needed to be built next. The agent built it. Measured it. Discovered whether it worked. Stopped when the metrics stopped moving.
The optimum is not what I would have designed. It is what the agent discovered by iterating on its own feedback. The agent read past the critical error at line 300,000 in Round 0. In the optimum, the tool counts it, the gate guarantees it appears, and the narrative leads with it. The agent discovered what it took to never miss it again.