Data Quality Audits over Messy Data
The agent had the file open. Two million rows, 48 columns, about 1.8 gigabytes of CSV. The task was simple: profile it, tell me what is wrong, and write a cleaning plan.
The agent read the first forty rows or so - all that fit comfortably in context. Then it summarized and continued. The report said 11 of the 48 columns are clean, a handful of nulls, the dataset “mostly fine with some gaps.”
But at row 1.2M, the category column drifted. Eight categories became fourteen. A new product line had shipped. The agent never scrolled far enough to see it.
The run journal was perfect. Linear. Zero re-entries. Zero stalls. The journal looked great. The output was wrong. Nothing in the system knew the difference.
The task
Here is the setup, exactly as it was given. The human hands the agent a large CSV with known messiness: nulls, duplicates, schema drift, outliers, inconsistent units. The human also hands it an eval criterion. The output is a quality report plus a cleaning plan, and two things have to be true of it: every column must be profiled, and every numeric claim in the report must be backed by a deterministic recompute.
The human does not design the workflow. The human gives the goal and the yardstick, then steps back. Everything from here on, the agent does itself.
The yardstick has three parts, and they answer three different questions:
| Signal | Question | Kind |
|---|---|---|
Run journal (getFsmRun) | How did the process run? Tokens, time per state, re-entries, stalls. | Fact |
| Completeness check (deterministic) | Is the output complete? Are all 48 columns profiled? | Fact |
| Claim check (deterministic recompute) | Is the output correct? Does every number in the report match a recompute? | Fact |
Critic (evalById) | Was the output any good? | Opinion |
Two facts, one opinion. The opinion is the weakest of the three, and the agent is told to state the numbers before it interprets them.
Round 0: the agent discovers the baseline is broken
No custom tools. No workflow. The agent reads the raw CSV, reasons over it, produces a report.
The journal is clean. Linear. No re-entries. No stalls. It reads like a perfectly engineered run.
But the output is incomplete. Run it again tomorrow and you get a different report because the model read a different window of the file. The report made 14 numeric claims: null rates, duplicate counts, outlier counts. The agent asserted them but never verified them.
The journal proved the process ran smoothly. It proved nothing about whether the output was right.
Round 1: the measurements reveal what the agent needs to build
The agent measures Round 0. Three signals, facts first.
- Tokens: 412,000. The raw CSV sat in context the entire run.
- Columns profiled: 11 of 48. The completeness check found only 11 columns touched. 37 columns were never looked at.
- Unverified claims: 14 of 14. The claim check looked for numeric assertions backed by a recompute. Found zero. The agent asserted 14 numbers and verified none of them.
- Critic score: 8 out of 10. “Looks thorough.”
The critic’s opinion and the deterministic facts are pointing in opposite directions. The critic thinks the report is thorough. The completeness check says 26.5 percent of the columns were profiled. The claim check says zero of the numeric assertions are verified.
Two gaps: the agent cannot hold the volume (37 columns missed), and the agent has no way to check its own work (14 unverified claims). The measurements name the problem. The agent now knows what to build.
Round 2: the agent discovers it needs a tool
The measurements revealed the gap: the agent cannot hold the volume. The solution is a tool.
The agent builds data_profile. It reads all 2M rows, all 48 columns. Outside context. For each column it computes: null rate, cardinality, type histogram, duplicates, outliers, unit normalization, schema drift.
The agent runs this tool and measures:
- Tokens: 38,000. Down from 412,000. The context holds the aggregated profile, not the raw CSV.
- Columns profiled: 48 of 48. The tool reads every row. Zero columns missed.
But the tool only computes. It says the category column has 14 distinct values after row 1.2M. It does not say whether that is a typo or a new product line. It does not propose a fix.
The tool solved the volume problem but not the interpretation problem. The agent needs something that can reason about what the data drift means.
Round 3: the agent discovers it needs a workflow with a gate
The agent now builds what Round 2 revealed was missing: a workflow that can interpret the profile and then check its own work.
Three states: profile runs the tool. classify reads the profile and writes the report: what is wrong, what it means, the cleaning plan. verify is the gate: it recomputes the profile and checks every numeric claim in the report against fresh numbers.
If the gate rejects, the report goes back to classify. The agent realizes this re-entrant loop is what makes the workflow able to fail. Without it, the journal stays clean while the output breaks.
Here is the shape of the whole thing. The tool owns the volume, the LLM owns the judgment, and the gate owns the boundary between them.
Here is the workflow, at the point where it converges. It is small on purpose.
fsm:
name: data-quality-audit
version: 1
initialState: profile
context:
profile: { type: string }
report: { type: string }
verdict: { type: string }
verdicts: { type: list, of: string } # one entry per verify pass; the list is the loop counter
agents:
auditor:
canTransition: [profile, classify]
guidance: full
states:
profile:
description: "Run data_profile over the full dataset. Set `profile` to the returned JSON summary."
requires:
- { name: profile, from: agent }
transitions:
- { toState: classify }
classify:
description: >-
Read the profile. Write the quality report: every issue, what it means,
and the cleaning plan. Every numeric claim must cite a column and a stat
from the profile. Set `report`.
requires:
- { name: report, from: agent }
transitions:
- { toState: verify }
verify:
description: "Recompute the profile and check every numeric claim in the report against it. Set `verdict` to the tool's string: approved or rejected with reasons."
requires:
- { name: verdict, from: { tool: audit_verify } }
transitions:
- toState: done
guards: [{ backend: cel, expression: "context.verdict == 'approved'" }]
- toState: classify # re-enter: report rejected, rewrite it
guards: [{ backend: cel, expression: "context.verdict != 'approved' && size(context.verdicts) < 3" }]
onTransition:
- { assign: { report: "" } } # consume the trigger so classify re-collects a fresh report
- toState: escalate
guards: [{ backend: cel, expression: "context.verdict != 'approved' && size(context.verdicts) >= 3" }]
done: { description: "Report approved by the verify gate. Claims checked: {{ context.verdicts }}.", terminal: true }
escalate: { description: "Report rejected three times. Verdict ledger: {{ context.verdicts }}. A human takes over with the full history.", terminal: true }
policy:
maxExchangesPerStage: 20
maxExchangesPerAgentPerStage: 10
checkpointEvery: 2
A few things to notice:
- The gate is deterministic.
verifydoes not ask the LLM whether the report is good. It recomputes the profile from the raw data and diffs the report’s claims against the fresh numbers. It returnsapprovedorrejected: <reasons>. That is a checkable fact, not an opinion. - The loop is bounded and observable.
verdictsis a list in context: the tool appends one entry per pass, so the counter is the data and it cannot disagree with the ledger it is counting. The budget issize(context.verdicts) < 3. When the budget is spent, the machine does not guess. It hands off to a human with the full history. - The re-entry edge consumes its own trigger. It clears
reporton the way back intoclassify, so the next pass re-collects a fresh report instead of routing on the stale one. A re-entrant state whose requirement is already present will not re-collect, so the edge has to blank the field it is about to re-gather.
The first time the gate fired, it caught a real error. The report claimed 12 percent nulls in customer_id. The recompute said 3.1 percent. The gate rejected the report, sent it back, and the corrected version came through on the second pass.
Before the gate, a clean journal meant nothing. After the gate, a clean journal means the report survived a recompute. That is the difference between a workflow that cannot fail and a workflow that can.
If you want the gate as plain code, here is the shape of it:
def verify(report, raw):
fresh = data_profile(raw) # recompute the whole profile
claims = extract_numeric_claims(report)
problems = []
for c in claims:
if fresh[c.column][c.stat] != c.value:
problems.append(f"{c.column}.{c.stat}: claimed {c.value}, recomputed {fresh[c.column][c.stat]}")
return "approved" if not problems else "rejected: " + "; ".join(problems[:8])
The rejected branch is the entire lesson. It does not trust the report. It recomputes.
This really happened
I keep this scenario as the main thread because it is clean to explain, but I want to point at the real artifact, because the pattern is not hypothetical.
In the same harness, there is a different task: a disease-etl workflow that fetches public health data, transforms it into twelve monthly records, writes them to a file, and verifies them before they are accepted. It went through five committed versions, and each one is archived.
The diffs between versions are small, and each one fixes a specific flaw the measurements exposed:
- v1 had the verify gate reading the file on disk. It was checking the wrong thing: the file, not the records in context.
- v2 fixed the gate to verify the records themselves, not the file.
- v3 added a
compact=trueflag to the fetch tool, because the raw payload was burning tokens. - v4 pinned the target year into the context so it flows through every tool call, because the year was getting lost between states.
- v5 told the worker to copy the raw string verbatim, because the model was re-serializing the nested object by hand and mangling the structure.
Five versions. Five archived files. Each one a response to a measured flaw, not a design preference. The archive is the proof that the evolution happened: .aegis/fsms/.history/disease-etl.1.yaml through disease-etl.4.yaml, with disease-etl.yaml at version 5.
That is not my data-quality scenario. It is a separate task that ran in the same harness, and it is the closest real cousin to the one above. The shape is identical: a deterministic core tool, a variable LLM layer, a verify gate that can reject, and a re-entrant loop that sends rejected work back.
Round 4: the agent discovers the optimum by tuning
The agent tunes, one change at a time. Measure. If a metric moves, it stays. If not, the agent stops.
Change one: The classify prompt considered all 48 columns. The agent realized the profile already flagged the problematic ones. Scope the prompt to only those. Re-run and measure: classify tokens dropped from 21,000 to 9,000. One number moved.
Change two: The profile tool did not normalize units before reporting. The report kept flagging unit mismatches the tool had already handled. The agent added unit normalization to the tool. Re-run and measure: false-positive rejections stopped. One number moved.
Nothing more. The agent ran the audit three more times. The gate did not reject. The metrics stayed flat for three consecutive runs. The agent discovered it had hit the ceiling and stopped.
The convergence signal was specific: the gate stopped rejecting. The gate is the only thing that can fail. When it stops failing, the workflow has converged.
The disciplines that run the whole loop are worth naming, because they are what keep the tuning honest:
- One change per round. If it batches, split it.
- Check before you build.
listFsmsandlistDynamicToolsbefore authoring anything new. - Report the regression. A loop that only ever reports improvement is broken.
- Report the ceiling. When three rounds show no measurable improvement, stop and say so.
The optimum: what the agent discovered
Here is the baseline against the converged optimum:
| Metric | Round 0 (baseline) | Optimum | Discovery |
|---|---|---|---|
| Tokens | 412,000 | 41,000 | The agent discovered it needs a tool to hold the volume |
| Columns profiled | 11 of 48 | 48 of 48 | The agent discovered a tool that reads every row |
| Unverified claims | 14 of 14 | 0 | The agent discovered it needs a gate to verify every claim |
| Critic score | 8/10 | 8/10 | The agent discovered the critic score was not the metric that mattered |
The critic score did not move. That is the point. The critic is a hint. The optimum is not a higher critic score. The optimum is zero unverified claims, every column profiled, and a gate that can reject and did not.
The agent discovered it had converged when the verify gate stopped rejecting across three consecutive runs. The measurements said “this is as good as it gets.” The agent reported the ceiling.
And here is what matters most:
No human designed the workflow, the tools, or the gate. The agent did.
The human gave the task: profile the dataset and write a cleaning plan. The human gave the eval criterion: every column must be profiled, every claim must be verified. Then the human stepped back.
The agent built data_profile because Round 1 revealed only 11 of 48 columns were profiled. The agent built the FSM because Round 2 revealed the tool could not interpret. The agent added the gate because Round 3 revealed the workflow needed to verify its own claims. The agent tuned the prompt and the tool because Round 4 revealed false positives.
Every decision was driven by a measurement. Every measurement revealed what needed to be built. The agent built it. Measured. Stopped when the metrics stopped moving.
The optimum is what the agent discovered by iterating, not what a human would have designed in advance.
Until next time…
The cleanest journal in the world is worthless if nothing in it can fail. The gate is what makes the journal mean something. Build the gate first, and the journal becomes a fact instead of a mood.
I keep the archived versions of every workflow I run this way, even the ones that look obvious in hindsight. The archive is where the learning lives: not in the final YAML, but in the diffs between the versions. If you are putting an agent in front of a dataset you do not trust, start with the gate. The rest is tuning.