The Agent That Builds Its Own Tools


I gave an agent a job and watched it figure out how to do it better, and better, until it stopped improving.


The premise: point an agent at a problem and let it iterate

The question this series answers is simple: What happens if you give an agent a task, a way to measure success, and the power to build its own tools and workflows?

The answer is surprising. The agent doesn’t just execute. It discovers. It runs the task, measures the output, identifies what’s broken, invents a solution, commits that solution to a workflow, and repeats. Each cycle builds on the last. Each measurement names a specific gap. Each gap prompts the agent to build something novel - something optimized for this problem, not a generic tool you would have written by hand. The agent keeps iterating until its own measurements say the work is done.

That’s the whole idea. A self-evolving agent harness lets you point an agent at a general problem and have it create novel solutions by iterating on its own feedback. The human gives the task and the yardstick. The agent does the rest: builds the tools, writes the workflow, measures itself, revises, and stops when the metrics converge.

This is not “an agent running a pipeline you designed.” This is “an agent discovering what pipeline would work best for this job.”


Round zero: the agent discovers the baseline is broken

The first job was simple to say. Pull a year of public health data for a specific disease, turn it into twelve monthly records, write them to a file. No design. No architecture. Just a transform.

I handed it to the agent. No custom tools. No workflow. No state machine. Just the agent, a big context window, and the raw data.

It worked. The agent read the data, reasoned over it, computed the months, and wrote a file. Twelve lines. One JSON object per month. It looked right. The run journal was immaculate. Linear. No re-entries. No stalls. Perfect.

The journal was clean because nothing in it could fail. There was no quality gate, no verify step, no mechanism to catch a mistake. So the journal proved nothing about the output. A clean journal is not a good outcome. It’s the absence of a check.

This is round zero: the baseline. The agent does the work by hand, in context, and produces an answer. The output looks good. But there’s no gate, so a mistake hides in the journal as easily as a success does. You can’t tell the difference by reading the process. You only learn by measuring.

The agent measures round zero. Discovers it’s broken. And figures out what to build next.


The discovery loop: iterate until the metrics stop moving

So how do you get from “the agent does it by hand” to “the agent built an optimized workflow”? You build a feedback loop and let the agent iterate.

The agent runs four phases, over and over, each time discovering something new about the problem and what would solve it.

RUN. The agent executes the workflow (or in round 0, the raw baseline). It forks a worker: deliberately simple, deliberately not clever. The worker is the subject of the experiment. If the worker is smart, you can’t tell whether the workflow works or the worker is compensating. So the worker is dumb on purpose.

MEASURE. The agent measures the run. Three distinct signals. The rule is iron: facts before opinion. The measurements reveal not just that something is wrong, but what is wrong and how bad it is.

DISCOVER. This is the key move. The agent looks at the measurements and realizes what needs to be built. Round 0 failed at high cost, so the agent discovers it needs a tool. A tool that counts but doesn’t interpret, so the agent discovers it needs an LLM layer too. A workflow with no gate, so the agent discovers it needs one. The discovery is data-driven, not intuition-driven.

BUILD. The agent authors the solution it discovered. It writes a tool for the deterministic part. It designs a workflow (a finite state machine) for the part that needs judgment. It adds a gate that makes the workflow able to fail. And when it commits a new version, the previous version is archived. A bad round is recoverable.

Then the loop repeats. The agent runs the new version, measures it, discovers what moved and what didn’t, and builds the next version. The agent keeps iterating until its own measurements say the work is done: the metric it set out to improve has plateaued across three consecutive runs.

convergence 3 rounds no improvement

MEASURE

DISCOVER

BUILD

RUN

CEILING

This is not a pre-designed process. The agent discovers the process itself, one measurement at a time. The human gives the task and the yardstick. Everything else: the tools, the workflow, the gate, the decision to stop. The agent figures it out.


Three signals for discovering what’s actually broken

The MEASURE step is where discovery happens. Most systems reach for one signal and call it done. But one signal lies. You need three, and they discover different things.

SignalWhat it revealsToolKind
EfficiencyHow the process actually rangetFsmRunFact (the run journal)
CorrectnessWhether the output is actually rightA deterministic recompute + diffFact (checkable)
QualityWhether the output is any goodevalById (LLM critic)Opinion (a hint, not a score)

Efficiency is the factual run journal. It records transitions, time per state, re-entries (rework), stalls. It’s authoritative on how the process ran, not whether it worked. A clean journal with no re-entries doesn’t mean success. It might mean nothing in the workflow can fail.

Correctness is a deterministic check. It recomputes the expected output from the raw input and diffs it against what the agent produced. It returns approved or rejected: <reasons>. This is checkable fact. It’s the only signal that can say “you missed 8 error classes” or “you merged two different people.” And this is the signal that drives the agent’s next discovery: if the gate rejects, something is broken that the agent didn’t build yet.

Quality is the LLM critic’s opinion. It’s the weakest signal: it drifts toward “looks thorough.” Treat it as a hint. But it’s the only one that can speak to the variable, judgment-laden parts. The other two can’t tell you if a narrative actually makes sense.

A run can be clean in the journal, pass the correctness check, and still be expensive or incomplete. It can have a perfect critic score and still be wrong. The agent needs all three, and must state the numbers before interpreting them. The agent reads the three signals and asks: what does this tell me to build next?


Every task has two halves

Here’s the question every task raises: why not just write a tool and be done with it? Why do you need the LLM at all?

Because every task has two halves, and they want different machines.

The deterministic core. The part that’s voluminous and rule-based. Parsing, aggregation, profiling, dedupe, normalization. A custom tool the agent builds crushes the LLM doing it by hand. Fewer tokens, because the tool processes the volume outside the context window. Faster. No context-window ceiling. Reproducible. For this half, you should write the tool. That’s the first half of the answer to “why not just write the tool?”

The variable layer. The part that’s genuinely variable and requires judgment. Root-causing, interpreting, classifying, resolving. A static tool can’t handle it, so the LLM reasons per item. A tool would be overfit and would fail on the new cases it hasn’t seen. For this half, the LLM’s generalization is the feature, not the bug. That’s the second half of the answer.

The design principle is that the agent learns to put the deterministic tool in its lane and the LLM in its lane, and to build a gate that checks the boundary between them.

The optimum is when each half is done by the thing that’s best at it, and the gate catches the cases where the variable half gets it wrong.

This is where the series builds on the Agents in Production work: the LLM sits inside a deterministic shell, and workflows are state machines, not task lists. Here the shell itself is built and tuned by the agent, measured against its own evals, until it’s optimal for the task.


A clean journal is not a good outcome

I didn’t discover the clean-journal trap by accident. I planted it.

I built a small workflow called summarize-doc: read a document, produce a summary, done. Two states. No verify step. No recompute. No gate. Nothing in the workflow that could fail.

Then I ran it and looked at the journal.

It was beautiful. Linear. Every state visited once, in order. No re-entries. No stalls. The kind of run you’d put in a slide deck.

And the output was worthless. The workflow had a planted flaw: it summarized the document by taking the first paragraph and calling it done. The journal had no record of that, because the journal only records process, and the process ran exactly as designed. The flaw was in the design, and a journal can’t see a flaw in the design.

That’s the point, and it’s the reason the series keeps coming back to it.

A workflow with no quality gate cannot fail, so a clean journal proves nothing about the output.

The journal tells you how the process ran. It does not tell you whether the output is right, complete, or any good. Those are the other two signals, and they have to come from outside the journal: a deterministic verify tool that recomputes and diffs, and a critic that can judge the variable parts.

This is the trap every round-zero run sits in. The agent does the work by hand, the journal looks clean, and the only thing standing between “looks great” and “actually great” is a check that wasn’t there.

So the discipline is: never accept a clean journal as evidence of a good outcome. Build the gate. Make the workflow able to fail. Then a clean journal finally means something.


The disciplines: how the agent avoids false turns

A loop with no rules is just thrashing. To discover the real solution instead of chasing noise, the agent follows four disciplines.

One change at a time. When the measurements reveal something that needs fixing, the agent changes exactly one thing. Re-runs. Measures. If the metric moves, the agent knows which change did it. If you change three things and a metric improves, you’ve learned nothing. One change per round keeps the signal clean.

Check before you build. The agent does not build in response to a hunch. Before authoring a new tool or a new workflow state, it checks: do the measurements actually name this gap? “Completeness is 60 percent, missing 8 classes” — a named gap. Build. “Tokens are high” — maybe, but is it the metric that matters? Check first. If the measurements don’t show a gap, don’t build.

Report the regression. If a change makes a metric worse, the agent says so explicitly. The version is archived, so recovery is always available. The reporting is what matters: the agent is learning from failed turns.

Report the ceiling. If three rounds in a row show no measurable improvement, the agent stops. It reports the ceiling: “the metric is not moving, this is as good as it gets from here.” It does not loop hoping something will change.

The disciplines are all about letting the data drive discovery. The agent states the numbers first, interprets them second. That keeps the agent honest.


This is real: five archived versions of one workflow

This series is all true. Not a thought experiment. Real runs, real metrics, real workflows that converged.

In the self-improve sandbox there’s a workflow called disease-etl. It pulls a year of public health data, transforms it into twelve monthly records, writes them to a .jsonl file, and verifies them against the raw data before they’re accepted.

That workflow went through five versions. The agent built version 1, measured it, discovered what was breaking, built version 2. Measured again, discovered a new problem, built version 3. Five times. Each version archived.

The discovery is visible in the diffs. Here’s version 1’s extract state:

extract:
  description: >-
    Call disease_fetch(year) and set `raw` to its JSON output
    (the object with the covid and influenza slices). Set `out_path`
    to out/disease_<year>.jsonl.

Version 5’s extract state:

extract:
  description: >-
    Use the `year` context field (the seeded target year) for everything
    in this state. Call disease_fetch(year=<year>, compact=true). The tool
    returns `raw` already serialized as a single JSON string of the
    {covid, influenza} object. Set `raw` to that `raw` string value
    VERBATIM: copy it exactly, do not re-serialize or reformat the nested
    object by hand. Set `out_path` to out/disease_<year>.jsonl.

In v1, the agent set raw to “its JSON output” — the agent would serialize it by hand. That’s an error class: dropped keys, misformatted nesting, fields lost in the copy. The verify gate would reject these.

By v5, the agent had learned: don’t re-serialize by hand. The tool returns a pre-serialized string. Copy it verbatim. Error class gone.

The agent also learned: context fields drift. By v5, the year flows explicitly through every tool call. A stale year can’t silently corrupt the transform.

And: the gate was checking the wrong thing. By v5, it checks the in-context records themselves, not the file on disk.

Each discovery came from a measurement that said “something is wrong.” The five archived versions are the record of what the agent learned.

No human designed this workflow or noticed these problems. The agent discovered them, one measurement at a time.


The four examples

The concept is only as good as the examples. Here are the four that follow, each one a full BUILD to RUN to MEASURE to REVISE arc on a different kind of data.

PartTaskDeterministic core (the tool the agent builds)Variable layer (the LLM’s lane)
Part 1: Log and Telemetry Triage at ScaleA large log volume to triagelog_triage: parse, dedupe by signature, aggregate counts and trendsRoot-causing, correlating signals, judging severity
Part 2: Data Quality Audits over Messy DataA large messy dataset to profileA deterministic profiling tool over the whole volumeClassifying what’s wrong, writing the cleaning plan
Part 3: Support Ticket Mining at ScaleThousands of free-text tickets to classify and aggregateDeterministic aggregation over extracted variablesExtracting variables and classifying per ticket
Part 4: Entity Resolution and Data FusionSeveral large overlapping record sets to dedupe and joinA deterministic blocking toolFuzzy resolution of the ambiguous matches

Each part follows the same arc: a round-zero baseline where the agent does the work by hand, an eval that names the gap, a deterministic core that crushes the volume, a variable layer with a quality gate, and a tune-to-optimum that stops when the metrics stop moving.

Each part also has a dedicated optimum section: the final metrics against the baseline, the specific signal that told the agent it had converged, and the explicit statement that no human designed the workflow or the tools. That’s the headline of the series, and each part drives it home with its own numbers.

Part 3 builds on Deterministic Shell, Probabilistic Core, and Part 4 on Stop Treating Agent Workflows Like Task Lists. If you’ve read the Agents in Production series, you know the shell. These four parts show the agent building the shell itself.


Until next time: what surprised me

The first time I watched an agent build a tool for me, I felt something I didn’t expect. Not that it works. It does. What surprised me was the discovery. The agent didn’t just follow a recipe. It measured a baseline, looked at the numbers, and realized: “I need a tool that reads the volume outside the context window.” That’s not something I told it to do. It figured it out because the measurement made the problem obvious.

Then it built that tool, ran it again, measured again, and realized the tool could count but couldn’t interpret. So it discovered it needed an LLM layer. Then it ran that and discovered it needed a gate. It didn’t know any of this in advance. Each discovery came from reading the metrics and asking “what would fix this?”

I’ve spent my career building shells: the pipelines, the state machines, the verify steps that sit around an LLM and keep it honest. The assumption was always that a human designs the shell. This series is what happens when you flip that. The agent designs the shell. The agent measures it. The agent improves it. The agent stops when the metrics say it’s done.

My job becomes three things: give the task, give the yardstick (the metrics that will drive discovery), approve the dangerous operations. Everything else is the agent’s. The agent discovers what tools to build, what workflow to design, what gates to add, and when to stop.

That’s the power of a self-evolving harness. You don’t give the agent a design. You give it a problem and a way to measure. And then you watch it discover the solution.


The four examples that follow

Each of the next four parts is a real run. A real task. Real metrics, from round zero to convergence. Each one shows an agent discovering a solution to a different kind of problem.

The pattern repeats, but the problems are different. Logs you need to triage. Messy data you need to audit. Tickets you need to classify. Records you need to resolve. The agent discovers what to build in each case, and you get to see the archived versions that prove it actually happened.

Read them and ask yourself: Is the agent’s solution as good as one I would have designed? I think you’ll find it’s not just good enough. I think you’ll find the agent got there by doing the work, by measuring, by iterating, by stopping when the metrics stopped moving. And I think you’ll be surprised at what it built.

I’ll see you in Part 1, where an agent reads a 40MB log by hand and misses the critical error at line 300,000. Then it discovers what to build to never miss it again.

Up next: Part 1 (Log and Telemetry Triage at Scale)