Entity Resolution and Data Fusion


Back to: Intro

It starts the way every bad merge starts: with a decision that looks obviously right.

The agent reads two records. J. Smith, 12 Main St and John Smith, 12 Main Street. Same person, obviously. It merges them into one customer and moves on.

Then it reads a third. John Smith, 480 Harbor Ave, Portland, OR. Different city. Different street. Different state. But the name matches and the agent is running on momentum. It merges that one too. One customer now carries three addresses in two different states.

There was no rule, no config, no human in the loop. The agent read the records and decided. It decided wrong.

This is the baseline breaking.


The task

Three overlapping record sets. Customer data pulled out of a CRM, a billing system, and an old marketing list, all exported as flat JSONL. Roughly 12,000 records in total, and they overlap: the same person shows up in two or three of the sets, with the name spelled differently each time.

The goal is one clean customer master. Every record resolved to a customer, or explicitly marked unmatched. No duplicates, no phantom customers, no one person split across two IDs.

And the yardstick, which I gave it up front and then stepped back: a deterministic check that recomputes the expected grouping from the raw input and diffs it against whatever the agent produced. A false merge or a false split is a checkable fact. If the agent merges two different people, the check says so. If it splits one person, the check says so. No opinions involved.

The human’s job here is exactly two things. Give the task. Give the yardstick. Then get out of the way.


Round 0: the agent discovers the baseline is broken

No custom tools. No workflow. Just the agent and 12,000 records.

The agent reads them one at a time, trying to match entities by hand. It keeps a mental list of “customers,” scanning for the next record that looks like one it has seen.

By record 4,000 the context is full. The list of names and addresses is a wall. The matching drifts. It merges J. Smith and John Smith on Main Street. A few hundred records later it merges that customer with the John Smith in Portland. It also splits one real customer across two IDs because the two records that belonged together never appeared close enough together in the stream to be noticed.

The output has false merges and false splits baked in. Run it again tomorrow and the merges come out slightly different.

The run journal is clean. Linear. No re-entries. Perfect process.

Two different people in one ID. One person split across two IDs. The journal proved nothing about whether the process was right.

Hand-merging 12,000 records by reading them in context does not work.


Round 1: the measurements reveal what needs to be built

The agent measures Round 0. Three signals, facts first.

  • Tokens: 41,200. The raw JSONL for all three record sets sat in context the whole way through.
  • False merges: 2. The Portland John Smith was folded into the Main Street customer. Two records with different streets, cities, states in the same group, with no linking evidence in the raw input.
  • False splits: 3. One real customer appeared in both the CRM and billing system. The raw input has a shared phone number. The output put them in different groups.
  • Unmatched records: 174. Records that should have been resolved to known customers but were not.
  • Completeness: 11,826 of 12,000 records. 174 fell off.
  • Critic score: thorough and well-reasoned. The critic thought the work was good.

The critic’s opinion and the deterministic facts are pointing opposite directions. The critic says thorough and well-reasoned. The correctness check says 2 false merges, 3 false splits, 174 unmatched.

The deterministic part (normalization, finding which records might match) is drowning in raw volume. The variable part (deciding whether two records actually match) has no gate and keeps making mistakes.

The agent now knows what to build.


Round 2: the agent discovers it needs a tool

The measurements revealed the gap: normalization and candidate-finding are drowning. These are pure computation - they do not need judgment. They need to run over the whole volume without a context window getting in the way.

The agent builds entity_block.

entity_block does the deterministic, high-volume part:

  • Normalization. Canonicalize names (initials, spacing, case, suffixes), standardize addresses (street abbreviations, unit formats, state codes), normalize phone numbers to a single form.
  • Exact-key dedupe. Drop records that are identical after normalization.
  • Blocking. Group records that might match by a blocking key. A pair that could be the same person must land in the same block. The key is a normalized surname plus a coarse location token. That cuts the candidate pairs from the full cross product down to a small set of same-block pairs.
  • Joins. Stitch the normalized fields so a record from the CRM and a record from the billing system can be compared field by field.

The tool processes the volume outside the context window. The agent’s context no longer holds 12,000 raw records. It holds the tool’s output: a list of blocks, each a short list of candidate records. Token cost drops hard, because the by-hand approach is gone.

Here is the architecture the tool sits in.

approved

rejected

Raw record sets

entity_block tool

Normalized records

Blocked candidate pairs

LLM fuzzy resolution

Proposed customer groups

verify gate

Customer master

The tool crushes the by-hand approach on the deterministic part. But it only computes. It can tell you that J. Smith and John Smith are in the same block. It cannot tell you whether they are the same person. That is the variable half, and a static tool would be overfit the moment it tried to answer it.

That is the second half of the “why not just write the tool?” question. For the deterministic half, you should write the tool. For the variable half, a tool would fail on the new cases. The LLM’s generalization is the feature, not the bug.


Round 3: the agent discovers it needs a workflow with a gate

The agent now builds what Round 2 revealed was missing: a workflow that can do fuzzy resolution and then check its own work.

A finite state machine with three states:

resolve. The LLM does fuzzy entity resolution over the blocked candidate pairs. It decides, per pair, whether they are the same person. This is the variable part. The LLM weighs a name, an address, a phone number, and context. A static rule would miss these judgment calls.

verify. This is the gate. It recomputes the expected grouping from the raw input and checks the proposed groups. Is every group actually the same person? Are there records that should have merged but did not? It looks for false merges and false splits. It returns approved or rejected: <reasons>.

The re-entrant loop. If verify rejects, the work goes back to resolve. The gate makes the workflow able to fail.

approved

rejected, re-enter

resolve

verify

done

This is the shape that repeats across the series: a deterministic tool in its lane, an LLM in its lane, and a gate that checks the boundary between them.

There is one subtle trap the agent had to handle, and it is worth naming. A re-entrant state’s requirement is satisfied the moment its field is present. If verify rejected but left the old resolution sitting in context, resolve would re-enter and find its requirement already met, and fall straight through without doing any new work. The loop would spin and nothing would change. The fix is to consume the trigger on the way back: blank the resolution field in onTransition so the next entry re-collects a fresh value. Small detail. Without it, the gate is a decoration.


Round 4: the agent discovers the optimum by tuning

The workflow works. The agent now discovers how to make it better. One change at a time. Measure. If the metric moves, keep it. If not, stop.

Change one: the blocking key. The agent measures and realizes the key is too coarse. Too many unrelated people in the same block. The LLM is wading through junk pairs. The agent tightens the key to a finer location token. Re-run and measure: candidate pairs per block dropped from a median of 41 to 9. Tokens per resolution pass dropped to match. One number moved.

Change two: the similarity threshold. The agent measures and realizes the LLM’s threshold for merging is too strict. Too many false splits. The agent lowers the threshold from 0.8 to 0.7. Re-run and measure: false splits dropped from 3 to 1. The phone number link is now caught. One number moved.

Change three: the threshold again. From 0.7 to 0.65. Re-run and measure: false splits dropped to zero. False merges stayed at zero. One number moved.

Convergence. The agent ran the same change a third time and got the same numbers. The metrics stopped moving. The agent discovered it had hit the ceiling and reported it.

The signal was a checkable fact. A false merge is “two records in the same group with no linking evidence in the raw input.” You can count them. You can watch the count go to zero and stay there. That is the sharpest signal in the whole series.


The workflow that won

Here is the workflow that won, as the agent committed it. It is a real Aegis FSM, valid against the DSL. Read it top to bottom and you can see every discipline from the rounds above encoded in the machine itself.

fsm:
  name: entity-resolve
  version: 4
  initialState: block

  context:
    blocking_key: { type: string }
    threshold:    { type: decimal }
    candidates:   { type: list, of: string }
    resolution:   { type: string }
    verdict:      { type: string }
    attempts:     { type: list, of: string }

  agents:
    resolver:
      canTransition: [block, resolve, verify]
      guidance: full

  states:
    block:
      description: "Call entity_block(blocking_key={{ context.blocking_key }}, threshold={{ context.threshold }}); set `candidates` to the blocked pairs."
      requires:
        - { name: candidates, from: agent }
      transitions:
        - { toState: resolve }

    resolve:
      description: "Fuzzy-resolve the blocked pairs in `candidates`; set `resolution` to the proposed customer groups."
      requires:
        - { name: resolution, from: agent }
      transitions:
        - { toState: verify }

    verify:
      description: "Call verify; it recomputes the expected grouping and checks `resolution` for false merges and splits. Set `verdict` to its verdict string exactly."
      requires:
        - { name: verdict, from: agent }
      transitions:
        - { toState: done, guards: [ { backend: cel, expression: "context.verdict == 'approved'" } ] }
        - toState: resolve
          guards: [ { backend: cel, expression: "size(context.attempts) < 3" } ]
          onTransition:
            - { append: { attempts: "${ 'rejected' }" } }
            - { assign: { resolution: "", verdict: "" } }
        - { toState: escalated, guards: [ { backend: cel, expression: "size(context.attempts) >= 3" } ] }

    escalated:
      description: "Three passes failed verification. Escalate for review."
      terminal: true

    done:
      description: "Customer master resolved and verified."
      terminal: true

Three things to notice.

Loop-bounding. The attempts list is the counter. Each rejected pass appends one entry, and the verify to resolve edge is guarded on size(context.attempts) < 3. When the counter hits three, the machine leaves the loop and parks at escalated instead of spinning forever. The counter is the data. There is no hidden integer to increment, because assign is literal-only and cannot do arithmetic.

Consume the trigger. The verify to resolve edge blanks resolution and verdict in onTransition. Without that, resolve would re-enter with the old resolution still sitting in context, find its requirement already met, and fall straight through without doing any new work. Blanking the field is what forces a fresh pass.

canTransition lists every state the agent acts in. block, resolve, and verify are all in the list, including the re-entrant resolve. Miss one and the run stalls permanently, because the agent has no state it is allowed to act in.

The guards are all CEL, in the same form the published Agents in Production series uses. One form, used consistently. That is a house rule, and it keeps the machine readable.


The optimum: what the agent discovered

Here is the baseline against the converged optimum:

MetricRound 0 (baseline)OptimumDiscovery
Tokens41,2006,400The agent discovered a blocking tool reduces context load
False merges20The agent discovered the verify gate catches false merges
False splits30The agent discovered the threshold tuning eliminates false splits
Unmatched records1740The agent discovered strict verify requires complete resolution
ReproducibleNoYesThe agent discovered the workflow is now deterministic

The token drop came from the deterministic core doing its job. The context holds blocked candidate pairs, not 12,000 raw records. The false-merge and false-split columns both went to zero. The unmatched column went to zero because every record is resolved or explicitly marked.

The signal that told the agent it had converged was the false-merge and false-split count hitting zero and staying at zero across two consecutive runs. The metrics stopped moving.

And here is what matters most:

No human designed the workflow, the tools, the blocking key, the similarity threshold, or the tuning strategy. The agent did.

The human gave the task: resolve and dedupe three overlapping record sets into one customer master. The human gave the yardstick: run journal, correctness check, completeness. Then the human stepped back.

The agent built entity_block because Round 1 revealed normalization and candidate-finding were drowning in volume. The agent built the FSM and gate because Round 2 revealed the tool could not make judgment calls about whether records actually matched. The agent tuned the blocking key because Round 3 revealed the key was too coarse. The agent tuned the similarity threshold because Round 4 revealed the threshold was causing false splits and false merges.

Every decision was driven by a measurement. Every measurement revealed what needed to be built. The agent built it. Measured. Stopped when the metrics stopped moving.

The optimum is what the agent discovered by iterating, not what a human would have designed upfront.


A closing note

This is the last one. Four examples, four different tasks, and the same loop running underneath every single one of them.

Logs, messy data, tickets, and now entities. Different domains, different failure modes. But the shape never changed. The agent starts by doing the work by hand and failing in a way the clean journal hides. The eval names the gap with a number, not an adjective. The agent builds the deterministic tool for the half of the job that is computation. It builds the LLM lane for the half that is judgment. It builds the gate that makes the workflow able to fail. And it tunes, one change at a time, until its own measurements stop moving.

I keep coming back to the thing that surprised me about this. It was not that the agent could build the workflow. It could. What surprised me was that it could stop. Most systems that can improve will keep improving, chasing a metric that has already flattened. This one ran the same change a third time, saw the same number, and reported the ceiling. That restraint, the willingness to say “this is as good as it gets from here,” is the part I would not have trusted an agent with a year ago.

The human’s job stayed the same size the whole way through. Give the task. Give the yardstick. Approve the dangerous operations. Everything else was the agent’s. I am not sure that line between “the human’s job” and “the agent’s job” is where it will stay for long, and I am not sure that is a bad thing.

If you are building agents that do real work, start where this series started. Give one of them a task it will do badly by hand, and a yardstick that can say so. Then watch what it builds.

Back to the start: Intro