The Data Driven Dispatch ← Four tellings
The Agentic Company

The man who beat his data pipeline by deleting it

He set out to stop his data from rotting and spent six months adding hours and parts to keep up. The fix, when it came, was to take things away.

The Dispatch Staff
Reconstructed from 2,040 local agent transcripts · June 12, 2026

The record showed the same company twice: one row marked canonical, one marked a duplicate. The dedup document that was supposed to settle which was which told Tristan to do two different things with them. It was late; the database held millions of rows. He read the document again.

He had written two duplicate-domain policies into it, on different days, that could not both be followed, and had not noticed until now. The bug was not in the data. It was in the rule. Duplicate-and-canonical confusion is one of the ordinary ways a contact database goes wrong, and the machinery he had built to catch it was now producing it.

For six months, the transcripts show, his response to a problem was to add: more hours, more workers, more rules. The fix that ended it ran the other way.

The rent on rot

Tristan runs Data Driven Partners, a lead-generation and cold-email firm. He has no formal engineering training. “Everything I have done has been me,” he told an A.I. he asked this spring to coach him on data architecture. The business runs on business-contact data — who runs which company, their email address, whether it still resolves — and for years he bought it the way the industry does. He rented it.

The trade rests on a fact its customers rarely price out: the product spoils. Studies put the decay of B2B contact data at two to three percent a month. People change jobs, companies fold, domains move. By that rate roughly a third of any database is wrong within a year. The incumbents do not hide this. They sell subscriptions to catalogs that are partly wrong, and they sell the corrections too.

The rent, Tristan came to argue in his sessions, bought a cycle: acquire data, enrich it, mail it, watch it go stale, pay again. Nothing he finished was likely to still be accurate a month later.

In December he decided to stop renting and build his own store. He worked at night, alone, one message at a time. The longest single conversation with a coding assistant in the logs runs 186 hours. December accounted for 304 hours of pipeline work, nearly all of it with him in the loop for each step.

Fifty-nine connections

By the end of winter the store existed: a Supabase database, queue-based enrichment workers, a domain finder, dashboards. In February he extended the logic to supply, deciding renting raw records made no sense either, and began scraping his own. He scaled the intake until the database’s connection pool strained — one transcript records 59 simultaneous connections. The store reached 37.2 million prospects.

The volume rose; his trust in it did not. Each record arrived stripped of its history: where it came from, how confident the source was, when it was last checked. The data was not so much wrong as unaccountable. A field that cannot say where it came from cannot be re-checked, and re-checking is the whole job.

In April he began selling — custom lists pulled from the store at a dollar a lead, 24-hour turnaround. The sums were small. The exposure was new: an error that had been a private annoyance was now a possible refund. The reconciliation machinery — the merges and the dedup rules that decided which of two near-identical companies was real — was where the errors clustered.

Boil the lake

What changed in May was the labor, not the goal.

Through April he had worked one conversation at a time, supervising each step. In May he changed the arrangement. He wrote standing orders for fleets of coding agents — one document opens, “Boil the lake. Do all the work to take something from 80% to 100%” — and ran them in parallel.

One swarm of agents ran 141 sessions looking for places where records were silently dropped. A second ran 209 sessions over ten days, turning what the first found into a set of data-quality rules. In May the systems logged 524 hours of agent work, against 4,615 messages from him.

Exhibit
Where the hours went, month by month
System-active hours of data work, by tool. The work didn’t shrink in May. It stopped being his.
Source: 2,040 local agent transcripts, parsed June 12, 2026. Hours sum gaps between events, capped at 10 minutes; they include autonomous agent runtime.

December’s 304 hours and May’s 524 are not the same hour. In December each hour had a person attached to it. By May most hours were machines executing written specifications while he directed. The total went up; his share of it went down.

Two things moved at once, and the transcripts do not cleanly separate them. His own competence rose over the seven months. So did the tools. The models he used in November often could not finish a single task without supervision; by May he was leaving fleets of them running overnight. Whether the gains came mainly from the operator or from the software underneath him is not something the logs settle. What the newer models removed was the gap between specifying a task and having it run unattended. That left more of the work resting on what he chose to point them at.

The flaws the swarms surfaced had not been hidden so much as postponed. By his own account in the sessions, he had known the data was leaking somewhere and had kept answering with more hours rather than restructuring. The transcripts record how long an untrained operator can substitute effort for structure before it stops working. Here it ran about six months.

Verb and noun

Late in May he opened a session that did not ask for a feature. He put a different proposition to the model: that the method itself — more hours, more machinery — was the problem. The distinction that came back was between a pipeline and a stored asset. A pipeline runs, finishes, and pushes its output downstream, where it ages. He had kept building that and kept watching it rot. The alternative was a store a process keeps current: never overwrite a record, append observations and compute the current value from them, and attach to every field a source, a confidence level, and a last-checked date — so that “is this still accurate?” becomes a query rather than a guess.

The diagnosis took an evening. A code review run alongside it concluded the architecture was already most of the way there: the schema of observations and golden records existed and leaked at a few seams. The two sessions that reframed the work ran 6.7 hours, against a prior 1,350 or so. The idea was cheap. What the 1,350 hours bought, on this telling, was an operator who would act on it instead of adding another rule.

Out of the loop

What runs in June is not operated by hand. Scheduled monitors check the morning’s data pulls. A daily job chases records that get stuck. A control-tower loop posts a summary of the operation during business hours. The system pulls its own raw material from an in-house scrape farm, declines to sell a record that cannot show its source and last-checked date, and flags its own failures. The reports are filed whether or not anyone reads them.

The caveats are substantial. The figures here are parsed from his own transcripts, not audited books. The leaks the swarms documented were written up for repair; the logs do not yet show them closed. In-house scraping carries terms-of-service and regulatory exposure that a licensed vendor would otherwise carry for its customers. And the whole operation’s judgment sits in one person — a dependency no schema removes.

It is also a lot of work for one person. The timestamps put him at the keyboard at 2 a.m. and again before 6. During the week the rules swarm peaked, another window shows him running revision loops on a song he was writing, “Hail Mary.”

Across the seven months the arrangement inverts. In December he worked the machine alone at night, message by message. In June the machines mostly exchange messages with each other — about stuck records, stale fields, the morning’s pulls — and he reads the summary, when he reads it. The loop he spent the winter inside now runs without him in it. What is left to him is choosing what it works on next.

He set out to fix a data pipeline. What he is left with is a procedure — written specifications, fleets of agents, scheduled checks — that is not specific to data. The records will keep decaying at two to three percent a month. The system now detects the decay, re-verifies, and charges for the corrected version. Whether that holds up outside his own telemetry is the open question.

Methodology — This account draws on 467 Claude desktop audit logs, 786 Codex rollout logs, 36 Claude Code transcripts, and 631 Cursor conversations (715,297 messages) on the operator’s machine, parsed June 12, 2026. Quotations are from his own session messages, lightly edited for capitalization and spelling. “System-active hours” sum the gaps between transcript events, capped at 10 minutes each; they include autonomous agent runtime and exclude idle time. Not covered: ChatGPT, DeepSeek, and Perplexity histories, and work done in hosted tools.