Get It In Writing

Get It In Writing

Consider a trip to the auditor.

You’ve received a notice from the ATO. Now you have to present everything to justify the last seven years of your tax filings.

In order to prepare for that audit, you will need to gather up all of your records. Hopefully you still have them.

If you have one, you’re going to want to bring your accountant along with you, because they’re well versed at answering the kinds of questions that come from the auditor.

The auditor is the judge of the matter. They decide whether your answers are sufficient. They decide whether your records are correct. They decide whether you get a fine or you get the all clear.

Any time you front an auditor you have to hope for the best. Because the auditor is not your friend.

It’s unnerving. These adversarial situations are designed this way by intent. Because adversarial frameworks are the most reliable way we have to get to the truth.

You can’t just walk up to the auditor and convince them that all of your math and all of your rationales are fine. It’s not that easy.

They’re going to go by the letter of the law and their rules and their rituals. They know their job is to lean very hard.

That's the way this is meant to work, because the harder they go, the closer everyone gets to the truth.

Out of very imperfect qualities comes - if not perfect knowledge - something that's 'good enough'.

Enough to satisfy the ATO, or your shareholders, or whoever needs to know the real state of affairs, minus human mistakes and foibles.

Without this process - an accumulation of processes for verification - no individual or business would ever really be able to trust financial claims made by another.

Business would be effectively impossible. Too risky.

That's the reason we're willing to bear the burden of the occasional audit: it prevents everything from unwinding into a chaotic fog of uncertainty.


We seem to have hit that fog with artificial intelligence.

For half a century we've had 'personal' computers that were dumb, but very accurate. They would add and subtract, multiply and divide, and do so flawlessly. The rare exceptions - like the famous Pentium division bug - stand out because they violate that essential quality of computing. It tells the truth.

Since the end of 2022, we've seen that change. Computers now are fantastically smart. In some areas, they're smarter than any of us - outside a tiny handful of super-elite experts.

There's just one problem: computers are no longer accurate.

Almost anyone who has used modern AI systems has witnessed a 'hallucination' - where the AI simply gets its facts wrong. The reasons for this are understood only in rough terms, but it boils down to this: making something that smart also makes it less accurate.

This isn't impossible to fix. We've worked out a range of ways to reduce hallucinations, the first and most obvious of which is a technique humans use all the time: when in doubt, consult an expert.

When an AI system generates an answer to a question, it 'checks the facts' against a system of record. That might be a database specific to an organisation, or a standard reference like Wikipedia, or the results of a broad search of the Web. When it presents its response, it 'backs up' its facts with links to the relevant references.

Hallucinations haven't disappeared, but they're significantly less common than they were a few years ago, because AI references these 'systems of record'.

That's a timely fix, because we're now building AI systems - agents - that can work on complex tasks.

The problem with agents is that those 'hallucinations' add up. Even at today's low hallucination rates, a long-running agent will eventually encounter one. It's just probability: the longer an agent runs, the higher the probability.

Hallucinations are fatal for agents: Put a hallucinated value into a spreadsheet and all the calculations go awry. Hallucinate a citation in a legal brief and the judge will fine you, then send you off for a disciplinary hearing.

Agents are interesting, tempting, but - for this reason - deemed too risky for the enterprise.

That would be a show-stopper for agents, but for one simple fact: we've built civilisation out of people who are unreliable, error prone, and - occasionally - of dubious ethics.

The foibles of humans offer the path forward for agents.


People are imperfect. That's not a judgement, it's an observation. We mis-remember, we forget, we misplace and misunderstand. We have a lot of words for the ways we get it wrong, because it's so universal.

More than five millennia ago, we worked out a workaround for imperfections: we got it in writing.

Writing was invented by accounting. Ponder that.

Nearly ten thousand years ago, we find examples of tokens for tracking goods, made out of clay.

These evolved into the first bit of writing we have - from around 3300 BC.

An inventory.

Markings incised on a tablet of wet clay, left to dry in the sun, became a permanent record of how many bushels of grains had been counted.

Not long after, we see the first ledger, then the first receipt - with the first name we can identify, because that receipt assigned goods to an owner.

In writing.

Writing doesn't create civilisation, but it does amplify civilisation by making commerce a precise enterprise. When you know how many bushels of grains you have, you can trade them. You can price your inventory. Give your customers receipts.

These aren't nice to haves: they're necessary elements to de-risk commerce. Without writing, commerce would depend on imperfect human memory; a castle built on shifting sands.

Civilisation arcs upward from Sumer in part because writing grants commerce the stability to begin amplifying human wealth.

The advent of writing is the beginning of commerce as we understand it, because it took imperfect human inputs and transformed them, through writing, into reliable, trustworthy foundations for commerce among strangers, at scale, beyond the reach of memory.

That's precisely the rigour we need to bring to our agents - transforming them into reliable, trustworthy elements for modern business.


To get there from here, we're going to need to recall that we're already very good at verification.

How can we know something is true? How can we inspect that truth? How can we test it?

Look to the tablets, the double-entry accounts, the audited business statements: We have millennia of practice in verification, nearly all of it directly and immediately applicable to agents.

Many times agents fail because, in effect, they get to 'grade their own homework'. Every agent runs a 'reflection' cycle after it completes an action, to satisfy for itself that the action worked. Until it knows the action completed successfully, it can't go on to the next action.

But can the agent know for sure? Will it mistake wrong for right - or even cheat the exam? All of these behaviours have been seen in agents, so we need to build for agents that can be imperfect in these ways. Just as people are.

This means agents need hard checks, verifications built into their operations, that the agent can not move beyond until the verification has been satisfied.

Just as agents should never grade their own homework, they should never build the systems that verify them. Verification needs to be done independently - and adversarially.

The verifier has to press the agent, testing every assertion, questioning every result. The agent has to be able to provide sufficient rationale to meet the standard of the verifier. Otherwise, the agent's work product will be dismissed as unreliable and risky.

We already do this with people. We know how. These practices are so old, so well understood and so necessary that we really don't think about them much - other than to note that we'd be taking ridiculous risks to operate without them.

So we don't name them. They're just what we do.

Now that we're bringing these practices to agents, we need a name.

VERIFICATION DESIGN.

Verification Design takes the best of human practices and brings them to artificial intelligence. It's the 'secret sauce' that's always been public, a necessary ingredient to launch this civilisation into its own upward arc, amplifying human wealth.

Verification Design is the flywheel AI needs to turn its imperfect raw material into a productivity multiplier the likes of which haven't been seen since the coming of steam power, three centuries ago.

As we bring Verification Design to our agents, the most important thing to remember is that we already know how to do this. It's in our cultural DNA. We're simply finding a new expression for one of the earliest things we learned: get it in writing.

Subscribe to The Watershed

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe