Home What's Possible The Audit Services Pricing Frontier Notes About Frontier Notes Start with an Audit
← All Notes

For AI to be useful, it has to be able to do things. Sounds obvious, right? A person using Claude to gather campaign data is asking it to do something. An agent putting a report together overnight is doing something. Take away the doing and what is left is a conversation.

But it also has to do things right.

That is both the challenge and the promise of building a successful AI ecosystem. Get it right, and it feels like magic. Get it wrong, and no one will trust or use anything. Something that works once is a demo, not a tool.

What follows is how I think this should be done, most of the time. It’s how we build, and it’s a good example of our broader philosophy of how AI systems should work. None of it is groundbreaking. But it is surprisingly rare.

Ask or act

To build these right, we must first recognize that there are basically two ways an AI (a chatbot, an agent, a system of agents) can use the tools we give it:

  • They can request some kind of information.
  • They can perform some kind of action.

I’ll refer to those as reading and doing below, although of course real requests are often a little of both and the distinction is much more nuanced than that.

An AI is reading when it wants to know how a client’s campaigns performed last month. It’s doing when it drafts next month’s campaign or changes a budget cap. Doing each of these properly requires two completely different, in fact polar opposite, approaches.

Reading: give the right information

The wrapper problem

You’ve connected Claude with your time tracking tool. Great! Then you ask a real question, like whether a client is going to run over budget this quarter, and it all comes crashing down.

A lot of off-the-shelf MCP tools are simply wrappers to APIs, the existing programmatic endpoints that software can use to call a service. Through a wrapper like that, your question becomes a call like:

GET /time_entries?client=1041&from=2026-07-01

and back come thousands of time entries.

Three problems arrive with them:

  • The context floods. How many discrete time entry records are there for one client? Hundreds? Thousands? APIs, by design, return a lot of data. That’s fine for an application, which takes the two fields it needs and ignores the rest. A model reads every single thing it is given, so the conversation fills with raw records before any actual thinking has happened.
  • The math is done by hand. Months of arithmetic, performed by a chat model in the middle of a conversation, which is the absolute worst place to do it. Ask again tomorrow and the answer can come out different, with nothing visibly wrong either time.
  • Even after all that, the data doesn’t hold the answer, and Claude doesn’t know it. Not one entry says “budget”; that lives in another system. Billable and non-billable hours look the same unless you know the rules that separate them. The answer is a trajectory computed from months of entries, and it isn’t sitting in any field.

The worst part is that you’ll get an answer anyway. Claude doesn’t know what it’s missing, so it works with what it has and hands back a confident, plausible number. There’s no error message and nothing to flag. You take it into a meeting, and it’s wrong in a way nobody can see.

And no matter what you layer on top, the wrapper itself never learns. It stays exactly as smart as the day it shipped.

A better way to do it

It starts from one governing rule: produce and return the smallest amount of the most relevant material. Everything extra you give a model is an excuse to get something wrong.


The key architectural point is that the ONLY thing the AI ever sends to the tool is the question. All the processing happens down the line. You can think of it as the question kicking off a team of experts:

  • The intent interpreter, who works out what’s actually being asked and what would fully answer it, then hands it to…
  • A team of code writers, who write real code, on the spot, using (say) that API we talked about, but with programs purpose-built to answer the actual question. Every program goes first to…
  • The inspector, who reads each one before it runs. Anything that fails the check never executes, and what passes runs in a sandbox where it can read but not touch, cut off at a hard time limit. The results go to…
  • The auditor, who compares what the writers produced, judging the numbers themselves rather than the style of the code, and picks the best one.

The team can also refuse. When a question is too big to answer reliably, it says so and asks for the question to be broken up. And all of it happens invisibly, in a matter of seconds, so that whoever asked gets exactly what they need, correctly, and nothing more.

And here is the coolest part. If you’re building this tool for your own company, or for a client, you can teach these experts exactly how that company works: which hours count against the budget, which number the client actually watches. That knowledge is far more effective here than in a skill riding on top of a chatbot. Written as clear, direct instructions, as close to the source of truth as they can get, it makes every single tool call as relevant as possible, for everyone and everything that calls.

So, back to the question we started with. Ask this tool whether the client will run over budget this quarter, and what comes back is the answer, not thousands of rows:

{
"as_of": "2026-08-31",
"budget_hours": 1200,
"hours_spent": 890,
"quarter_elapsed": "62%",
"budget_consumed": "74%",
"projected_finish": "106% of budget",
"primary_driver": "design revisions"
}

That’s what the chatbot gets back. Exactly what it needs, and nothing it doesn’t. The flood of records and the hand math happened where they belong, below the conversation, and all that surfaces is the answer.

Doing: do the right thing

Everything so far has been about asking. Now the AI wants to change something. That could be an action (publish this thing) or writing (change this data) or both, and either way, we’ve left the world of wrong answers and entered the world of wrong outcomes, the kind that can’t be fixed by asking again. Ambiguity is dangerous here in a whole new way: whatever the model composes goes to the platform live, on the first call.

Three questions decide this side. Does the AI know how to do the right thing? How do we protect against it doing the wrong thing? And how do we give it the context it needs to make the best decision possible, and to self-audit what it’s about to do?

The key architectural point here is the exact opposite of before. EVERYTHING IS STRUCTURED.

The menu is fixed ahead of time: every action the tool can take was named by a person, and nothing is composed on the fly. If nobody defined the action, there’s no way to ask for it.

And before taking any action, the model has to walk through a carefully controlled progressive reveal:

  • The contract is the instruction manual for one action, and the AI has to ask for it before doing anything else. It spells out exactly what the operation requires: which fields, what formats, what’s optional, what the platform will and won’t accept, and the order things have to happen in. The AI builds its request against that spec, so by the time it asks for anything real, the request already has the right shape.
  • The preview is a dry run. The AI’s first call carries its full, finished request, but nothing executes. Instead, the tool answers with a plain statement of what it’s about to do: create this campaign, in this account, with this budget, in a paused state. Does this look okay? The same checks that guard the real operation run against the preview, so any problem surfaces now, while it’s still just words. The reviewer, whether that’s a person reading it or an agent comparing it against its instructions, catches the mistake here, and the AI revises and previews again until it’s right.
  • The go is a separate act, given only after that review, with ownership checked at that exact moment. For anything that matters, decisions stay with people.

Underneath all of it, the guards aren’t rebuilt inside each action. They’re inherited from the shared path every action runs through, so there’s no forgotten check waiting to be found. And everything is recorded, with alerts rationed, because a system that pages about everything gets ignored.

When this all works, it’s quick and quiet. You ask for next month’s campaign, you read a preview of exactly what will happen, and you say go. The decision never left your hands, and every step is in the log.

Where the magic comes from

The two approaches could not look more different, and they are the same idea. All the real work, the judgment, the math, the checks, the guardrails, happens below the conversation. What surfaces is small and clean: an answer you can act on, a preview you can approve.

That’s also what makes these tools more than tools. Most of the current thinking puts the discipline in the harness, the layer wrapped around the model that guides it and checks it, and a harness protects exactly one consumer. I’d put the discipline lower. A piece that carries its own judgment and its own limits is a building block: nothing stacked on top of it needs to re-supply its safety, so a chatbot can lean on it today, an agent can lean on it unattended at three in the morning, and whatever we build next starts from there instead of from zero. The more atomic each piece and the more discipline it carries inside itself, the more safely the whole thing stacks. That’s how an AI ecosystem grows without growing fragile.

It is most of the work. Making a tool exist takes a fraction of our time; making it give the right information and do the right thing takes the rest. And better models won’t shrink that share, because a better model makes wrong answers more convincing, not rarer. This layer is the design, not a phase.

But get it right, and nobody ever sees any of it. Someone asks, and the right answer comes back. Someone says go, and the right thing happens. It feels like magic.