Why I'm Building Outcome
First in a series on how we're building Outcome, and why.
A few years ago I watched someone I respect run a decision that mattered — real budget, real people, a deadline that wasn’t going to move — using a spreadsheet, three browser tabs, and a phone call to the one person who actually understood the data. She got it right. But she got it right despite her tools, not because of them. The information she needed existed. It was just scattered across systems that didn’t talk to each other, in formats nobody had time to reconcile, behind a wall of effort too high to climb before the deadline.
That gap — between the data an organization already has and the decision it actually needs to make — is the thing I’ve come to Outcome to close. This is the first of a few posts about how we’re going about it. I want to start with the part that matters most, which isn’t a feature. It’s a belief about how this technology should be used by the people who serve the public.
You can’t run a county on a chatbot
There’s a story the industry is telling right now, and it goes like this: the models are getting good enough that you can point one at your hardest problems and let it run. Spend more, get more. Let the agents loose.
For a lot of consumer software, that’s fine. The cost of a chatbot being confidently wrong about a dinner reservation is a bad dinner. But the people we build for don’t live in that world. When a county allocates emergency resources, when a public defender decides which cases need which evidence, when a nation watches its coastline — being confidently wrong isn’t an awkward moment. It’s something a real person lives with. And a system that sounds certain while being wrong is worse than no system at all, because it launders a guess into something that looks like an answer.
The deeper point will outlast whatever we’re calling these tools this year: an institution can’t delegate judgment to probability. The word “chatbot” might sound quaint in five years. The problem won’t. As long as the consequences of a decision land on real people, someone accountable has to be able to stand behind it — and “the model was pretty sure” has never been a thing anyone could stand behind.
So we started from a different question than “how powerful can we make it?” We started from “how do we make something a serious person can actually trust enough to act on?” Those turn out to be very different design problems. The first one points you toward bigger models and more autonomy. The second one points you somewhere the hype isn’t looking yet.
I think that’s the real fork in the road for this whole industry, and almost everyone has taken the other branch. Most AI platforms begin with the model — here’s what it can do, now where can we apply it? — and work outward toward your problem. We began at the other end, with the institution: how do the people who serve the public actually earn trust, divide authority, and stand behind a decision — and how do we make that faster without breaking the parts that make it legitimate? Start from the model and you build something that asks the institution to bend around it. Start from the institution and you build something that bends around it instead. That one choice, made at the very beginning, shapes everything that follows.
The most sustainable line of code is the one you don’t have to run
Here’s the contrarian part, and I’ll say it plainly: throwing a large language model at every problem is usually the wrong tool, and often the expensive, fragile, untrustworthy one.
Most of what an organization needs from its data isn’t a feat of language understanding. It’s arithmetic, aggregation, comparison, lineage — the boring, provable work that traditional analytics has done reliably for decades. When you do that work with plain, deterministic methods, you get something precious: an answer you can check. It computes the same way every time. You can trace it to its source. No one has to take it on faith.
So that’s what we do for the heavy lifting. We reserve the large models for the genuinely hard part — the places where you actually need to read meaning, weigh ambiguity, draft a recommendation a human will judge. And even there, we don’t let the model have the last word with itself. The system keeps checking its own conclusions against the data they came from, and as that data changes, the picture changes with it. The goal isn’t an oracle that pronounces. It’s a partner that does the legwork and shows its work — and is the first to tell you when the work no longer holds.
There’s a practical reason this matters beyond trust, and it’s one I care about a lot: precision is also what makes this affordable. A platform that’s genuinely lean to run is a platform that can be priced for a rural county or a small nation, not just a Fortune 500. The organizations with the hardest problems and the smallest budgets have been priced out of serious tooling for a long time. Building efficiently isn’t only good engineering. It’s how you put this capability in the hands of the people who’ve never been able to afford it.
The machine helps you see. You decide.
The last belief is the one I’d want a public servant to hold us to.
The good institutions — the ones that outlast the people who run them — share power on purpose. A council sets the rules. Someone acts within them. A second person signs before anything consequential goes through. Someone independent gets to watch. That’s not bureaucracy for its own sake. It’s how a community keeps power answerable.
We’re building Outcome to work the same way. The system can help you see faster and move faster, but the authority to decide stays with people, and it stays distributed across people — not concentrated in a model, and not concentrated in whoever happens to own the model. Every recommendation traces back to its source. Every consequential action carries a name beside it. The technology earns its place by making good judgment faster, not by replacing it.
I’ve started to think of this as the difference between two architectures, and it’s worth naming plainly. One is plenary: all the information flows up to a single authority, that authority decides, and the decision flows back down. It’s a clean design, and it’s the one most AI is quietly built toward — one system that sees everything and decides everything. The other is constitutional: different authorities hold different responsibilities, different permissions, different oversight, and they coordinate from a common operating picture toward action no single one of them could take alone. The thing worth noticing is that the plenary model is the wrong shape whether the authority at the top is a model, a CEO, or a central command — and that almost no institution that has to answer for itself actually runs that way. A hospital doesn’t. A bank doesn’t. A court doesn’t. They distribute authority on purpose, because concentrating it is a risk, not an efficiency. We built Outcome for the second architecture, because that’s the one real institutions actually live in. I’ll give that idea its own post before long, because I think it’s the most important design decision we’ve made.
There’s a word for this that I want to reclaim a little. Everyone in this industry talks about sovereignty now, and they almost always mean one narrow thing: where your data physically sits. That matters, and we do it — your information stays in your environment, full stop. But it’s the shallowest version of the question. Sovereignty isn’t really about where the data lives. It’s about who stays in control — and control isn’t one thing, it’s four. Your data stays yours. Your rules govern what the system does with it. Your people own the decisions that carry consequences. And your own structures of approval and oversight govern what actually gets done. Lose any one of those and the others start to feel like a courtesy.
The test of all this is whether the boring guardrails hold at the exact moment the system is most sure it’s helping. A recommendation can’t bypass procurement policy because it looks urgent. A sheriff can’t reach sealed records because the model decided they were relevant. A disaster doesn’t suspend oversight — it follows the emergency authorities the institution already defined. That’s the difference between a tool that respects an institution and one that quietly asks the institution to get out of its way. I’ll write more about what that looks like in practice in a later post, because it’s the part I think the rest of the field is going to get wrong.
I’ll say the harder version of this plainly, because it’s why I took the job. For years now, the people with the most power have had the best tools — systems that vacuum up enormous amounts of information, model the world and the people in it, and produce confident, stochastic judgment calls that never have to justify themselves. The leaders those tools serve are increasingly answerable to no one. And whether the thing whispering to a powerful person is a sycophant on the payroll or a model trained to tell them what they want to hear, the failure is the same: the answers drift further and further from reality, and the people who actually know — the expert, the field officer, the person closest to the work — get overruled in favor of whatever just came out of Silicon Valley. Concentrated power plus unaccountable, unjustified judgment is a house of cards. It looks like strength right up until it doesn’t.
Some companies build for the people at the top of that structure — the ones answerable to no one. We build for the other side of it: for the institutions that have to still be standing, and still be accountable, when the house of cards comes down and the real work of rebuilding begins. That’s not a doom prophecy; it’s a bet on who’s still doing serious work in five years. We think it’s the people who never stopped being accountable to anyone.
But here’s the part that surprised me, and it’s why I’m actually optimistic. Those people — the ones doing the most consequential work — mostly aren’t using these tools at all yet. Maybe they’ve had a model write them a Python script. Otherwise they’re getting it done the way they always have: a spreadsheet, a fax machine, a phone call to the one person who actually knows. And they’re right to be cautious. Nobody has handed them something they’d be willing to stake a real decision on. The loudest part of this industry has been busy spending money to look busy; the people with something real on the line took one look at the slop and quietly kept using the spreadsheet. That’s not them being behind. That’s them being correct.
So that’s who we actually compete with. Not a flashy demo. A fax machine, a spreadsheet, and a phone call. The toolchain I described at the top of this post — the scattered systems, the wall of manual effort, the one person who understood the data — that isn’t history. It’s still how a great deal of the most important work in the world gets done. That’s the real incumbent. And that sets the bar exactly where it should be. The question was never “is this smarter than last quarter’s model.” It’s “is this trustworthy enough to put down the spreadsheet.” That’s a much harder bar, and it’s the only one that matters to the person whose name is on the decision.
I think a lot of this industry is about to spend a few years learning, the expensive way, that capability without accountability doesn’t survive contact with a real institution. We’d rather start there.
What’s next
This is the first post in a series. In the ones to come I’ll go deeper on the parts I’ve only gestured at here — how we keep data in your hands instead of ours, what it actually means to build separation of powers into software, and the engineering behind doing more with less. I’m writing them partly to be useful and partly to be held accountable: it’s easy to write down values, and harder to build them into a product. We intend to do the harder thing, in the open, and to show our work.
If you do this kind of work — the kind where being right matters because someone lives with the consequences — I’d genuinely like to hear what you’re up against. That’s the whole reason I’m here.
— James