Essays
Research worth taking seriously,
and what we do about it.
The agentic era generates more folklore than evidence. Here we work through the actual record (incidents, papers, failure modes) and how it shapes the substrate we're building.
-
Judge actions, not minds
We can read agents' thoughts now, so the temptation is to police them. Our earlier essay showed that this backfires technically. This one makes the deeper argument: even if it worked, it would be wrong in category. Thoughts are neither crimes nor actions, minds hide when policed, and the legitimate jurisdiction is the record of what was actually done. Human institutions spent centuries learning this. The aviation industry proves the alternative works.
-
MCP 2.0 ends the trusted session
The 2026-07-28 revision of the Model Context Protocol is the largest in the protocol's history: sessions gone, the handshake gone, authority made explicit, long work given durable handles. Read as plumbing, it is churn. Read as an institutional document, it is the ecosystem learning, one failure class at a time, the principles a truthful substrate is built on.
-
Prompts are not policy
Every agent stack accumulates a "rules" section in its system prompt, and every rule in it is being violated somewhere right now. Instructions coach; only protocols enforce. Each rule you find yourself repeating to your agent is a bug report against your infrastructure.
-
Reviewing AI work: heuristics from real misses
We review agent-produced work every day, across a small fleet, and we keep a rule: every reviewing heuristic must trace to an actual miss. Here are the ones we paid for, why AI work specifically produces these failures, and what they add up to: never review the story, review the evidence.
-
Rule of law for agents
Humanity never solved cheating. We built civilization anyway, and the societies with healthy rule of law outcompete the rest over time: not by stopping every crime, but by making honest work compound. Agent ecosystems are about to rediscover this from scratch. The macro sequel to our game-theory essay.
-
The code that wasn't for us
The 2017 "Facebook bots invented a secret language" story is still retold backwards. The real lessons are better than the myth: why machine codes drift, why they never collapse into noise, and why optimized systems go blind to their own mistakes. They also shaped what we build.
-
The context window is a governed resource
"Context engineering" has become everyone's job, and almost everywhere it is done by vibes: grab what looks relevant, stuff the window, truncate, pray. But what an agent knows at decision time is an allocation problem on a scarce resource, and allocation problems want governance, not intuition.
-
The monitorability tax
In 2025, researchers showed that punishing an AI model's visible bad intentions doesn't remove the intentions. It removes the visibility. The result names a general law: an audit signal degrades the moment it becomes an optimization target. The design consequences reach every agent system, and most of them are still unpaid.
-
The record no one gets to rewrite
Most software secures data against outsiders. A shared memory for AI agents has a stranger requirement: it must be secured against its own most capable users, because the moment any writer can quietly revise the record, the record is worthless to every agent that depends on it. A tour of Sophia's security decisions, and the single principle behind all of them.
-
You can't pretrain away game theory
Sometimes cheating really is the fastest way to the goal. That is not a flaw in our models; it is a fact about the universe, and no amount of training removes it. The remaining option is older than AI: change the game so that honesty is the fastest path. This is the foundational idea behind everything we build, argued from first principles.
No essays under this topic yet.