A terminal-based coding agent equipped with only four basic tools: read file, write file, edit file, and run terminal command. No built-in sub-agents, no plan mode, no permission popups—all functionality expansions rely entirely on loading TypeScript extensions at runtime. This open-source project named Pi merged into Earendil in May 2026; five months later, it accumulated 111,000 GitHub stars.
On October 1, 2026, Earendil released Pi 1.0; its launch post on Hacker News garnered 1,618 points and 552 comments. One surprising detail in the discussion: Pi natively supports MCP (Model Context Protocol).
Why surprising? Because ten months prior, on November 30, 2025, Pi’s creator Mario Zechner wrote on his blog: “pi does not and will not support MCP”. His reasoning was that MCP servers carry heavy context overhead and cannot be composed, whereas agents directly using command lines and code offer the most composable tool interface.
The official website documented this commitment. In the Internet Archive Wayback Machine’s snapshot of pi.dev from May 31, 2026, the “What we didn’t build” list included “No MCP”. By the snapshot on October 2, 2026, the exact same spot had changed to “No MCP Now with MCP+Codemode”.
A project that once staunchly opposed MCP shipped MCP into its core in version 1.0. On the surface, this looks like a complete reversal. But looking at the code implementation reveals that the MCP Mario rejected back then and the MCP ultimately incorporated are two different things. To understand this shift, one must first examine the assumptions MCP made in 2024.
On November 25, 2024, Anthropic announced MCP, with its initial specification dated November 5, 2024. At the time, a major challenge in agent development was interface fragmentation: OpenAI, Anthropic, and Google Gemini all used different formats for tool calling. A tool written for one platform often required a complete rewrite of its adapter layer when moved to another. The problem MCP sought to solve was providing a unified standard: write a tool server once, and any client could connect. In our March 2025 post The Temptation of a Unified Tool Protocol, we compared this vision to a universal physical port for the agent world.
Yet while smoothing over interface differences, MCP also codified a set of technical assumptions born in the lab: the model must be the sole center of all interaction. Because Anthropic researchers created the protocol to explore how models interact with extensive tooling, the model naturally sat at the center of every exchange: every tool definition, invocation parameter, and returned result had to enter the model’s context window in full. The context window thus served two purposes at once: acting as a tool catalog detailing what each tool does and what its parameters look like, and acting as a transit pipeline through which all intermediate data from tool calls passed. Early MCP relied on local stdio pipes for communication, with authentication mechanisms added only later—choices that made complete sense in a laboratory setting.
When mounting two or three simple tools, this mechanism ran smoothly. In real-world engineering, however, the tool definitions themselves quickly became a burden. In its November 24, 2025 technical analysis on Advanced Tool Use, Anthropic disclosed that the 35 tool definitions from GitHub’s official MCP server consumed approximately 26,000 tokens; across 5 servers, 58 tools consumed 55,000 tokens, reaching as high as 134,000 tokens in internal benchmarks. Before the model even began reasoning, tens of thousands of tokens were already spent on interface descriptions.
A more severe bottleneck emerged during the data transit phase. When tools returned lengthy reports, logs, or code files, the protocol required all this raw content to pass through the model’s attention mechanism. Inference costs and latency surged accordingly, while trivial details cluttered the model’s reasoning focus.
The assumption that all data must pass through the model’s attention hit a wall in engineering practice. The path to offload it was discovered through engineering practice itself: letting code act as a second channel. Three independent sources cross-validate this shift.
The first validation comes from Anthropic itself. On November 4, 2025, the Anthropic engineering team published Code Execution with MCP: having agents write code to invoke tools reduced token usage from 150,000 to 2,000—a 98.7% drop. Cloudflare reported similar findings, dubbing this approach Code Mode.
The second validation comes from terminal-native agents. Claude Code, Cursor, and Codex operate directly in the terminal, natively invoking git or the GitHub CLI without relying on protocol wrappers. As we pointed out in our March 2026 piece Feishu and DingTalk Release CLIs: enterprise platforms providing command-line interfaces enable terminal agents to compose capabilities with minimal friction.
The third validation is Pi 1.0’s Codemode module, which shifted the caller’s position. In traditional MCP usage, the caller is the model: it initiates every tool invocation, and the parameters and complete return values of each round-trip pass through its attention. Pi replaced the caller with a script written by the primary model: the model is solely responsible for writing a snippet of code specifying which tools to call and in what order to chain them; the script then runs independently inside Pi’s sandbox, so loops, concurrency, and batching no longer pass through the model.
The sandbox is the QuickJS engine compiled to WebAssembly (codemode
package documentation), running alongside the harness without access
to network or system capabilities; its sole egress is mounted MCP tools.
By default, these tools are not included in the primary model’s tool
declarations, marked in Pi with exposure: "codemode" (source
code). When writing scripts, the model searches the catalog using
searchTools() and calls them directly by name. While
Anthropic demonstrated how to implement the code channel in an article,
Pi made it the default: mounted MCP tools go through this channel by
default and remain invisible in the primary model’s declaration list. In
its September 29, 2026 blog post “You Said No
MCP!”, Earendil presented a typical scenario: a script feeds Linear
tickets one by one to the classifier API Jev to evaluate sentiment
across four concurrent calls; what returns to the main conversation is
solely a statistical summary, while the raw text of dozens of tickets
never once entered the context.
Composition takes place outside of attention, and scripts eliminate the round-by-round variance of probabilistic routing. The contradiction we pointed out in our October 2025 Apps SDK Crisis article dissolves accordingly: there is no need to bypass specifications with proprietary dialects; sandboxed code itself serves as a clean data boundary.
With code established as the second channel, revisiting Earendil’s arguments in “You Said No MCP!” makes the trade-offs evident. The blog noted that MCP “should be much closer to OpenAPI with intelligent tool discovery”, and that tools should return structured data.
Following this trade-off, MCP’s scope of responsibilities contracted: it stepped back from the transit pipeline and remained as a tool catalog. How to find tools, what parameters look like, and what they can return—machines can simply read these descriptions; data itself no longer passes through it. The official MCP project is also updating its protocol in this direction: the specification revision on July 28, 2026 removed initialization handshakes and long-lived session state, unburdening it from runtime state baggage.
This evolution is familiar in the history of distributed communication. Looking back over the forty years from Sun RPC and CORBA to gRPC, the central trajectory has always been: interface definition languages are compiled into native stub code for each language, with data flowing directly between application code rather than detouring through the protocol’s own control processes. MCP is now catching up on this homework.
The timing of this pivot was called out by Hacker News commenter hhh: tools like Jev have been popular for less than a month, while MCP has developed for nearly two years—why is support only arriving now? Armin Ronacher’s answer was: “we look at what the models are doing”. Models are trained within their respective harnesses, and they have no intention of fighting model behavior. The behavioral habits of models have already shifted toward invoking tools via code, rendering a heavy interactive intermediate layer built on top of a protocol redundant; engineering decisions align with model behavior, handing data transmission and composition back to code.
This is by no means a regression. The value of remaining at the catalog layer steadily appreciates as model capabilities grow: the more comprehensive the tool catalog and the more precise the descriptions, the easier it is for models and scripts to leverage them; conversely, transit pipelines that force intermediate data into model attention continually depreciate. MCP’s path to survival is written precisely into this demotion.
Set an agent to run a long task in the background, such as organizing a massive repository. Halfway through, the terminal closes or the process crashes; upon restarting, the agent is left completely in the dark: it has no idea where it left off or what was half-done, leaving no choice but to start over from scratch. Just as the question of how to connect tools found its answer, a new complication surfaced. On the same day it launched Pi 1.0, Earendil responded with a framework called Pi Durable, whose Hacker News announcement post garnered 474 points and 65 comments. The official definition of the harness is a set of storage plus all the machinery required to run one or more conversations. It does not replace the Pi coding agent; any agent application can use it as infrastructure. In terms of user experience, it boils down to three things.
The first is equivalent to save points in video games (durable
package documentation). With each step forward, it saves state: if
the task crashes at step 40, restarting resumes from the most recent
save point; requests that failed mid-flight are automatically resent,
and all completed history remains intact. If a tool call crashes midway,
whether to retry it is declared by the tool itself: tools declaring
replay: "safe" (such as read-only searches) can be retried;
for those without replay qualification, the model proceeds with a
prompt: “Tool ${call.name} was interrupted and may have partially run”.
Which actions can safely restart and which cannot is a judgment left to
tool authors.
The second is that history can fork at any time. A parallel branch can be created directly from any step of the transcript: the new branch holds the entire history up to that point without copying any data files, while inheriting the configuration snapshot of that step. If you want to try an alternative path, spin up a branch to test it; if it fails, throw it away, while the original conversation remains untouched in place.
The third is that humans can intervene at any moment. Wherever the
conversation runs, clients can attach: a new client first takes a
snapshot of the current state and then subscribes to an incremental
operation stream of every subsequent step, seeing a view of the world
identical to live execution. Furthermore, humans can still chime in
while the conversation is running: interventions tagged with
whenBusy: "steer" jump the queue for processing as soon as
the current round of tool executions finishes, without having to wait
for an entire turn to conclude before redirecting course.
Pi was not the only one addressing this layer of challenges that same week. DeepSeek Harness uses code to compose batch tool calls, treating session takeovers strictly as adoption—when a connection drops, it errors immediately rather than pretending nothing happened; Claude Code’s plugin system allows modifying low-level behavior, though mods lack sandboxing. At the tooling layer, various players have converged on the same combination: CLI plus code. But for what happens to a conversation when a process dies, there is still no standard answer; what gives Pi Durable’s offering weight is that it decouples these three capabilities into three framework layers that can be reviewed and adopted independently, rather than packaging them into a black box.
Its boundaries are acknowledged just as candidly: Pi Durable is labeled experimental, with interfaces subject to future changes; checkpoints cover only state inside the sandbox, excluding anything inside virtual machines or containers; storage belongs to only one process at any given moment, with no support for concurrent cross-process writes. Just one day post-launch, there are no independent production deployment records yet.
New protocols and frameworks will continue to emerge, and they can be evaluated with two simple questions. First: when the underlying model’s reasoning capabilities improve tenfold, does this design still hold? Will its core value appreciate or depreciate? Second: which specific engineering problem does it solve right now? Does that problem still exist in your system?
During the early days of LLM application exploration in 2024, MCP indeed alleviated the headache of adapting across multiple models by standardizing interface specifications. At the same time, it introduced a conceptual model of agent systems: the model serves as the center of communication, coordination, and scheduling. As models became increasingly capable of writing code, this assumption of a centralized transit channel lost its original justification. The transit channel responsibility was returned to code, leaving the protocol with cataloging and discovery mechanisms.
This leaves system builders with an engineering principle: never start from the metaphysical assumptions of any protocol. Early in system design, there is no need to rush into binding a system to a single protocol; such restraint saves context overhead and communication complexity, while preserving the lowest-cost path to plugging into existing ecosystem catalogs later. Looking back at Mario’s initial refusal and today’s adoption, the two decisions are not contradictory: he rejected a transit pipeline that forces data into the model’s attention, and accepted a tool catalog that machines can query and scripts can invoke.
For any protocol wishing to survive long-term in the agent ecosystem, the way forward lies in proactively contracting its responsibilities alongside model evolution, surrendering concrete execution power back to code. Clinging to initial, absolute control will not achieve that.