AI AgentAI Products & PlatformsIndustry & Competition

Building a Manus in 2026: What Infrastructure You No Longer Need to Hand-Roll

Open Manus, type a task into the prompt box: research evaluation benchmarks for generative video models over the past month, compile an analysis report with a comparison table, and hit Enter. On your screen, everything looks quiet, save for real-time thought logs unfolding and draft text continually updating. But in the server rooms behind the glass, an entire mechanical system is rapidly grinding into gear.

Hand a task to Manus, and what happens next feels familiar when broken down step by step. According to mechanisms disclosed in the official Manus blog post, Understanding Manus sandbox - your cloud computer (accessed 2026-09-20), once the system receives an instruction, it spins up an isolated cloud virtual machine sandbox specifically for that task. This is a cloud computer equipped with dedicated compute, an independent network stack, an isolated file system, a browser, and a general-purpose toolchain. Tasks remain isolated within their virtual machine environments, never interfering with one another as they advance independently in parallel.

What Happens Behind a Manus Task

Inside this newly provisioned virtual machine, the model enters a tightly wound core loop. As the Manus team reviewed in their official technical blog post, Context Engineering for AI Agents: Lessons from Building Manus (accessed 2026-09-20), the logic of this agent loop is compact: after receiving user input, the model selects an action from a predefined action space in each iteration based on the current context. That action executes inside the VM sandbox to produce an observation. Both the action and observation are then appended to the end of the context and fed into the next round of model decision-making. This cycle repeats until the system determines the task is complete.

In the real world, this loop runs heavy. According to Manus, a typical complex task consumes around 50 tool calls on average. Over dozens of iterations, as context length grows, large language models tend to suffer from attentional drift, gradually forgetting the initial objective or looping inside dead ends. To counter this degradation of attention, Manus adopted an engineering tactic to steer attention inside the sandbox. It maintains a todo.md file on the virtual machine’s disk, repeatedly checking and updating this checklist during execution to cross off completed steps. The team calls this mechanism recitation: by continually restating the overarching objective at the very end of the context, it forces the overall plan back into the model’s active attention zone, preventing it from getting lost in ultra-long contexts.

Along with those 50 tool calls comes high-frequency throughput of massive, heterogeneous data. Verbose DOM trees from web scraping, dozens of downloaded PDF pages, and error stacks from terminal commands would quickly exhaust even a 256K or larger window if dumped wholesale into context, driving up inference costs and introducing painful latency. Manus treats the file system as the ultimate context, taking advantage of physical disk properties: near-infinite capacity, natural persistence, and direct read-write access for the agent. The model learns to offload intermediate artifacts to disk files, retaining only file paths or source URLs in the context window. This compression strategy preserves recoverability, allowing the model to shrink context size significantly without dropping critical facts.

When execution hits an unexpected snag, a command line error, or a web page timeout, the system keeps the failed execution trace and error messages intact within the context. It does not erase them, nor does it rely on sampling randomness to retry blindly. In the view of the Manus team, confronting real-world friction and demonstrating self-healing capability is precisely what separates genuine agent behavior from brittle scripts.

A Manus task passes sequentially through five stages: task input, cloud VM, agent loop, file memory, and delivered result; missing any single stage prevents completion

Stripping away the product shell, the underlying engineering required to sustain such a task boils down to a clear procurement checklist:

  1. Isolated Compute: Providing an independent, secure compute environment and sandbox for each concurrent task, complete with its own file system, network, and runtime.
  2. Session Affinity and Intra-Task State: For a long-running task spanning dozens of calls and over ten minutes, subsequent requests must route reliably to the same machine, maintaining memory and disk state steadily throughout execution.
  3. Cross-Task Long-Term Memory: When a user starts a new task three days later, the system must retrieve historical preferences and background context. This memory must exist independently of any single task’s ephemeral compute node.
  4. Tool Integration and Credential Management: Tasks must interface with external search, enterprise services, and various APIs. The sandbox must securely store user-authorized credentials and private API tokens without exposing plaintext secrets.
  5. Observability: Across an automated chain spanning dozens of steps, ops systems must maintain full-link visibility into which step stalled, which tool call returned an error, and how context consumption and latency were distributed.

These five requirements land squarely on some of the hardest problems in modern distributed systems. In 2024, almost no off-the-shelf cloud service packaged solutions for these challenges. By 2026, the landscape is shifting substantially.

In 2024, You Had to Build This Checklist Yourself

Looking back at the path taken by the first generation of general-purpose autonomous agent products, pioneer teams paid a heavy price for this foundation. The Manus team confessed in their official technical blog that they tore down and rebuilt their agent framework four full times just to find a workable context layout and scheduling logic. They coined a self-deprecating term for repeatedly groping through architectural unknowns, tweaking prompts, and debugging engineering issues: “stochastic graduate student descent,” likening it to a graduate student trying their luck through trial and error.

Rewriting the framework was merely code logic; managing the lifecycle of the underlying sandboxes was a much harder battle. According to Manus blog posts, the team had to manage a complex cloud VM scheduling system. Sandboxes had to support on-demand provisioning, automatic sleep when users stepped away to conserve resources, and automatic wake-up when users returned. Continuously idle sandboxes were automatically reclaimed, with a retention window of 7 days for free users and 21 days for pro users.

After a sandbox was reclaimed, if the user revisited a historical task, the backend scheduling system had to spin up a new sandbox and selectively restore final deliverables, user-uploaded attachments, and generated web pages or slide decks, while leaving behind intermediate scraping scripts and temporary logs. If a virtual machine suffered an unrecoverable system failure mid-task, the backend also had to automatically provision a fresh sandbox to take over.

From Manus’s official launch in March 2025, to Meta completing its acquisition on December 29 of the same year (publicly reported at over $2 billion), to its August 2026 blog post, A Note to Our Users (accessed 2026-09-20), announcing an impending spin-off back to independent operations along with regulatory data purges, the trajectory of this first-generation pioneer left a clear trail. Back then, cloud provider shelves were empty. Developers had to hand-craft everything from low-level VM lifecycle orchestration, network isolation, and storage state machines all the way up to the user interface.

The turning point arrived in 2026. After two years of production trial and error, public cloud providers began standardizing this custom foundation and putting it on their shelves.

Among these commercial offerings, Amazon Bedrock AgentCore from AWS is one of the most complete industrial examples assembled against this checklist. According to AWS official documentation, What is Amazon Bedrock AgentCore (accessed 2026-09-20), the platform rolled out a managed matrix spanning Runtime, Memory, Gateway, Identity, Code Interpreter, Browser, and Observability.

A clear caveat is necessary here: discussing AgentCore does not imply that Manus uses AWS services in production. There is currently no public evidence indicating any direct connection between the two. We use AgentCore as a reference specimen because it serves as a mirror: when a major cloud provider attempts to package the infrastructure needed to build a Manus into standardized products, it reveals how they categorize, abstract, and price these fragmented requirements.

Before making the comparison, let’s clarify a mental model that often causes confusion. Many developers ask: cloud providers already offer Lambda functions for executing code and container sandboxes for security analysis, so why build a dedicated AgentCore Runtime? The dividing line comes down to two entirely different system design philosophies: the jail cell versus the apartment building.

Traditional code execution sandboxes act like jail cells, typical examples being Code Interpreters and ephemeral Lambdas. The system assumes code running inside them is untrusted, originating from third parties or generated on the fly by an LLM, with potential risks of malicious attacks, infinite loops, or resource abuse. The sandbox’s primary mission is strict confinement and stripping external network access, destroying the container immediately after a few seconds of simple computation. The AgentCore Runtime hosting the core logic is a different beast: it resembles an apartment building for welcoming guests. Residing in its apartments is the developer’s own highly trusted agent process. This process requires a stable, long-term residence, capable of holding connections open for hours, accessing external networks freely, safely holding sensitive credentials authorized by users, and orchestrating external tools from its own private space.

Starting from this framing of a long-running, sticky, and trusted apartment building, cloud infrastructure begins addressing the agent checklist item by item. The following five capabilities unfold along this primary axis.

Item-by-Item Comparison: How AgentCore Meets the Checklist

With the five checklist items in hand, let’s examine how AgentCore addresses each with standardized cloud components. Across this cloud shelf, you will find both out-of-the-box replacements and architectural trade-offs that diverge sharply from vertical products.

1. Isolated Compute: From Bare VMs to Managed microVMs

The foremost requirement for an autonomous agent is an independent compute environment capable of running Python logic, reading and writing files, and initiating outbound network calls. Without such a dedicated space, multi-step tool orchestration lacks a physical anchor.

Within AgentCore, the component addressing this core requirement is AgentCore Runtime. AWS official documentation, Runtime how-it-works (microVMs) (accessed 2026-09-20), positions it as a managed compute layer. It is neither a model nor does it contain specific prompts or decision logic; rather, it provides a secure, serverless runtime environment purpose-built for dynamic AI agents and their tools.

The onboarding process for developers is straightforward: provide a standard arm64 container image stored in ECR, or upload a zip archive bundling your dependencies. Your application simply listens on port 8080 inside the container and exposes two standard HTTP endpoints: POST /invocations to handle invocations, and GET /ping for health checks.

Under the hood, each incoming session gets a dedicated lightweight microVM, achieving isolation across CPU, memory, and the root file system. While official AWS documentation avoids explicitly naming the underlying virtualization engine in product overviews, Marc Brooker, an engineer on the AWS Firecracker team, previously analyzed this dedicated-microVM session architecture in detail on his blog. Corroborated by broader technical discussions across the community, we infer with high confidence that it runs on Firecracker, the lightweight virtualization engine powering modern AWS serverless architecture.

In terms of resource limits, a single microVM session currently supports up to 2vCPU and 8GB of memory, with container image sizes capped at 2GB, sustaining continuous task runs for up to 8 hours. This spares developers the hassle of spinning up EC2 instances, mounting Docker daemons, and writing host isolation scripts by hand.

2. Session Affinity and Intra-Task State: Two Contrasting Storage Philosophies

A long task involving 50 tool calls cannot tolerate state fragmentation: if intermediate files from call one remain on Machine A while call two routes to Machine B, the agent’s contextual awareness collapses instantly.

AgentCore solves this by introducing a runtimeSessionId parameter into its invocation specification, requiring a string length of at least 33 characters, and detailing its sticky routing rules in official documentation, Runtime sessions (accessed 2026-09-20). As long as the session has not terminated and the client continues passing the same session identifier on subsequent calls, the gateway reliably routes traffic to the same dedicated microVM.

It is precisely in handling intra-task state and storage that public cloud providers and vertical product teams diverge. Looking back at the lifecycle of Manus Sandbox, Manus treats the sandbox as the user’s personal cloud computer. It supports long periods of sleep and resumption, and across reclamation and rebuilding mechanisms spanning days or even weeks, the system identifies finished deliverables and restores them automatically.

AgentCore Runtime adheres to the public cloud ethos of statelessness and ephemerality. The official documentation explicitly warns that in-process memory and the local file system within a microVM are ephemeral by default and must never be relied upon for durable storage. The default idle timeout for a microVM is just 15 minutes (configurable from 60 seconds to 8 hours). Once a task explicitly terminates or reaches its idle timeout, the system immediately tears down the microVM and clears its memory. Even if a client sends a subsequent invocation bearing the same session identifier, the system provisions an entirely blank environment.

To balance statelessness and persistence, AgentCore introduced tiered storage. In preview, it offers session storage with configurable mount paths, retaining file system data across session pauses and resumptions, which expires after 14 days of inactivity and is wiped when the agent version updates. For permanent storage of large files shared across tasks, developers must connect external S3 Files or EFS network file systems via VPC. Cloud providers treat compute environments as transient banquets, while vertical products strive to furnish personalized studies with persistent memory; the two operate under vastly different idle costs and trust boundaries.

3. Cross-Task Long-Term Memory: Decoupling Memory from Compute

A mature agent cannot act amnesic, asking the same onboarding questions every time a user starts a new session. Formatting habits, preferred data representations, and personal background established across past tasks need to persist as long-term assets.

Instead of burying cross-task memory inside the compute sandbox’s file system, AgentCore decouples it into a dedicated cloud service: AgentCore Memory. Structurally, short-term session interaction summaries and long-term user profiles are converted into structured vector records and managed by an independent storage engine.

The immediate benefit of this decoupling is freeing up expensive compute resources. The compute microVM can be torn down without hesitation the moment a task finishes, while long-term user preferences accumulate continuously in an independent, low-cost storage layer. When the next session starts in a fresh microVM, the application simply queries the Memory service via API to inject relevant memory slices into the context. Compute and memory scale independently.

4. Tool Integration and Credential Management: Secure Delegation and Unified Gateways

In products like Manus, users frequently authorize the agent to access private GitHub repositories, Google Drive folders, or internal corporate databases. Manus Sandbox therefore allows users to configure sensitive personal API tokens directly. How securely those credentials are kept dictates the security posture of the entire system.

In enterprise settings, injecting high-privilege keys directly into volatile container environments as plaintext environment variables poses significant leakage risks. AgentCore addresses this with a bidirectional inbound and outbound authentication and gateway architecture.

On the inbound side, Runtime natively integrates with AWS IAM SigV4 signature verification and supports seamless OAuth 2.0 authentication with major external identity providers, such as Amazon Cognito, Okta, and Microsoft Entra ID. This verifies the legitimacy of every incoming invocation.

On the outbound side, AgentCore provides a dedicated Identity credential management component. When an agent acts on behalf of an end user against external systems, underlying API keys and OAuth refresh tokens are held securely by the Identity service. The runtime environment receives only scoped, temporary operational credentials, removing the risk of long-lived secrets lingering in an ephemeral container’s file system. Paired with the AgentCore Gateway, the system registers and mounts fragmented external tools, enterprise microservices, and modern Model Context Protocol endpoints under a unified surface, leaving tool metadata discovery and traffic routing to the gateway.

5. Observability and Dedicated Sandboxes: End-to-End Tracing and Isolating Dirty Work

The final item on the checklist covers observability and the isolation of high-risk actions. In a 50-step tool execution chain, model hallucination or tool timeout at any single stage can stall the entire system; operators need visibility into what went wrong.

AgentCore Observability builds heavily on AWS CloudWatch and X-Ray, generating end-to-end distributed traces for each complex agent execution. Input parameters, execution duration, token counts, and exception traces for every tool call are tagged with session identifiers and archived.

For specific high-risk execution tasks, AgentCore provides turnkey managed Browser and Code Interpreter components. This reflects the jail cell versus apartment building division described earlier: the developer’s core orchestration logic resides stably in the Runtime apartment, while untrusted work—such as browsing an unknown external URL or executing a freshly generated Python math script—is offloaded to an external Browser or Code Interpreter sandbox. Even if the peripheral execution environment runs out of memory or hangs maliciously, the core agent host remains unaffected.

Manus’s five requirements map directly to AgentCore components: isolated compute to Runtime microVMs, session affinity to sticky routing, long-term memory to Memory, tool credentials to Gateway and Identity, and observability to Observability

The Infrastructure Bill Is a Confession of Agent Workload Shape

Architecture diagrams sketch an ideal skeleton, but the monthly invoice is the most candid confession of an agent’s actual workload profile. By examining billing metrics, we can reverse-engineer a cloud provider’s core engineering assumptions about this new class of workload.

Looking at AWS’s official Amazon Bedrock AgentCore Pricing page (accessed 2026-09-20), AgentCore Runtime establishes four fundamental billing rules:

  1. Per-second metering granularity, with a 1-second minimum duration for microVM compute;
  2. CPU compute is billed on active utilization, waiving CPU charges during I/O wait times spent waiting on network responses or model inference;
  3. Memory is billed based on peak per-second consumption with a minimum granularity of 128MB, spanning every second of the session lifecycle, including idle wait intervals;
  4. Scale to zero: resource consumption and billing drop to zero when no active sessions exist.

This bill is more honest than any feature list because it voices AWS’s assumptions about agent workload shapes. The two core billing rules correspond to two distinct physical realities.

Why does AWS waive CPU fees during I/O wait times? Because unlike traditional compute-heavy workloads, an agent’s wall-clock time is overwhelmingly spent waiting. As noted in the Manus blog post cited earlier, Manus observes an average input-to-output token ratio as high as 100:1. Across an entire task lifecycle, a vast majority of time is spent waiting for foundation models to run prefill inference over massive contexts, or waiting on external web pages to load and API endpoints to respond. If billed continuously for VM cores on an hourly basis, developers would pay heavily for idle waiting.

Why must memory be billed continuously by the second? Because as long as an allocated microVM remains awake, its container processes, context cache, and intermediate disk mappings physically hold a chunk of host RAM that the underlying scheduler cannot reallocate. This physical lock on memory dictates that the meter must keep running.

On its pricing page, AWS provides an illustrative benchmark for one million sessions: assume a production system processes 1,000,000 agent sessions in a month, with each session averaging 10 minutes, 90% of which is spent in idle I/O wait, and actual 1 vCPU compute time totaling just 60 seconds. Under V2 runtime on-demand rates, the final monthly bill is $6,703, which amortizes to roughly $0.006703 per session.

Dissecting that $0.0067 cost reveals a striking ratio: CPU charges account for only $0.002127, while memory charges reach $0.004576—more than double the compute cost. In agent-driven architectures, memory residency takes center stage, supplanting CPU cycles as the dominant line item on infrastructure bills.

This economic quirk, driven by unique workload shapes, spawned two major engineering headaches in the initial V1 architecture of AgentCore Runtime. According to the official AWS launch post, The new AgentCore runtime: Elastic, optimized, and consistently fast starts (accessed 2026-09-20), the first was high-watermark memory billing, where sessions were billed based on peak memory usage up to that point without reducing costs if memory was freed; the second was cold-start latency that deteriorated with image size, jumping from roughly 5.4 seconds for a 200MB image to nearly 30 seconds for a 2GB image.

AWS overhauled AgentCore Runtime V2 to address cold starts and memory economics through three core mechanisms: on-demand memory paging with idle cooldown reclamation, capturing compact snapshots of pre-warmed containers at deployment time, and instantaneously restoring execution environments directly from snapshots when new sessions arrive.

Benchmarks use a bare echo agent baseline without model or tool invocations. The test in AWS’s launch blog was initiated from an EC2 client, including cross-region public network round trips, with 5,000 cold invocations per agent. The table below draws from a separate benchmark in the official AWS sample repository, where the client ran outside AWS (measuring round-trip time over the public internet), testing 500 sessions per container image size and 5,000 sessions for zip packages, showing P75 latencies:

Container Image Size V1 Runtime Cold Start Latency (P75, AWS Benchmark) V2 Runtime Cold Start Latency (P75, AWS Benchmark)
200MB image 5.37 seconds 1.94 seconds
500MB image 7.41 seconds 2.13 seconds
750MB image 11.46 seconds 2.11 seconds
1GB image 15.48 seconds 2.13 seconds
2GB image 29.72 seconds 2.16 seconds

This snapshot mechanism imposes explicit coding constraints on developers, as detailed in AWS’s official guide, Optimize runtime V2 performance (accessed 2026-09-20). Because containers restored from snapshots share the hostname localhost and a main process PID of 1, random seeds, system timestamps, dynamic credentials, and instance network identifiers must be resolved dynamically at the request level, rather than baked in during pre-snapshot global initialization.

V2 does not guarantee cost savings across all scenarios. For lightweight applications deployed via zip bundles that offer little room for on-demand paging, AWS’s self-test showed P75 cold-start latencies of 2.85 seconds on V1 versus 1.96 seconds on V2; factoring in V2’s higher unit pricing, upgrading may not reduce costs. Finally, a clear distinction must be made: snapshot restoration currently applies only to accelerating new instance initialization. Mid-flight memory snapshot suspension and resumption for long-running tasks—flagged officially as a suspend/resume capability—remains unreleased.

AWS benchmark data shows V1 cold-start latency scaling with image size from 5.4 seconds to 29.7 seconds, while V2 remains flat at roughly 2 seconds

The Dividing Line and Honest Relevance

Once the AgentCore industrial puzzle is assembled, the dividing line between product and infrastructure emerges with clarity, delineating the boundaries of each domain.

Consider what cloud providers hand us through AgentCore: lifecycle scheduling for lightweight microVMs, a snapshot engine flattening P75 cold-start latency to roughly two seconds in AWS benchmarks, intra-session sticky routing gateways, a per-second billing pipeline that waives CPU charges during I/O waits, and enterprise-grade inbound and outbound credential proxies. This is heavy industrial foundation work, forged from years of effort by infrastructure engineers.

Yet even if developers purchase this entire foundation off the shelf, they still will not have a competitive Manus. Infrastructure handles low-level operational grind, but the core assets determining an agent’s intelligence ceiling remain rooted above the dividing line, within context engineering and model decision layers.

Above this foundation, AgentCore provides none of the critical mechanisms that truly govern agent performance. It gives you no tightly tuned core loop orchestrating 50 tool calls; no recitation mechanism updating todo.md to combat attentional drift; no KV-cache context layout engineered for prefix stability; and certainly no self-healing strategies for failed scrapes or CLI errors, prompt optimizations across messy workflows, evaluation benchmarks, or the final user product experience. The hard-won secrets that the Manus team accrued after tearing down their system four times—as recounted in their July 2025 blog post—remain entirely within this layer.

Looking across the broader industry landscape in 2026, AgentCore does not stand alone. Based on our specialized technical survey completed in July 2026, software engineering exploration into agent execution isolation has reached full-scale commercialization. From E2B providing lightweight Python code sandboxes for autonomous agents, to Microsoft’s Azure Dynamic Sessions pool; from Cloudflare Sandbox running on edge networks, to GitHub Copilot’s managed cloud development agent environments, infrastructure providers are racing to stake their claims. AgentCore’s significance lies in being backed by a premier public cloud ecosystem, packaging the entire checklist—isolated compute, persistent state, long-term memory, tool gateways, and secure identity—into an industrial-grade managed suite delivered end-to-end for the first time.

Two distinct solutions exist for the same problem. For tight-knit groups with high trust, fixed scale, and state requiring multi-month maintenance, the sensible answer is a long-term apartment rental: persistent state, monthly billing, and mutual trust among users. AgentCore, by contrast, targets public cloud multi-tenant workloads with millions of unvetted, bursty, mutually untrusted, short-lived sessions. Its DNA resembles a budget business hotel: prioritizing stateless elasticity, enforcing stringent identity checks, nickel-and-diming idle resources, and evicting expired sessions decisively. Neither approach is inherently superior; they simply cater to different audiences.

Frankly speaking, most front-line engineers and developers reading this article do not need to build a full-fledged, multi-purpose Manus by hand, nor do they need to rush into migrating workflows to a heavy public cloud monolith. Engineering priorities for autonomous agents vary by use case; blindly lifting and shifting to the cloud only adds cognitive burden.

Yet grasping these shifts carries concrete practical value. As LLM application paradigms move from one-shot QA to autonomous agents running for dozens of minutes across dozens of steps, the center of gravity in software engineering is shifting rapidly toward underlying infrastructure.

What this article aims to deliver is a portable five-item checklist and an engineering map outlining operational boundaries. When you next hear a startup announce funding for a “novel agent cloud platform,” or watch a major cloud provider host a splashy launch event for new agent building blocks, you will not be distracted by the buzzwords.

You can pull out this checklist and assess the fundamentals: Does it cover isolated compute, session affinity, long-term memory, credential management, or observability? Is its virtualization model building a jail cell to police untrusted code, or operating an apartment building to house trusted processes? What assumptions does its pricing model make about idle wait states? And above the clear boundary of infrastructure, how solidly grounded are your own product core and context design? These questions will provide you with a dependable frame of reference.