Open Manus, type a task into the prompt box: research evaluation benchmarks for generative video models over the past month, compile an analysis report with a comparison table, and hit Enter. On your screen, everything looks quiet, save for real-time thought logs unfolding and draft text continually updating. But in the server rooms behind the glass, an entire mechanical system is rapidly grinding into gear.
Hand a task to Manus, and what happens next feels familiar when broken down step by step. According to mechanisms disclosed in the official Manus blog post, Understanding Manus sandbox - your cloud computer (accessed 2026-09-20), once the system receives an instruction, it spins up an isolated cloud virtual machine sandbox specifically for that task. This is a cloud computer equipped with dedicated compute, an independent network stack, an isolated file system, a browser, and a general-purpose toolchain. Tasks remain isolated within their virtual machine environments, never interfering with one another as they advance independently in parallel.
Inside this newly provisioned virtual machine, the model enters a tightly wound core loop. As the Manus team reviewed in their official technical blog post, Context Engineering for AI Agents: Lessons from Building Manus (accessed 2026-09-20), the logic of this agent loop is compact: after receiving user input, the model selects an action from a predefined action space in each iteration based on the current context. That action executes inside the VM sandbox to produce an observation. Both the action and observation are then appended to the end of the context and fed into the next round of model decision-making. This cycle repeats until the system determines the task is complete.
In the real world, this loop runs heavy. According to Manus, a
typical complex task consumes around 50 tool calls on average. Over
dozens of iterations, as context length grows, large language models
tend to suffer from attentional drift, gradually forgetting the initial
objective or looping inside dead ends. To counter this degradation of
attention, Manus adopted an engineering tactic to steer attention inside
the sandbox. It maintains a todo.md file on the virtual
machine’s disk, repeatedly checking and updating this checklist during
execution to cross off completed steps. The team calls this mechanism
recitation: by continually restating the overarching objective at the
very end of the context, it forces the overall plan back into the
model’s active attention zone, preventing it from getting lost in
ultra-long contexts.
Along with those 50 tool calls comes high-frequency throughput of massive, heterogeneous data. Verbose DOM trees from web scraping, dozens of downloaded PDF pages, and error stacks from terminal commands would quickly exhaust even a 256K or larger window if dumped wholesale into context, driving up inference costs and introducing painful latency. Manus treats the file system as the ultimate context, taking advantage of physical disk properties: near-infinite capacity, natural persistence, and direct read-write access for the agent. The model learns to offload intermediate artifacts to disk files, retaining only file paths or source URLs in the context window. This compression strategy preserves recoverability, allowing the model to shrink context size significantly without dropping critical facts.
When execution hits an unexpected snag, a command line error, or a web page timeout, the system keeps the failed execution trace and error messages intact within the context. It does not erase them, nor does it rely on sampling randomness to retry blindly. In the view of the Manus team, confronting real-world friction and demonstrating self-healing capability is precisely what separates genuine agent behavior from brittle scripts.
Stripping away the product shell, the underlying engineering required to sustain such a task boils down to a clear procurement checklist:
These five requirements land squarely on some of the hardest problems in modern distributed systems. In 2024, almost no off-the-shelf cloud service packaged solutions for these challenges. By 2026, the landscape is shifting substantially.
Looking back at the path taken by the first generation of general-purpose autonomous agent products, pioneer teams paid a heavy price for this foundation. The Manus team confessed in their official technical blog that they tore down and rebuilt their agent framework four full times just to find a workable context layout and scheduling logic. They coined a self-deprecating term for repeatedly groping through architectural unknowns, tweaking prompts, and debugging engineering issues: “stochastic graduate student descent,” likening it to a graduate student trying their luck through trial and error.
Rewriting the framework was merely code logic; managing the lifecycle of the underlying sandboxes was a much harder battle. According to Manus blog posts, the team had to manage a complex cloud VM scheduling system. Sandboxes had to support on-demand provisioning, automatic sleep when users stepped away to conserve resources, and automatic wake-up when users returned. Continuously idle sandboxes were automatically reclaimed, with a retention window of 7 days for free users and 21 days for pro users.
After a sandbox was reclaimed, if the user revisited a historical task, the backend scheduling system had to spin up a new sandbox and selectively restore final deliverables, user-uploaded attachments, and generated web pages or slide decks, while leaving behind intermediate scraping scripts and temporary logs. If a virtual machine suffered an unrecoverable system failure mid-task, the backend also had to automatically provision a fresh sandbox to take over.
From Manus’s official launch in March 2025, to Meta completing its acquisition on December 29 of the same year (publicly reported at over $2 billion), to its August 2026 blog post, A Note to Our Users (accessed 2026-09-20), announcing an impending spin-off back to independent operations along with regulatory data purges, the trajectory of this first-generation pioneer left a clear trail. Back then, cloud provider shelves were empty. Developers had to hand-craft everything from low-level VM lifecycle orchestration, network isolation, and storage state machines all the way up to the user interface.
The turning point arrived in 2026. After two years of production trial and error, public cloud providers began standardizing this custom foundation and putting it on their shelves.
Among these commercial offerings, Amazon Bedrock AgentCore from AWS is one of the most complete industrial examples assembled against this checklist. According to AWS official documentation, What is Amazon Bedrock AgentCore (accessed 2026-09-20), the platform rolled out a managed matrix spanning Runtime, Memory, Gateway, Identity, Code Interpreter, Browser, and Observability.
A clear caveat is necessary here: discussing AgentCore does not imply that Manus uses AWS services in production. There is currently no public evidence indicating any direct connection between the two. We use AgentCore as a reference specimen because it serves as a mirror: when a major cloud provider attempts to package the infrastructure needed to build a Manus into standardized products, it reveals how they categorize, abstract, and price these fragmented requirements.
Before making the comparison, let’s clarify a mental model that often causes confusion. Many developers ask: cloud providers already offer Lambda functions for executing code and container sandboxes for security analysis, so why build a dedicated AgentCore Runtime? The dividing line comes down to two entirely different system design philosophies: the jail cell versus the apartment building.
Traditional code execution sandboxes act like jail cells, typical examples being Code Interpreters and ephemeral Lambdas. The system assumes code running inside them is untrusted, originating from third parties or generated on the fly by an LLM, with potential risks of malicious attacks, infinite loops, or resource abuse. The sandbox’s primary mission is strict confinement and stripping external network access, destroying the container immediately after a few seconds of simple computation. The AgentCore Runtime hosting the core logic is a different beast: it resembles an apartment building for welcoming guests. Residing in its apartments is the developer’s own highly trusted agent process. This process requires a stable, long-term residence, capable of holding connections open for hours, accessing external networks freely, safely holding sensitive credentials authorized by users, and orchestrating external tools from its own private space.
Starting from this framing of a long-running, sticky, and trusted apartment building, cloud infrastructure begins addressing the agent checklist item by item. The following five capabilities unfold along this primary axis.
With the five checklist items in hand, let’s examine how AgentCore addresses each with standardized cloud components. Across this cloud shelf, you will find both out-of-the-box replacements and architectural trade-offs that diverge sharply from vertical products.
The foremost requirement for an autonomous agent is an independent compute environment capable of running Python logic, reading and writing files, and initiating outbound network calls. Without such a dedicated space, multi-step tool orchestration lacks a physical anchor.
Within AgentCore, the component addressing this core requirement is AgentCore Runtime. AWS official documentation, Runtime how-it-works (microVMs) (accessed 2026-09-20), positions it as a managed compute layer. It is neither a model nor does it contain specific prompts or decision logic; rather, it provides a secure, serverless runtime environment purpose-built for dynamic AI agents and their tools.
The onboarding process for developers is straightforward: provide a
standard arm64 container image stored in ECR, or upload a zip archive
bundling your dependencies. Your application simply listens on port 8080
inside the container and exposes two standard HTTP endpoints:
POST /invocations to handle invocations, and
GET /ping for health checks.
Under the hood, each incoming session gets a dedicated lightweight microVM, achieving isolation across CPU, memory, and the root file system. While official AWS documentation avoids explicitly naming the underlying virtualization engine in product overviews, Marc Brooker, an engineer on the AWS Firecracker team, previously analyzed this dedicated-microVM session architecture in detail on his blog. Corroborated by broader technical discussions across the community, we infer with high confidence that it runs on Firecracker, the lightweight virtualization engine powering modern AWS serverless architecture.
In terms of resource limits, a single microVM session currently supports up to 2vCPU and 8GB of memory, with container image sizes capped at 2GB, sustaining continuous task runs for up to 8 hours. This spares developers the hassle of spinning up EC2 instances, mounting Docker daemons, and writing host isolation scripts by hand.
A long task involving 50 tool calls cannot tolerate state fragmentation: if intermediate files from call one remain on Machine A while call two routes to Machine B, the agent’s contextual awareness collapses instantly.
AgentCore solves this by introducing a runtimeSessionId
parameter into its invocation specification, requiring a string length
of at least 33 characters, and detailing its sticky routing rules in
official documentation, Runtime
sessions (accessed 2026-09-20). As long as the session has not
terminated and the client continues passing the same session identifier
on subsequent calls, the gateway reliably routes traffic to the same
dedicated microVM.
It is precisely in handling intra-task state and storage that public cloud providers and vertical product teams diverge. Looking back at the lifecycle of Manus Sandbox, Manus treats the sandbox as the user’s personal cloud computer. It supports long periods of sleep and resumption, and across reclamation and rebuilding mechanisms spanning days or even weeks, the system identifies finished deliverables and restores them automatically.
AgentCore Runtime adheres to the public cloud ethos of statelessness and ephemerality. The official documentation explicitly warns that in-process memory and the local file system within a microVM are ephemeral by default and must never be relied upon for durable storage. The default idle timeout for a microVM is just 15 minutes (configurable from 60 seconds to 8 hours). Once a task explicitly terminates or reaches its idle timeout, the system immediately tears down the microVM and clears its memory. Even if a client sends a subsequent invocation bearing the same session identifier, the system provisions an entirely blank environment.
To balance statelessness and persistence, AgentCore introduced tiered
storage. In preview, it offers session storage with
configurable mount paths, retaining file system data across session
pauses and resumptions, which expires after 14 days of inactivity and is
wiped when the agent version updates. For permanent storage of large
files shared across tasks, developers must connect external S3 Files or
EFS network file systems via VPC. Cloud providers treat compute
environments as transient banquets, while vertical products strive to
furnish personalized studies with persistent memory; the two operate
under vastly different idle costs and trust boundaries.
A mature agent cannot act amnesic, asking the same onboarding questions every time a user starts a new session. Formatting habits, preferred data representations, and personal background established across past tasks need to persist as long-term assets.
Instead of burying cross-task memory inside the compute sandbox’s file system, AgentCore decouples it into a dedicated cloud service: AgentCore Memory. Structurally, short-term session interaction summaries and long-term user profiles are converted into structured vector records and managed by an independent storage engine.
The immediate benefit of this decoupling is freeing up expensive compute resources. The compute microVM can be torn down without hesitation the moment a task finishes, while long-term user preferences accumulate continuously in an independent, low-cost storage layer. When the next session starts in a fresh microVM, the application simply queries the Memory service via API to inject relevant memory slices into the context. Compute and memory scale independently.
In products like Manus, users frequently authorize the agent to access private GitHub repositories, Google Drive folders, or internal corporate databases. Manus Sandbox therefore allows users to configure sensitive personal API tokens directly. How securely those credentials are kept dictates the security posture of the entire system.
In enterprise settings, injecting high-privilege keys directly into volatile container environments as plaintext environment variables poses significant leakage risks. AgentCore addresses this with a bidirectional inbound and outbound authentication and gateway architecture.
On the inbound side, Runtime natively integrates with AWS IAM SigV4 signature verification and supports seamless OAuth 2.0 authentication with major external identity providers, such as Amazon Cognito, Okta, and Microsoft Entra ID. This verifies the legitimacy of every incoming invocation.
On the outbound side, AgentCore provides a dedicated Identity credential management component. When an agent acts on behalf of an end user against external systems, underlying API keys and OAuth refresh tokens are held securely by the Identity service. The runtime environment receives only scoped, temporary operational credentials, removing the risk of long-lived secrets lingering in an ephemeral container’s file system. Paired with the AgentCore Gateway, the system registers and mounts fragmented external tools, enterprise microservices, and modern Model Context Protocol endpoints under a unified surface, leaving tool metadata discovery and traffic routing to the gateway.
The final item on the checklist covers observability and the isolation of high-risk actions. In a 50-step tool execution chain, model hallucination or tool timeout at any single stage can stall the entire system; operators need visibility into what went wrong.
AgentCore Observability builds heavily on AWS CloudWatch and X-Ray, generating end-to-end distributed traces for each complex agent execution. Input parameters, execution duration, token counts, and exception traces for every tool call are tagged with session identifiers and archived.
For specific high-risk execution tasks, AgentCore provides turnkey managed Browser and Code Interpreter components. This reflects the jail cell versus apartment building division described earlier: the developer’s core orchestration logic resides stably in the Runtime apartment, while untrusted work—such as browsing an unknown external URL or executing a freshly generated Python math script—is offloaded to an external Browser or Code Interpreter sandbox. Even if the peripheral execution environment runs out of memory or hangs maliciously, the core agent host remains unaffected.
Architecture diagrams sketch an ideal skeleton, but the monthly invoice is the most candid confession of an agent’s actual workload profile. By examining billing metrics, we can reverse-engineer a cloud provider’s core engineering assumptions about this new class of workload.
Looking at AWS’s official Amazon Bedrock AgentCore Pricing page (accessed 2026-09-20), AgentCore Runtime establishes four fundamental billing rules:
This bill is more honest than any feature list because it voices AWS’s assumptions about agent workload shapes. The two core billing rules correspond to two distinct physical realities.
Why does AWS waive CPU fees during I/O wait times? Because unlike traditional compute-heavy workloads, an agent’s wall-clock time is overwhelmingly spent waiting. As noted in the Manus blog post cited earlier, Manus observes an average input-to-output token ratio as high as 100:1. Across an entire task lifecycle, a vast majority of time is spent waiting for foundation models to run prefill inference over massive contexts, or waiting on external web pages to load and API endpoints to respond. If billed continuously for VM cores on an hourly basis, developers would pay heavily for idle waiting.
Why must memory be billed continuously by the second? Because as long as an allocated microVM remains awake, its container processes, context cache, and intermediate disk mappings physically hold a chunk of host RAM that the underlying scheduler cannot reallocate. This physical lock on memory dictates that the meter must keep running.
On its pricing page, AWS provides an illustrative benchmark for one million sessions: assume a production system processes 1,000,000 agent sessions in a month, with each session averaging 10 minutes, 90% of which is spent in idle I/O wait, and actual 1 vCPU compute time totaling just 60 seconds. Under V2 runtime on-demand rates, the final monthly bill is $6,703, which amortizes to roughly $0.006703 per session.
Dissecting that $0.0067 cost reveals a striking ratio: CPU charges account for only $0.002127, while memory charges reach $0.004576—more than double the compute cost. In agent-driven architectures, memory residency takes center stage, supplanting CPU cycles as the dominant line item on infrastructure bills.
This economic quirk, driven by unique workload shapes, spawned two major engineering headaches in the initial V1 architecture of AgentCore Runtime. According to the official AWS launch post, The new AgentCore runtime: Elastic, optimized, and consistently fast starts (accessed 2026-09-20), the first was high-watermark memory billing, where sessions were billed based on peak memory usage up to that point without reducing costs if memory was freed; the second was cold-start latency that deteriorated with image size, jumping from roughly 5.4 seconds for a 200MB image to nearly 30 seconds for a 2GB image.
AWS overhauled AgentCore Runtime V2 to address cold starts and memory economics through three core mechanisms: on-demand memory paging with idle cooldown reclamation, capturing compact snapshots of pre-warmed containers at deployment time, and instantaneously restoring execution environments directly from snapshots when new sessions arrive.
Benchmarks use a bare echo agent baseline without model or tool invocations. The test in AWS’s launch blog was initiated from an EC2 client, including cross-region public network round trips, with 5,000 cold invocations per agent. The table below draws from a separate benchmark in the official AWS sample repository, where the client ran outside AWS (measuring round-trip time over the public internet), testing 500 sessions per container image size and 5,000 sessions for zip packages, showing P75 latencies:
| Container Image Size | V1 Runtime Cold Start Latency (P75, AWS Benchmark) | V2 Runtime Cold Start Latency (P75, AWS Benchmark) |
|---|---|---|
| 200MB image | 5.37 seconds | 1.94 seconds |
| 500MB image | 7.41 seconds | 2.13 seconds |
| 750MB image | 11.46 seconds | 2.11 seconds |
| 1GB image | 15.48 seconds | 2.13 seconds |
| 2GB image | 29.72 seconds | 2.16 seconds |
This snapshot mechanism imposes explicit coding constraints on
developers, as detailed in AWS’s official guide, Optimize
runtime V2 performance (accessed 2026-09-20). Because containers
restored from snapshots share the hostname localhost and a
main process PID of 1, random seeds, system timestamps, dynamic
credentials, and instance network identifiers must be resolved
dynamically at the request level, rather than baked in during
pre-snapshot global initialization.
V2 does not guarantee cost savings across all scenarios. For lightweight applications deployed via zip bundles that offer little room for on-demand paging, AWS’s self-test showed P75 cold-start latencies of 2.85 seconds on V1 versus 1.96 seconds on V2; factoring in V2’s higher unit pricing, upgrading may not reduce costs. Finally, a clear distinction must be made: snapshot restoration currently applies only to accelerating new instance initialization. Mid-flight memory snapshot suspension and resumption for long-running tasks—flagged officially as a suspend/resume capability—remains unreleased.
Once the AgentCore industrial puzzle is assembled, the dividing line between product and infrastructure emerges with clarity, delineating the boundaries of each domain.
Consider what cloud providers hand us through AgentCore: lifecycle scheduling for lightweight microVMs, a snapshot engine flattening P75 cold-start latency to roughly two seconds in AWS benchmarks, intra-session sticky routing gateways, a per-second billing pipeline that waives CPU charges during I/O waits, and enterprise-grade inbound and outbound credential proxies. This is heavy industrial foundation work, forged from years of effort by infrastructure engineers.
Yet even if developers purchase this entire foundation off the shelf, they still will not have a competitive Manus. Infrastructure handles low-level operational grind, but the core assets determining an agent’s intelligence ceiling remain rooted above the dividing line, within context engineering and model decision layers.
Above this foundation, AgentCore provides none of the critical
mechanisms that truly govern agent performance. It gives you no tightly
tuned core loop orchestrating 50 tool calls; no recitation mechanism
updating todo.md to combat attentional drift; no KV-cache
context layout engineered for prefix stability; and certainly no
self-healing strategies for failed scrapes or CLI errors, prompt
optimizations across messy workflows, evaluation benchmarks, or the
final user product experience. The hard-won secrets that the Manus team
accrued after tearing down their system four times—as recounted in their
July 2025 blog post—remain entirely within this layer.
Looking across the broader industry landscape in 2026, AgentCore does not stand alone. Based on our specialized technical survey completed in July 2026, software engineering exploration into agent execution isolation has reached full-scale commercialization. From E2B providing lightweight Python code sandboxes for autonomous agents, to Microsoft’s Azure Dynamic Sessions pool; from Cloudflare Sandbox running on edge networks, to GitHub Copilot’s managed cloud development agent environments, infrastructure providers are racing to stake their claims. AgentCore’s significance lies in being backed by a premier public cloud ecosystem, packaging the entire checklist—isolated compute, persistent state, long-term memory, tool gateways, and secure identity—into an industrial-grade managed suite delivered end-to-end for the first time.
Two distinct solutions exist for the same problem. For tight-knit groups with high trust, fixed scale, and state requiring multi-month maintenance, the sensible answer is a long-term apartment rental: persistent state, monthly billing, and mutual trust among users. AgentCore, by contrast, targets public cloud multi-tenant workloads with millions of unvetted, bursty, mutually untrusted, short-lived sessions. Its DNA resembles a budget business hotel: prioritizing stateless elasticity, enforcing stringent identity checks, nickel-and-diming idle resources, and evicting expired sessions decisively. Neither approach is inherently superior; they simply cater to different audiences.
Frankly speaking, most front-line engineers and developers reading this article do not need to build a full-fledged, multi-purpose Manus by hand, nor do they need to rush into migrating workflows to a heavy public cloud monolith. Engineering priorities for autonomous agents vary by use case; blindly lifting and shifting to the cloud only adds cognitive burden.
Yet grasping these shifts carries concrete practical value. As LLM application paradigms move from one-shot QA to autonomous agents running for dozens of minutes across dozens of steps, the center of gravity in software engineering is shifting rapidly toward underlying infrastructure.
What this article aims to deliver is a portable five-item checklist and an engineering map outlining operational boundaries. When you next hear a startup announce funding for a “novel agent cloud platform,” or watch a major cloud provider host a splashy launch event for new agent building blocks, you will not be distracted by the buzzwords.
You can pull out this checklist and assess the fundamentals: Does it cover isolated compute, session affinity, long-term memory, credential management, or observability? Is its virtualization model building a jail cell to police untrusted code, or operating an apartment building to house trusted processes? What assumptions does its pricing model make about idle wait states? And above the clear boundary of infrastructure, how solidly grounded are your own product core and context design? These questions will provide you with a dependable frame of reference.