Four agent runtimes, measured from the inside
Claude Cowork, Microsoft 365 Copilot Cowork, the GitHub Copilot harness, and Codex run tasks on very different machines. Here is what each runtime actually provides.

On this page8 sections
Two unrelated products now share the name Cowork: Anthropic's Claude Cowork and Microsoft 365 Copilot Cowork. Both run your task inside a sandbox whose full hardware allocation the vendor does not publish. That is the part worth paying attention to.
Every one of the new agent products executes work on a real machine. A Linux kernel, a shell, a filesystem, a process tree, a memory limit that will kill your task when you cross it. Without published allocations, anyone sizing an agent workload is doing it against hardware they cannot see.
So I measured four of them from the inside: Claude Cowork, Microsoft 365 Copilot Cowork, a Copilot Studio agent on the GitHub Copilot harness, and Codex. I used read-only inspection of what each sandbox reports about itself, plus a small number of tests that push until something breaks.
The short version is that the machines are nothing alike, and in three of the four, the limit that will actually stop your task is smaller than the number the sandbox reports.
Brain, hands, session#
The useful mental model comes from Anthropic's engineering write-up on managed agents, which describes the architecture as decoupling the brain (the model and its harness) from the hands (the sandboxes and tools that perform actions) and the session (a durable log of events).
Microsoft arrives at the same place with different vocabulary. Its Copilot Studio harness documentation defines a harness as a runtime between what you design and the model that reasons, and names the GitHub Copilot harness, the standard harness, and the Copilot chat harness.
Two vendors, two vocabularies, one architecture. The model decides; a separate machine acts; a log outside both holds the state.
That third element is easy to skip past, and it changes how agents behave. Anthropic explains that the session is an append-only log outside the harness. The harness can recover from failure by reading that log and resuming from the last event. The context window stops being the container for the work. It becomes an index into it.
Why the wrapper stopped being enough#
For most of the last two years, an agent was a model with a list of tools. You defined the connectors, the knowledge, the topics. The model picked one, something executed an API call, and the result came back into the conversation. That pattern works, and for a large number of agents it is still the right choice.
It runs out in four specific places.
Intermediate state has nowhere to live. An agent processing sixty invoices cannot hold them all in context. With a filesystem, it writes them down and works in batches.
The action space is a fixed menu. A wrapper performs only the verbs someone defined in advance. Inside a runtime, the agent can write the verb nobody anticipated: twenty lines of Python to reshape a spreadsheet that no connector covers.
Determinism has nowhere to go. Models are unreliable at arithmetic across five hundred rows. Code is exact. A runtime lets the mechanical part go to code while the judgment stays with the model. This is the argument I find most convincing, because it improves correctness rather than convenience.
Recovery needs a workspace. Long tasks fail halfway. An agent with a filesystem can inspect what it produced, find the broken step, and redo it. An agent with only API calls receives a 400 and starts improvising.
Add the fact that Word, Excel, PowerPoint, and PDF are produced properly only by running libraries somewhere, and the sandbox stops being a feature. It becomes the substrate.

The same five properties, before and after the sandbox became the substrate.
What the four runtimes actually provide#
Measured on 17 and 18 August 2026, from inside each sandbox.

Enforceable limits across the four runtimes. No GPU is exposed in any of them.
The full readout:
Claude Cowork#
- CPU: 2 vCPU enforced; 2 vCPUs visible
- Memory: 5.84 GiB enforced; 8.00 GiB visible
- Storage: 252 GiB filesystem; 30 GiB writable
- Swap: none
- GPU / VRAM: none
- Isolation: Firecracker microVM
- Kernel: 6.18.5-fc-v20
Microsoft 365 Copilot Cowork#
- CPU: 0.3125 vCPU quota enforced; 1 vCPU visible
- Memory: 4.10 GiB enforced; 4.53 GiB visible
- Storage: 49.2 GiB filesystem; 46.7 GiB writable
- Swap: none
- GPU / VRAM: none
- Isolation: resource-controlled container
- Kernel: 6.1.146.1-microsoft-standard
GitHub Copilot harness#
- CPU: 1 vCPU enforced; 1 vCPU visible
- Memory: 2.17 GiB enforced; 2.17 GiB visible
- Storage: 9.75 GiB filesystem; 7.44 GiB writable
- Swap: none
- GPU / VRAM: none
- Isolation: KVM VM, Cloud Hypervisor
- Kernel: 6.12.8+
Codex#
- CPU: 8 vCPU enforced; 9 vCPUs visible
- Memory: 14 GiB enforced; 15.93 GiB visible
- Storage: ~63 GiB filesystem; ~54.3 GiB writable
- Swap: none
- GPU / VRAM: none
- Isolation: Kata-style VM, Cloud Hypervisor
- Kernel: 6.18.35
Start with what is identical. No GPU is exposed in any of the four. The sandbox is where the agent acts, not where it thinks. Every token of model inference happens elsewhere. If you were hoping to run a local embedding model beside your agent to save a round trip, none of these four exposed the hardware for it in these observations.
Then the shapes. Codex is sized to build software: eight cores and 14 GiB, comparable to a developer laptop. Claude Cowork behaved like a two-core machine in this observation: I measured 1.98× scaling across two workers and two jiffies of steal time across an entire boot. That does not establish dedicated physical cores. Microsoft 365 Copilot Cowork inverts the ratio entirely, trading CPU for capacity. Its shape suggests a runtime intended to pull in many files and produce documents rather than compile code. The GitHub Copilot harness is the smallest of the four and is sized to orchestrate tool calls.
One detail the vendor does not document: Claude Cowork identifies itself as a Firecracker microVM on an AWS Nitro host. The ACPI OEM ID is FIRECK, the kernel is a custom 6.18.5-fc-v20 build, and /proc/iomem carries an AMZNC10C device. Anthropic's Cowork architecture documentation describes isolated per-session sandboxes without naming the cloud-session virtualization technology, so this is a measurement rather than a restatement.
The gap that matters#
Here is the part I did not expect, and the reason this article exists.
In three of the four runtimes, the number the sandbox reports is not the number that will stop your task.

What each sandbox reports, against what it enforces.
Claude Cowork presents an 8.00 GiB virtual machine. The cgroup ceiling on the shell is 5.84 GiB. I verified this by allocating 64 MiB at a time until the kernel intervened: the last successful allocation was 5,632 MiB, and the process was OOM-killed just past it against a hierarchical_memory_limit of 6,265,872,384 bytes. A workload sized against the 8 GiB the VM advertises dies at roughly 73% of it.
The disk is more dramatic. statvfs reports a 252 GiB filesystem. Available space is 30 GiB. The remaining 210 GiB is reserved to uid 65534 under a resv_strict mount, so root cannot reach it either. df will show plenty of capacity used and almost nothing available, which is a confusing failure mode if you have not seen it before. Deletes keep working after writes start failing.
Codex shows nine online vCPUs and 15.93 GiB, and enforces eight CPU-equivalents and 14 GiB. Microsoft 365 Copilot Cowork shows a full logical CPU and 4.53 GiB, and enforces a 0.3125 vCPU quota and 4.10 GiB.
Only the GitHub Copilot harness reports what it enforces in this sample.
The practical rule is simple and worth internalizing. Read the cgroup, not the VM. nproc, /proc/meminfo, and df describe the machine that was built. The cgroup describes the machine you are allowed to use, and it is the one that decides whether your task finishes.
Nobody gets open internet#
None of these four observations exposed unrestricted internet access from inside the sandbox. The restrictions differ by product and are enforced outside the workload.
Anthropic documents that Claude Cowork cloud sessions cannot reach private, internal, link-local, cloud-metadata, or Anthropic-internal systems, and that egress passes through a mandatory proxy with allow-listed destinations. My probe matched that behavior: pypi.org and the npm registry returned 200, and github.com was reachable, while example.com, api.ipify.org, and 1.1.1.1 all failed with a connection reset. Package registries and code hosts passed; those general-internet probes did not.
Microsoft goes further for the Copilot Studio sandbox. Its engineering team states that the sandbox has no outbound network path: code running there cannot call an API or send an email regardless of what it imports. Everything leaves through configured tools and knowledge. In this Codex observation, traffic was mediated by a parent container rather than exposed directly to the sandbox.
This is the right design and worth naming, because it inverts a common assumption. Giving an agent a shell does not give it the internet. The dangerous capability is not code execution alone; it is code execution plus egress. These architectures separate the two deliberately.
What this changes when you build#
Check the machine before you promise the capability. Any task that loads a large dataset into pandas has a hard ceiling, and on the GitHub Copilot harness in this observation, that ceiling was 2.17 GiB. Model quality is not the binding constraint there.
Sizing is now a cost variable, not only a capability one. Microsoft manages Cowork through usage-based Copilot Credits. Anthropic is more explicit, pricing Managed Agents at $0.08 per session-hour of active runtime on top of token rates. When the machine is metered, a slow sandbox is not merely annoying.
Ask which isolation boundary you are getting. A Firecracker microVM, a KVM sandbox under Cloud Hypervisor, and a resource-controlled container represent different implementation choices. Guest observations can identify clues, but they do not establish every tenancy or hardware-boundary guarantee your security reviewer will care about. Ask the vendor for the guarantee, not only the mechanism you can see.
Do not assume the sandbox persists. Anthropic states that each cloud Cowork session gets its own sandbox, created at session start and destroyed at session end, sharing no state with other sessions. Microsoft describes its Copilot Studio sandbox as a working area rather than permanent storage. Anything you want to keep has to leave through a tool.
Method and limits#
These are the resources exposed to the sandbox, not the host underneath. From inside the Copilot Studio sandbox, the Azure Instance Metadata Service refused connections, so the parent VM SKU and region could not be confirmed from there.
This is one observation per runtime, not a benchmark suite. Allocations can vary by tenant, region, plan, and load, and several of these products are in preview or beta, so any of these numbers can change without notice. A branded CPU model string does not prove dedicated physical cores: the Copilot Studio sandbox advertises an 80-core EPYC and receives one core. And "no GPU exposed" describes the sandbox, not the physical host.
The memory ceiling, disk reservation, and egress allowlist on Claude Cowork are the three findings I pushed on until something failed. The rest is read-only inspection and should be treated as such.
The short version#
The harness stopped being a wrapper around tool calls and became a machine with a kernel, a memory limit, and a firewall. The four vendors sized those machines very differently, and three of them enforce a limit smaller than the one they display.
Before you design an agent around a runtime, look at what the runtime gives you. Then look again at what it actually lets you use.
I help organizations harness the power of generative AI, with a focus on Microsoft AI and agent frameworks.
Stay in the loop
Continue with your email for subscriber access and future article updates.
Related Articles
The Road to Autopilots Is a Stack, Not a Ladder
A four-layer architecture model that separates chat, API agents, harness runtimes, and Microsoft Autopilots by who owns each responsibility.
Read moreRunning local models inside Microsoft's agent sandboxes
A small ONNX model ran locally in two Microsoft agent runtimes. The ceiling on what you can put there is packaging, not compute.
Read moreStay in the loop
Continue with your email for subscriber access and future article updates.