Running local models inside Microsoft's agent sandboxes
A small ONNX model ran locally in two Microsoft agent runtimes. The ceiling on what you can put there is packaging, not compute.

On this page8 sections
I loaded a quantised BERT-tiny model into a Copilot Studio agent on the GitHub Copilot harness and into Microsoft 365 Copilot Cowork. Both ran it on CPU inside their own sandbox. In these two runs, neither used an external endpoint for inference.
The task I gave it was prompt-injection detection, and both runtimes returned the same decision. But the classification result is almost beside the point. What the runs demonstrate is that these sandboxes will load model weights from disk and execute inference locally, which is a capability nobody advertises and almost no agent design currently uses.
That is worth sitting with. The new Microsoft agent products each execute a task inside a Linux sandbox with a shell, a filesystem, and a Python runtime. Most people treat that runtime as the place where the platform's tools happen to run. It is also a place where your own code and your own models can live, for as long as the task lasts.
This article is the experiment, the ceiling I ran into, and what I think changes when that ceiling moves.
Three runtimes, three results#
I started where it was easiest. Codex gives a task 8 enforced vCPU and 14 GiB of RAM, which is enough to be careless with. I loaded LFM2.5-230M-UD-IQ2_M.gguf, a 105 MB quantisation of Liquid AI's LFM2.5-230M, and asked it for a joke.
It worked on the first attempt. Prompt processing ran at 155.9 tokens per second, generation at 138.6, and the whole thing finished in 0.98 seconds. On CPU, in a cloud sandbox, from a general-purpose language model that had no business being there.

Then I moved to the two runtimes I actually build on, and the picture changed. The GGUF model did not run in either. An INT8 ONNX model did, in both.
So the payload became a BERT-tiny classifier, packaged with its tokenizer and calibrated thresholds. The Copilot Studio agent loaded it from its sandbox filesystem, ran inference through ONNX Runtime, and returned BLOCK at 0.9999665 for an obvious injection string. Cowork ran the same model as a local skill and returned BLOCK at an injection score of 1.0. I did not verify whether Cowork rounds its output, so I would not read anything into the difference between those two numbers.
Same weights, two unrelated products, both executing them locally.

Why this works at all#
The sandbox is a real machine, and I measured what each one provides a few days ago. The GitHub Copilot harness gave me 1 vCPU and 2.17 GiB of RAM. Cowork enforced a 0.3125 vCPU quota against 4.10 GiB. Codex provided 8 vCPU and 14 GiB. None of the three exposes a GPU.
That last point sets the shape of everything else. There is no accelerator in any of these runtimes, so the sandbox is where the agent acts, not where it thinks. Any model you run there is a CPU model, and that pushes you toward small, quantised, task-specific things rather than general reasoning.
The second constraint is stricter and more interesting. Microsoft describes the Copilot Studio sandbox as having no outbound network path, and is blunt about what that means: code there cannot call an API "no matter what it imports." You cannot fetch weights at runtime. You cannot pip install an inference engine. Whatever you intend to run has to be inside the package before the task starts.
That turns a question about compute into a question about logistics, and logistics is where I hit the wall.
Packaging is the ceiling, not compute#
My assumption going in was that small models were a CPU story. A third of a core is not much, and a 230M-parameter model on that budget is not appealing.
That is not what stops you.
Microsoft documents Cowork's skill packaging precisely. A plain .md skill file can be up to 1 MB. A .zip or .skill archive, with a SKILL.md at its root plus any companion files the skill uses, can be up to 10 MB compressed, up to 50 MB uncompressed, and up to 100 files.
That is the whole door. Here is what has to fit through it, smallest first.
- A BERT-tiny INT8 classifier, the thing I actually shipped, lands in single-digit megabytes. It fits.
- The ONNX Runtime wheel is 20.8 to 23.1 MB. It does not fit.
all-MiniLM-L6-v2embeddings at INT8 are 23 MB. They do not fit.- The DeBERTa-v3-base model behind most off-the-shelf detectors is 184M parameters, tens of megabytes even at INT8. It does not fit.
LFM2.5-230M-UD-IQ2_M.gguf, the file that ran happily in Codex, is 105 MB. It misses by an order of magnitude.

The compressed limit binds harder than it looks, because every artifact on that list is already-quantised binary weight data and compresses poorly. Zipping a GGUF does not buy you a factor of ten.
So the reason a 230M-parameter model does not run in Cowork is not that a third of a core is too slow to serve it. It is that 105 MB cannot be delivered through a 10 MB door. The runtime has 4.10 GiB of RAM sitting there, perfectly capable of holding the model. There is no way to hand it over.
That reframing is the practical takeaway. Before asking whether the CPU can run your model, ask how the weights, the tokenizer, the thresholds, and any runtime dependency reach the sandbox. Delivery becomes the binding constraint before inference ever starts.
A useful thing falls out of that arithmetic#
The list above contains a second result, and I did not go looking for it.
ONNX Runtime for Linux ships as a prebuilt manylinux wheel of 20.8 to 23.1 MB. No compilation, which is exactly why it travels well into a sandbox with no compiler. But at 21 MB it cannot fit inside a 10 MB compressed archive, and it cannot be installed at runtime either, because there is no network to reach PyPI.
The classifier nevertheless ran.
If it arrived as a skill archive, ONNX Runtime was already in that sandbox before my skill got there. That is a deduction rather than a documented fact, and it depends on the delivery assumption. It is consistent with what Microsoft's own engineering team says about the Copilot Studio sandbox, which ships with libraries for documents, spreadsheets, PDFs, data, charts and images and was counted at 99 Python libraries in July 2026, with no public inventory.
If it holds, it is good news, and it is the part I would most like other people to verify in their own tenants. You are not shipping an inference engine. You are shipping weights, and the entire budget goes to the model.
What the tier is good for today#
Under 10 MB compressed, on CPU, the honest answer is narrow classification. That is less limiting than it sounds, because a great deal of what agents do repeatedly is classification wearing a costume.
Injection and jailbreak detection is the case I tested, and a public size ladder already exists for it. TestSavantAI publishes its detector family in tiny, small, medium, base and large variants, built respectively on BERT-tiny, BERT-small, BERT-medium, DistilBERT and DeBERTa, each with an ONNX export. Only the bottom of that ladder fits a Cowork skill archive today, which makes model selection a packaging decision as much as an accuracy one.
The same size class covers intent and topic routing, language identification, toxicity scoring, sentiment, and small entity models that can catch an identifier before it lands in a summary. Any of those, run locally, saves a round trip and keeps the content inside the runtime.
Two sentences of honesty before the interesting part. Microsoft already ships built-in protection against user and cross-domain prompt injection, so a local classifier is an addition rather than a gap-filler. And a model inside the sandbox has no authority over the layer that created the sandbox, so it cannot veto a tool call the orchestrator has decided to make; that requires an orchestrator-level control.

What becomes possible#
Everything from here is my forecast rather than a measurement. I think this tier gets substantially more interesting, and the reason is that the gap is small.
The distance between what fits and what is genuinely useful is about one order of magnitude. A 10 MB door against a 105 MB general-purpose model. That is close enough that either side could close it within a product cycle, and both sides are moving.
Models are getting smaller faster than most architecture assumptions account for. Liquid AI reports LFM2.5-230M decoding at 213 tokens per second on a Galaxy S25 Ultra and 42 on a Raspberry Pi 5. When a vendor targets a phone and a hobbyist board as first-class deployment surfaces, a cloud sandbox with a full core and gigabytes of RAM is a comfortable target by comparison. The quantisation ladder for that same model already spans 462 MB down to 105 MB, and the aggressive end of that ladder keeps improving.
Limits, meanwhile, are lines in a validator rather than capacity decisions. A CPU quota reflects what a vendor is willing to pay for. An archive size limit reflects what someone thought a skill would need. The first is expensive to change and the second is not.
If either moves, three things follow.
Embeddings arrive first. MiniLM-class models at 23 MB INT8 are the nearest miss on the whole list, roughly a factor of two away. Local embeddings would mean semantic search and reranking over the working set without a retrieval service, inside a runtime that already has the documents open.
Generation becomes possible for narrow work. Not reasoning. Extraction, normalisation, rewriting a field, turning a messy cell into structured output. The jobs where a round trip to a frontier model costs more in latency than the task is worth. My Codex run is the existence proof that a 230M model is fast enough on CPU to sit inside a loop.
The runtimes accumulate engines. If ONNX Runtime really is already preinstalled, that is a strong hint about where this goes. Sandbox base images will grow inference runtimes the same way they grew PDF and spreadsheet libraries, because the platform benefits when your payload is only weights.
There is a cost dimension too, though I would hold it loosely. Microsoft prices a Cowork task partly on runtime, which means local inference is not free the way local inference on your laptop is free. It also means a decision made in milliseconds inside the sandbox competes against one that costs a network round trip plus a model call. Which of those wins depends on what Microsoft's "runtime" input actually meters, and Microsoft has not published that.
What this is not#
One observation per runtime, one tenant, one day, on preview-era products that can change without notice. Treat all of it as a reproducible starting point rather than a specification.
The size budget is verified for Cowork specifically. I have no primary source for an equivalent binary-delivery budget in the Copilot Studio harness, and the knowledge-source limits people usually quote are a different pathway, meant for retrieval over documents rather than for placing a binary in a sandbox. Do not assume the two products share a limit.
I also did not isolate why GGUF failed. Missing engine, missing compiler, and delivery size are all plausible and I tested none of them against each other. And on the classifier itself, one obvious injection string is a smoke test, not a security evaluation. Small guardrail models of this class have documented failure modes in both directions, both evasion and over-defence, in literature I have cited but not read in full. If you deploy one, calibrate it on your own traffic and measure false positives as seriously as detections.
Go and try it#
The recipe is short enough to be worth stating. Take a small ONNX model. Quantise it to INT8. Package it with its tokenizer and any thresholds, inside whatever delivery mechanism your runtime offers. Keep the whole thing under 10 MB. Call it from your skill and see what the sandbox does.
I expect most people who try this will find something more interesting than a classifier, because the constraint selects for creativity. A 10 MB budget and one CPU core is roughly the machine a lot of genuinely useful software was written for.
The result worth carrying away is not that an agent sandbox can run any model. It cannot, and the reason is a packaging limit rather than a hardware one. The result is that there is a real small-model tier inside these runtimes right now, most designs are not using it, and it is one order of magnitude away from being somewhere you would want to put real work.
I help organizations harness the power of generative AI, with a focus on Microsoft AI and agent frameworks.
Stay in the loop
Continue with your email for subscriber access and future article updates.
Related Articles
The Road to Autopilots Is a Stack, Not a Ladder
A four-layer architecture model that separates chat, API agents, harness runtimes, and Microsoft Autopilots by who owns each responsibility.
Read moreFour agent runtimes, measured from the inside
Claude Cowork, Microsoft 365 Copilot Cowork, the GitHub Copilot harness, and Codex run tasks on very different machines. Here is what each runtime actually provides.
Read moreStay in the loop
Continue with your email for subscriber access and future article updates.