The Road to Autopilots Is a Stack, Not a Ladder
A four-layer architecture model that separates chat, API agents, harness runtimes, and Microsoft Autopilots by who owns each responsibility.

On this page7 sections
The road to Autopilots is often drawn as a clean progression:
Chat → Reasoning → Agent harnesses → Autopilots
It is a useful launch slide. It is a poor architecture diagram.
The sequence combines four different questions: How capable is the model? Who keeps the control loop moving? Where does the work execute? Whose identity and authority does the system carry?
Once those questions are separated, the road looks less like a ladder and more like a stack:
- In Chat, the human owns continuation.
- In an API-wrapped agent, the application owns the loop.
- In an agent runtime or harness, the runtime packages the execution environment.
- In an Autopilot, the organization assigns identity and accountability.
That distinction matters because these layers overlap. A chat product can run on a sophisticated harness. A background agent can be highly autonomous without being an Autopilot. An Autopilot can pause for human approval and still remain an Autopilot.
These are four ownership questions, not four levels of intelligence. The responsibilities accumulate; adding a runtime does not remove the application team's obligations.
1. Chat: the human owns continuation#
I would collapse Chat and Reasoning into one layer.
Reasoning models clearly change what a conversation can achieve. They can hold a longer line of thought, make better judgments, and solve more complex problems. But that is a capability improvement inside the model. It does not, by itself, change the operating boundary around the model.
In Chat, a person still creates the rhythm of work. We open the conversation, provide context, inspect the answer, correct it, and decide whether another turn should happen. Even when Chat can search the web, run code, or call a tool, the interaction remains centered on a human session.
That is not a lesser form of AI. It is often the right one for ambiguous work where intent changes as we learn. Strategy, exploration, and judgment benefit from having the person inside the loop on every turn.
The defining question is simple: if the person stops prompting, does the work stop moving?
If yes, you are still primarily in Chat.
2. API-wrapped agent: the application owns the loop#
The first structural change happens when software, rather than the user, decides what the next turn is.
“API-wrapped agent” is my shorthand here, not an official product category. It describes an application that gives a model instructions, tools, and enough state to pursue a bounded task. The model chooses an action, the application executes it, the result goes back to the model, and the loop continues until the task finishes, fails, or reaches an approval gate.
Anthropic's practical definition is similarly compact: agents are typically models using tools in a loop, guided by feedback from the environment. OpenAI's SDK documentation makes the ownership split even more explicit. With the Agents SDK, the SDK runs the loop, while the server still owns deployment, tool implementations, state storage, and approval decisions.
That is a much larger step than adding a better prompt. The application now needs stopping conditions, error handling, authorization, state, observability, and a way to resume after human review.
Yet the execution surface can remain narrow. The agent may only see five predefined functions. It may have no filesystem, no isolated shell, no durable workspace, and no way to recover an intermediate artifact beyond whatever the application persisted for it.
The model has agency over the next action. The application still owns the world in which that action is possible.
For example, a release-summary agent might call functions to list completed tickets, retrieve pull requests, and save a draft. It can finish useful work without a general-purpose shell. That is an illustrative design, not a claim about a tested deployment.
3. Agent runtime or harness: the system provides an execution environment#
A harness moves the boundary again.
Microsoft now defines a harness as the runtime between the builder's design and the model. It decides when to call the model, what context and components to send, how to interpret the response, and which tools to invoke.
The more capable version adds something an API loop alone does not guarantee: a place for the work to live while it is happening.
That may include a sandbox, files, memory, skills, session state, retries, recovery logic, credential mediation, and observability. Instead of merely calling a sequence of APIs, the agent can inspect its own intermediate work, create artifacts, run code, revise a failed step, and continue across a longer task.
There is a terminology wrinkle here. In the broad sense, the software running an API tool loop is already a harness. The distinction I am drawing is between an application-specific loop and a reusable execution environment, not between products with and without an SDK. OpenAI's sandbox documentation makes this explicit: an application-hosted SDK can run the harness while sandbox compute handles files and commands.
A workspace also does not guarantee persistence. Cowork's documented architecture separates a temporary sandbox from saved session files. For work that must resume later, define what happens to artifacts, checkpoints, and approvals when compute disappears.

The difference is the execution environment, not whether the system can save state. Arrows return execution results to the model; durable storage is a separate design choice.
For the release-summary example, a runtime earns its place when the agent needs to inspect a repository, run checks, compare files, and revise an artifact. Those capabilities shape what the same model can accomplish.
One result from Princeton's HAL CORE-Bench Hard leaderboard makes the point unusually visible. Claude Code with Claude Opus 4.5 recorded an automated score of 77.78%, while CORE-Agent with the same model scored 42.22%. HAL also reports 95.5% after manual validation for Claude Code. Each entry represents only one run, so it would be careless to generalize the percentages. This is an illustrative comparison, not a controlled experiment establishing that the harness alone caused the gap. It is a reason to evaluate the complete system, rather than treating the model name as its performance specification.
The operational owner matters too. Microsoft's Agent Framework documentation separates hosting from protocol. A managed platform can operate the container, scaling, session lifecycle, and platform integration. In a self-hosted setup, your application remains responsible for routes, identity, authorization, policy, storage, deployment, and scaling.
So “runtime” does not automatically mean “vendor managed.” A team can build its own harness. But once it does, it has crossed the same architectural boundary and inherited the same responsibilities.
4. Autopilot: the organization assigns identity and accountability#
The move from a harness to an Autopilot is different from every earlier transition.
It is not mainly about a stronger model, a longer loop, or a better sandbox. In Microsoft Foundry's current definition, an Autopilot is a persistent, named member of the organization with its own identity, security profile, standing role, and responsible manager.
Every Foundry agent receives an agent identity. An Autopilot also receives an Entra agent user account. That second identity gives it its own email, calendar, OneDrive, Teams presence, and position in the organization chart. It can act in Microsoft 365 as itself, rather than borrowing whichever user happens to be present.
This is the cleanest boundary in the whole stack.
Memory, planning, proactivity, and learning describe capabilities. They do not establish who is acting. Microsoft says the same about autonomy: a background service agent can run on its own all day without being an Autopilot, while an Autopilot can be constrained by triggers, audiences, approvals, permissions, and policy.
Identity changes the governance question from “which app registration called this API?” to “which persistent actor performed this work, what was it allowed to do, and who answers for it?”
That is why Microsoft's Autopilot lifecycle sounds deliberately organizational. A developer publishes a blueprint. A tenant administrator approves it. A manager hires and onboards an instance, grants access to team resources, observes its work, and eventually offboards it. The manager is also the person obliged to stop the instance when it misbehaves, even when someone else caused the fault.

Conceptual identity model, not a deployment or token-flow diagram. A capability declaration does not grant access to every resource.
The lifecycle also separates consented operations from access to actual team resources. Having a mail tool or mailbox does not automatically authorize sending. Both the allowed operation and its resource scope need an explicit decision.
This is no longer only a software deployment. It is a delegation design. In the release-summary example, the organization might give the system a standing coordination role and its own account, while still requiring a manager to approve external messages.
Why the ladder breaks#
The ladder metaphor suggests every system moves neatly from reactive to proactive, from simple to advanced, and finally to autonomous. Current products already break that sequence in several directions.
A chat experience can sit on a harness. Copilot Studio explicitly has a Copilot chat harness, so interface and runtime are different axes.
A non-Autopilot can act proactively. Standard-harness agents in Copilot Studio can react automatically to events after publication. Today those triggers use the maker's credentials. That credential detail is precisely why starting without a prompt and acting under your own identity are not the same thing.
A background service agent can be autonomous. Microsoft's Foundry taxonomy says so directly. It can act under application permissions without receiving an agent user account or becoming a participant in Microsoft 365.
And an Autopilot does not need unlimited autonomy. It can operate behind approvals and within a tightly scoped audience. What makes it an Autopilot is not freedom from humans. It is durable organizational identity.

Identity determines who acts; policy determines how far it may act. These are illustrative operating policies, not four new product tiers.
Product language is evolving too. In its September 25 announcement, Microsoft renamed Scout to Autopilot and announced an expansion of its private preview. That product name and Foundry's documented identity category should not be treated as universal industry terminology. Foundry's overview still limits Autopilot blueprints to hosted agents. These are published product statements, not proof of availability in a particular tenant.
Four questions to ask before choosing a layer#
Instead of asking, “How agentic should this be?”, I would use four more concrete questions.

Choose the smallest sufficient architecture. Adding capabilities does not erase the ownership questions underneath.
Who owns the next step? If a person should inspect and steer every turn, use Chat. If software should continue through a bounded tool loop, you need an agent.
Where does working state survive? If the task needs files or code execution, define its workspace. If it must resume later, specify how artifacts, checkpoints, and approval state are saved and restored. A tool list is not a workspace, and a workspace is not a persistence guarantee.
Whose authority is used? Acting on behalf of the current user, acting as a backend service, and acting as a named organizational participant are three different security models. Do not hide that decision behind the word “autonomy.”
Who is obliged to stop it? An owner who can deploy software is not automatically a manager who watches the work. Persistent delegation needs an explicit operational accountability path.
The smallest sufficient layer is usually the best design. Chat keeps human judgment close. An API agent adds bounded initiative without forcing a general-purpose runtime. A harness earns its complexity when work needs an inspectable execution environment and deliberate recovery. An Autopilot is justified when the work itself needs a persistent organizational actor.
The final step is institutional#
The most important transition on the road to Autopilots is not from one model generation to the next.
Chat gives AI a conversational surface. An agent gives it a control loop. A harness gives it somewhere to work. An Autopilot gives it a place in the organization.
That last move carries the highest promise and the largest obligation. The system can stay with a team, act across surfaces, and keep work moving under its own scoped authority. In return, the organization has to decide who grants that authority, who watches the result, who pays for its operation, and who stops it when the operating envelope no longer holds.
The road is real. But it is built from layers of ownership, not rungs of intelligence.
I help organizations harness the power of generative AI, with a focus on Microsoft AI and agent frameworks.
Stay in the loop
Continue with your email for subscriber access and future article updates.
Related Articles
Running local models inside Microsoft's agent sandboxes
A small ONNX model ran locally in two Microsoft agent runtimes. The ceiling on what you can put there is packaging, not compute.
Read moreFour agent runtimes, measured from the inside
Claude Cowork, Microsoft 365 Copilot Cowork, the GitHub Copilot harness, and Codex run tasks on very different machines. Here is what each runtime actually provides.
Read moreStay in the loop
Continue with your email for subscriber access and future article updates.