The short version
- Researchers are meeting at SOSP to examine operating system primitives designed specifically for autonomous AI agents
- The workshop focuses on isolation scheduling memory observability reliability and security for long-running agent workloads
- Accepted work includes research on capability control risk budgets agent context and OS abstractions for self-evolving agents
A research community focused on operating systems for AI agents is meeting in Prague around a question that becomes harder to ignore as agents move from demonstrations into long-running services. Traditional operating systems were designed around processes threads files sockets and resource controllers. Agentic workloads introduce another layer because the software generating actions is itself adaptive and can change how it uses those resources over time.
The second AgenticOS workshop at SOSP brings together systems and AI researchers working on this problem. The organizers describe a gap between existing operating system primitives and the requirements of agents that plan invoke external tools collaborate maintain long-lived state and interact continuously with their environment.
An agent needs guarantees about effects not just execution
One of the central arguments at the workshop is that an operating system traditionally does not judge whether a program’s actions make sense. It isolates the program and enforces resource rules while the application remains responsible for its own logic. An AI agent changes that relationship because it generates its own actions and can make decisions across a long sequence of operations.
The workshop’s keynote frames this as a missing runtime contract for agents. An order can time out after a remote system has already accepted it and a retry can create a duplicate. A context compaction step can lose a standing instruction. A code review approval can become detached from the code that was actually merged. These failures can occur even when the model’s individual decision looks reasonable.
That is why the research agenda includes more than sandboxing. Accepted work covers semantic profiling scheduling reliability budgets capability controlled runtimes formal representations of agents and mechanisms for controlling what an untrusted agent can externalize from its exploration environment.
The research agenda reaches below the agent framework
The workshop program includes work on GPU profiling under coding agent workloads and scheduling methods that account for reliability budgets. Other papers examine volatile agent context and OS runtimes for embodied agents. The goal is to establish mechanisms that can be enforced below the model and agent framework so applications do not have to rebuild the same safety and reliability logic independently.
This is a different direction from building another orchestration library. An orchestration layer decides which model or tool should run. An agent operating system would need to define the guarantees around execution itself including isolation provenance state management resource limits and the consequences of irreversible actions.
The emerging question is what an operating system must guarantee when its primary user is an AI agent rather than a human
AgenticOS workshop
If this research matures it could change the infrastructure stack used by autonomous software. Instead of treating an agent as another application process developers could have system primitives that understand agent identity delegated authority context lifetime and risk. That would make safety and reliability properties part of the execution environment rather than optional features added by each application.
The work remains research rather than a finished operating system standard. The importance of the workshop is that multiple groups are now examining similar problems from different directions. Isolation scheduling observability capability control and irreversible actions are becoming systems questions as much as AI questions.
The workshop also highlights the need for better semantics around state. Traditional processes have memory and files but the operating system does not normally know whether a piece of state represents a user instruction a temporary plan or an irreversible commitment. Agents routinely mix those categories and can carry context across tools and sessions.
Resource scheduling presents another challenge. An agent can spend compute on repeated experiments or wait on external services while holding resources that other agents need. A scheduler designed for agent workloads may therefore need to understand deadlines reliability budgets and the cost of abandoning or retrying a task.