Written by
Sijie GuoCEO and Co-Founder, StreamNative, Apache Pulsar PMC Member
Rui FuStaff Software Engineer, StreamNative
Pengcheng JiangStaff Software Engineer, StreamNative
Guangning E
Open-Sourcing Orca Agent Engine: Building an Open Runtime for Managed Agents

Today, we are excited to open-source Orca Agent Engine (OAE), a declarative runtime for building, deploying, and governing AI agents in your own environment. It is available under Apache 2.0 in Developer Preview, together with the CLI, SDKs for TypeScript, Python, and Go, and ready-made skills and cookbooks.
We introduced Orca Agent Engine a year ago. Since then, we have been running agents on it, gathering feedback, and improving it. The goal is straightforward: you choose the harness and model, set the rules for how agents run, and keep the control plane and session records in your environment. Here are its most important pillars:
-
An open runtime for managed agents. Declare and run agents in your own environment, where the control plane and event log stay under your control.
-
Choice of harness. Run the agent loop with the Claude Agent SDK, Codex, or Pi.
-
Choice of model. Use Anthropic, OpenAI, Google Gemini, DeepSeek, or MiniMax, with support for OpenAI-compatible endpoints. Choose the model in configuration.
-
Governance that lives outside the agent. Apply guardrails that only tighten, keep provider credentials outside the sandbox, and record each session in an event log.
-
Compatible with the Claude Managed Agents API. Use the official Anthropic SDKs for supported operations by changing their base URL; the conformance matrix documents differences.
-
Integrated with the tools you already run. Sessions export as OpenTelemetry traces to your observability and evaluation stack.
Agents are a new kind of workload
Building with language models has evolved from individual model calls to agent frameworks and harnesses. A harness is the program that runs the agent loop, managing model calls, tools, and context. Managed agent APIs put that harness behind an API call, with long-running sessions, hosted tool execution, and state that outlives a request.
An agent is a new kind of workload. It can run untrusted code against production systems, so isolation cannot be an afterthought. It may work for hours and need to survive a restart, so a session cannot be treated like a single request. It needs an identity, an isolation boundary, a budget, and a durable record of what it did. In Orca, that session event log is also the basis for resuming interrupted work.
What production agents actually need
Three problems shape how we design Orca for teams moving agents from demos toward production.
Cost is an attribution problem before it is a price problem. Agents spend in bursts, across several providers, sometimes on credentials issued to individuals, so an aggregated bill can obscure costs per team or per feature. Teams need to connect spending to individual sessions and set limits before costs accumulate.
Control requires more than permission to call a tool. An agent asked to clear a backlog might delete everything in it and report success. A tool allowlist alone cannot determine whether that is an acceptable outcome or enforce a session budget. Limits need to account for the action, its arguments, and the session’s state, and they need to be enforced outside the agent’s control.
Choice helps teams manage cost and control. Not every task needs a frontier model, yet changing models or harnesses can require code changes and new integrations. Reducing that work makes it practical to match the model and harness to the task while keeping runtime policies consistent.
Why we built a new runtime
Managed agent APIs simplify operations, while open runtimes give teams more deployment options. We built Orca around the combination we wanted: a declarative API, choice of harness and model, governance outside the agent, and a control plane and event log that teams operate in their own environment.
The runtime should make it easier to adopt a new model or harness while keeping the policies, records, and tools around it. That is the layer Orca takes on.
Choose your harness and your model
- SDK, CLI and UI clients reach the Registry.
- The Registry works with the Harness Server. It never runs agent code itself.
- The Harness Server executes built-in tools in a Sandbox.
- The Harness Server and the Sandbox both reach the AI Gateway, which holds the credentials.
- The AI Gateway calls MCP tool servers and LLM providers on their behalf.
- Transcripts and audit logs land in the Event Store, backed by Kafka, Postgres or Pulsar.
Choice starts in the declaration. In Orca Agent Engine, the harness and model are fields in the agent configuration. Changing the harness mode currently requires creating a new agent; it is not an in-place switch on an existing agent. Depending on the harness, it runs beside the sandbox or inside it. Model and tool calls routed through the AI Gateway use credentials resolved at the gateway. You can configure a supported provider or an OpenAI-compatible endpoint for a model you host yourself.
That changes what you can do. A routine task can run on a smaller, lower-cost model while more demanding work uses a frontier model. You can choose another supported harness through the same runtime API, subject to that harness’s capabilities. Agent configurations are versioned, and each session keeps the version it started with, so later configuration changes do not rewrite past runs.
Govern agents from outside the agent
Control means limits the agent cannot edit, so Orca Agent Engine enforces them outside the agent’s editable configuration. Before an action runs, the applicable guardrails evaluate it against rules set for the organization, workspace, agent, and session. Each rule returns allow, ask, or deny, and the strictest answer wins: a lower scope can tighten a rule but cannot loosen one set above it.
Cost gets the same treatment. Budgets are rules too: at a soft threshold, the agent asks before it spends more; at the hard cap, it is stopped. Switching to a cheaper model requires manual configuration; Orca does not automatically downgrade the model. Provider credentials stay in a vault and are resolved at the gateway, rather than passed into the sandbox. Each session’s record captures its events and approval requests, giving teams a record they can inspect.
Fit into the stack you already run
- Build, test and monitor happen in the tools you already use.
- Run and govern are provided by Orca.
- The Orca Agent Engine sits at the centre as the runtime.
- Orca exports to your evaluation and observability platforms; it does not ask you to migrate to it.
Choice also covers the tools around the runtime. The agent lifecycle already has frameworks and harnesses for building, evaluation platforms for testing, and observability platforms for monitoring. Orca focuses on running and governing agents, and connects to those tools through trace export.
Orca implements the Claude Managed Agents API, so the official Anthropic SDKs can use supported operations by changing their base URL. A generated conformance matrix documents coverage and differences. Sessions export as OpenTelemetry traces to an OpenTelemetry backend or a Langfuse-compatible endpoint. You decide whether to include prompts and tool output. The export is one way: Orca sends traces without importing your evaluation datasets or scores.
Why we opened it
Models and harnesses will keep changing. We believe teams should be able to inspect and extend the runtime that connects them, and retain control of the records their agents produce.
We would like to see others build on Orca, including services that compete with ours. More harnesses, new guardrails, and storage and sandbox backends are all places to contribute. Support for framework-based agents is next, and we would like to build it with the community.
Get started today
Orca Agent Engine is available in Developer Preview. Start with the quickstart: bring up the local stack, configure an agent, and run your first session with an Orca SDK or a supported Anthropic SDK workflow. Try it with your own workload. Open an issue, share what you build, and help us shape the runtime.
For more on the architecture and contribution opportunities, read our companion post, Inside Orca Agent Engine: governing managed agents from outside the agent.
-
Orca Agent Engine: github.com/orca-ae/orca-agent-engine
-
Documentation and quickstart: docs.runorca.ai
-
Claude Managed Agents conformance matrix: conformance-matrix.md
-
The CLI: github.com/orca-ae/orca-cli
-
SDKs: TypeScript, Python and Go
-
Skills and cookbooks: github.com/orca-ae/orca-skills and github.com/orca-ae/orca-cookbooks
About the authors

Sijie Guo
CEO and Co-Founder, StreamNative, Apache Pulsar PMC Member
Sijie’s journey with Apache Pulsar began at Yahoo! where he was part of the team working to develop a global messaging platform for the company. He then went to Twitter, where he led the messaging infrastructure group and co-created DistributedLog and Twitter EventBus. In 2017, he co-founded Streamlio, which was acquired by Splunk, and in 2019 he founded StreamNative. He is one of the original creators of Apache Pulsar and Apache BookKeeper, and remains VP of Apache BookKeeper and PMC Member of Apache Pulsar. Sijie lives in the San Francisco Bay Area of California.

Rui Fu
Staff Software Engineer, StreamNative
Rui Fu is a software engineer at StreamNative. Before joining StreamNative, he was a platform engineer at the Energy Internet Research Institute of Tsinghua University. He was leading and focused on stream data processing and IoT platform development at Energy Internet Research Institute. Rui received his postgraduate degree from HKUST and an undergraduate degree from The University of Sheffield.

Pengcheng Jiang
Staff Software Engineer, StreamNative
Pengcheng Jiang is a software engineer at StreamNative. He mainly focuses on the Compute platform, including Pulsar Functions, IO Connectors, and Kafka Connects. Before joining StreamNative, he worked at Naver China and was in charge of the Serverless Platform. Pengcheng got his Master's degree from the China Academy of Telecommunications Technology (CATT) and a Bachelor's degree from Beihang University(BUAA).

Guangning E
Guangning E is an Apache Pulsar committer and the main contributor to Apache Pulsar IO and Apache Pulsar Manager. He works as a senior software engineer at StreamNative, where he specializes in cloud platform, cloud computing, and big data related fields.
