
OpenAI Agents SDK vs Claude Code: Pick the Seat, Not the Logo
Table of Contents
OpenAI Agents SDK vs Claude Code is a seat choice, not a model bake-off. The OpenAI Agents SDK is a Python package I embed so a product can run turns, tools, guardrails, and handoffs for someone using that product. Claude Code is Anthropic's agentic coding tool: it reads a repo, edits files, and runs commands from a terminal, an IDE, a desktop app, or a browser. I am William Spurlock, AI Systems Architect and Fractional AI CTO at Spurlock Studios LLC. I have spent 20,000+ hours on agentic systems and shipped 600+ automations, with 500+ still live. The logo on the invoice does not tell me which seat to fill.
If you still want the plain definition of an agent before the runtime split, start with what an AI agent is. I assume you already know an agent is more than a chatbot with a longer prompt. What is left to decide is which runtime owns the loop.
I will not put a customer refund inside Claude Code. I will not start a repo change in the Agents SDK unless a sandbox is part of the run. Those two refusals are the whole post.
What is the OpenAI Agents SDK? #
The OpenAI Agents SDK is the library I put inside a product when the runtime should own the loop. On March 11, 2025, OpenAI described it as an open-source upgrade over Swarm, the experimental SDK from the year before, with four pieces: agents, handoffs, guardrails, and tracing. As of September 30, 2026, the Agents SDK docs still install that package with pip install openai-agents and still list agents, handoffs, and guardrails as the small primitive set, plus built-in tracing so you can see the run.
An agent, in this SDK, is a model with instructions and tools. Handoffs let one agent pass the conversation to a specialist. Guardrails validate input and output, and the docs say input checks can run in parallel with the agent and fail the run when they do not pass. Tracing is how I look at the path later. I do not treat a green chat bubble as proof the loop behaved.
The models page I read on September 30, 2026 says the recommended OpenAI path is OpenAIResponsesModel, which calls the Responses API. There is a second path, OpenAIChatCompletionsModel, for the Chat Completions API. I keep new product work on the Responses path. I wrote up that API on its own in the Responses API and Agents SDK note. The SDK is the loop around it. The API is the model call underneath.
| Piece | What I expect in a real run |
|---|---|
| Agent | Instructions, tools, and a loop that keeps going until the task is done |
| Handoff | Control moves to a specialist inside the same run |
| Guardrail | A check that can stop the run when input or output fails |
| Trace | A record I can open after the customer is gone |
| Session | Working context kept across turns in that loop |
When the agent does not name a model, the same models page says the SDK uses gpt-5.6-luna with reasoning effort set to none and verbosity set to low. If I want the frontier setting, the page says I can set the model to gpt-5.6-sol on the agent or on the run config. I write that name into the job ticket. A forgotten model field is still a choice. It is the default, not a mystery.
Jobs that belong in this seat:
- A customer is in your product, and the agent may create or update a record there.
- You need a specialist handoff, and you need the trace of that handoff in one run.
- You need a guardrail that can fail the run before a tool writes.
- The operator is not sitting at a keyboard approving each shell command.
Sandbox agents are the exception I plan on purpose. The September 30, 2026 docs describe them as specialists inside isolated workspaces, with manifest-defined files and resumable sessions. That is how I would let the SDK touch files. It is not how I would replace a coding seat that already knows git.
What is Claude Code as a production agent runtime? #
Claude Code is the operator seat: a person, a repo, and a tool that can edit files and run commands. As of September 30, 2026, Anthropic's Claude Code overview calls it an agentic coding tool that reads your codebase, edits files, runs commands, and connects to your development tools. The same page lists the surfaces: terminal, IDE, desktop app, and browser. Most of those surfaces want a Claude subscription or an Anthropic Console account. The terminal CLI, VS Code, and JetBrains also support third-party providers, per that overview.
This is not a library I import into a checkout page. I open it where the code lives. The overview says CLAUDE.md in the project is read at the start of every session, and that Claude Code can also read an AGENTS.md if the repo already has one. Skills package a repeatable workflow the team can name. Hooks run a command before or after an action, such as a format step after an edit. The Model Context Protocol (MCP) is how Claude Code reaches a ticket system, a drive, or a tool I own.
| Surface | What I use it for |
|---|---|
| Terminal | Edit files, run commands, and stay in the repo from a shell |
| IDE | Diffs and the editor I already have open |
| Desktop app | Side-by-side sessions and scheduled tasks on the machine |
| Browser | A long job on a repo I do not have locally |
The skills docs say Claude Code skills follow the Agent Skills open standard, and Claude Code adds invocation control and subagent execution on top. I keep the authoring rules in my Claude Code skills guide. A skill with no name the team can say out loud is a prompt someone will lose.
Files I expect in a repo before I call Claude Code a runtime, not a demo:
- A
CLAUDE.mdthat states the stack, the commands that are allowed, and the paths that are off limits. - One skill for a job the team repeats, such as a review or a release note.
- One hook that blocks a tool I do not want, or that formats after an edit.
- An MCP server only after the read path is the one I meant.
Subagents are the parallel workers. The subagent docs say a custom subagent can set tools, a permission mode, hooks, and skills, and that the skills field injects the full skill text at startup. I use that when one session should not hold the whole repo in its head. I do not spawn a second agent because the first one felt slow.
There is a third door, and I keep it separate. The same overview says the Agent SDK is how you build your own agents powered by Claude Code's tools, with control over tool access and permissions. That is Anthropic's embed path. Claude Code the product is the seat. The Agent SDK is the library when the product itself must drive those tools. Mixing the names is how a founder buys a CLI and thinks they shipped a customer agent.
What breaks if you point Claude Code at customer chats? #
Claude Code assumes a developer is in the loop. A customer chat does not. The permission prompts, the local settings file, and the hook that runs on your machine are built for an operator who can see the diff. A stranger in a widget cannot approve a shell command, and should not be asked to. If I point that runtime at refunds, billing, or a public inbox, I have put a coding seat in a product doorway.
When this goes wrong, the model usually gets blamed. The trust boundary is the real problem.
| Wrong move | What actually breaks |
|---|---|
| Customer chat inside Claude Code | Tool approval sits with a developer session, not with the product |
| Hook only in user settings | The block lives on one laptop, not on the customer path |
| No guardrail on the first SDK agent | Later specialists inherit input that was never checked |
| SDK edits the repo with no sandbox | File writes happen in the product process, not in an isolated workspace |
A few boundaries I write down before anyone connects a channel:
- A hook in
~/.claude/settings.jsonfollows the operator. It does not wrap a public chat. - Project hooks in
.claude/settings.jsoncan be committed. Local settings in.claude/settings.local.jsonstay on that machine. I do not pretend a local file protects every customer. - Managed policy and plugin hooks are the org-wide copies. If the company needs the block, it does not live only in my home directory.
- Skill hooks last for the rest of the session once the skill runs. Subagent hooks last while that subagent runs. The hooks docs spell out that split as of September 30, 2026.
The Agents SDK has the mirror-image mistake. The handoff docs say handoffs stay inside a single run, input guardrails apply only to the first agent in the chain, and output guardrails apply only to the agent that produces the final output. If I picture every specialist re-checking the customer, I am wrong. The first agent is the door. The last agent is the exit. The middle hops are not a second security review unless I add tool guardrails on purpose.
I also will not staff a repo refactor with the Agents SDK and call it Claude Code. The SDK can run sandbox agents in an isolated workspace. Claude Code already has the git status, the diff, and the pull request on the surfaces listed above. Pick wrong and I either rebuild a coding seat inside a product loop, or I expose a product loop on a laptop session. Either way the week is gone.
How do I choose between the OpenAI Agents SDK and Claude Code? #
Name the person the agent serves, then name the system it may write. The seat follows those two nouns. A customer plus your product database is the OpenAI Agents SDK. An operator plus a git repo is Claude Code. A product that must drive Claude Code's own tools is Anthropic's Agent SDK, which the overview keeps distinct from the CLI.
I do not start from the model name. The September 5 frontier board on this site already picks models. This page picks the runtime. A stronger model in the wrong seat still writes to the wrong place.
| Job | Seat I pick | Why |
|---|---|---|
| Customer asks for a refund in the product | OpenAI Agents SDK | The loop, the guardrail, and the trace live in the app |
| Change the repo and open a pull request | Claude Code | The tool already edits files and runs commands on that repo |
| Product feature that must use Claude Code's tools | Anthropic Agent SDK | The overview separates that library from the CLI session |
| SDK job that must touch real files | Agents SDK sandbox agents | Docs describe an isolated workspace, not a bare process write |
| One model call, no handoff, no session | Responses API directly | The SDK docs say to own the loop yourself when the job is that short |
The order I actually use:
- Write one sentence: who talks to the agent, and what record it may change.
- If the person is a customer, stop looking at Claude Code.
- If the record is a file in a repo, stop looking at a bare Agents SDK runner.
- If the product must call Claude's coding tools, read the Agent SDK note on the overview before you install the CLI for that job.
- Write the model name down. Unset on the Agents SDK means
gpt-5.6-lunaon the models page I read September 30, 2026. Say so in the ticket. - Add the check from the next section before anyone calls it production.
Non-OpenAI models are a separate switch, not a reason to change seats. The models page says to start with the SDK's built-in provider points, and that LiteLLM is a best-effort beta adapter you install as openai-agents[litellm]. I only add that adapter when a named provider is a requirement. I do not add it because a demo used a different logo.
My opinion, held loosely and used on real jobs: the Agents SDK wins when the customer never sees a terminal. Claude Code wins when the artifact is a diff. If a vendor pitch says one runtime covers both, ask them where the guardrail runs and where the hook file lives. If they cannot point at a doc, they are selling a logo.
How do I check an agent runtime before I call it production? #
I call it production when one run leaves a trace I can open and a bad input gets stopped. Token spend stays on the bill. It is not the pass line. I want the end state in the system the agent was allowed to touch, and I want a planted failure that the guardrail or the hook actually catches.
For the Agents SDK, the minimum I will accept:
- One trace for a single run, including the handoff if a specialist took over.
- One planted input that trips the input guardrail on the first agent.
- A confirmation that the handoff did not start a second run. The handoff docs say it stays inside the one run.
- The model string written on the ticket, either the explicit
gpt-5.6-solor the defaultgpt-5.6-lunaif nobody set one. - If files were touched, the sandbox workspace name, not a shrug.
For Claude Code, the minimum I will accept:
CLAUDE.mdexists, and a fresh session is supposed to read it. That is what the overview says happens at session start.- One skill a teammate can invoke by name.
- One hook that blocks a tool or a path I listed as off limits.
- If a subagent exists, its tool list is explicit. The default of "whatever the parent had" is how a specialist gets a shell I did not mean to give it.
| Check | Agents SDK | Claude Code |
|---|---|---|
| Where the loop lives | Inside the product process | On a terminal, IDE, desktop, or browser session |
| Stop a bad input | Input guardrail on the first agent | Hook or permission mode before the tool runs |
| Prove the path | Built-in trace for that run | Diff plus the hook log for that session |
| Repeatable instruction | Agent instructions in code | CLAUDE.md plus a named skill |
| File access | Sandbox agent, if files are in scope | The repo the session was opened in |
Things I refuse to count:
- A screenshot of a happy answer with no trace id.
- A model leaderboard that never names the seat.
- "We installed it" with no planted failure.
- A customer flow running on someone's laptop session.
The 35,000+ hours saved for clients came from jobs where the seat was boring and the check was real. I have also watched a clever demo die because the hook lived in one home directory and the customer path had nothing. I keep the check short. If I cannot run it in an afternoon, the runtime is not ready for a queue.
What goes on the job card before install? #
Six lines, filled in this order: who, write, seat, stop, model, proof. I do not run pip install openai-agents or the Claude Code install script until each line has a noun in it. A blank line means I am still shopping for a logo.
| Line | Customer refund in the product | Operator changing the repo |
|---|---|---|
| Who | A customer inside the product | An operator in the repo |
| Write | The order record | A pull request |
| Seat | OpenAI Agents SDK | Claude Code |
| Stop | Input guardrail on the first agent | A hook on a forbidden path |
| Model | gpt-5.6-luna unless the ticket names gpt-5.6-sol | Claude subscription or Anthropic Console account |
| Proof | One trace from a planted bad refund | One diff plus the hook firing |
How I use the card:
- If Who says customer and Seat says Claude Code, I change the seat. I do not change the model first.
- If Write says a file in the repo and Seat says a bare Agents SDK runner, I either move the job to Claude Code or I name a sandbox workspace.
- If the product must call Claude Code's tools with no one at the keyboard, Seat becomes Anthropic's Agent SDK. The CLI install is the wrong line.
- If Model is blank on an Agents SDK job, I write
gpt-5.6-lunaand the date I read the models page, September 30, 2026. - If Proof is "the answer looked fine," the card is not done.
The refund column is a blank pattern. Fill the write line with your own order record. This page names no company and no payback figure. The repo column is the same six lines on the other seat.
Two mistakes I send back:
- The card names three seats. One job gets one seat. A second runtime is a second card.
- The proof line cites a benchmark from a launch post. I want a trace or a hook from this job, not a score from someone else's demo.
When the card is full, the install is short. The Agents SDK docs still say pip install openai-agents. The Claude Code overview still points macOS, Linux, and WSL at the native install script, with Homebrew's stable cask about a week behind. The card tells me which of those lines I am allowed to type.
I keep the filled card next to the trace or the diff. If the seat on the card and the seat in the run disagree, the run does not count.
FAQ #
Is the OpenAI Agents SDK the same thing as the Responses API? #
No. The Responses API is the model call. The Agents SDK is the loop around it. As of September 30, 2026, the models page says the recommended OpenAI path is OpenAIResponsesModel, which uses the Responses API. Use the API alone when you want to own the loop yourself for a short call. Use the SDK when you want it to manage turns, tools, guardrails, handoffs, or sessions.
Did the Agents SDK replace Swarm? #
OpenAI presented the Agents SDK as the upgrade over Swarm, not as a rename of the same experiment. The March 11, 2025 announcement calls Swarm an experimental SDK and lists agents, handoffs, guardrails, and tracing as the Agents SDK improvements. I do not start a new product on Swarm. I install openai-agents and follow the current docs.
Which model does the Agents SDK use when I forget to set one? #
The models page says an Agent with no model uses gpt-5.6-luna, with reasoning effort none and verbosity low. That is the default I read on September 30, 2026, aimed at high-volume runs. Frontier work on that page means setting the model to gpt-5.6-sol yourself. If the ticket does not name one of those, I treat the run as the default and I say so.
Do handoffs in the Agents SDK stay inside one run? #
Yes. The handoff docs say handoffs stay within a single run. Input guardrails apply only to the first agent in the chain. Output guardrails apply only to the agent that produces the final output. I put the strict input check on the door, not on a specialist I hope will be careful. Tool guardrails are a separate setting when I need a check on each function call.
Can the Agents SDK call a non-OpenAI model? #
Yes, through the SDK's provider points, and LiteLLM only when those points are not enough. The models page says to start with built-in integration for a non-OpenAI provider, and that LiteLLM is a best-effort beta adapter installed with openai-agents[litellm]. I validate tool calling on the provider I will actually ship. A beta adapter is not a promise that every feature matches the OpenAI path.
Where do Claude Code hooks live? #
The hooks docs list seven places: user settings, project settings, local project settings, managed policy, plugin hooks, skill frontmatter, and subagent frontmatter. User settings apply to all of your projects on that machine. Project settings can be committed. Local project settings stay off git when Claude Code writes them there. Skill hooks stick around for the rest of the session. Subagent hooks end when that subagent ends. I pick the scope that matches who must be protected.
What is Anthropic's Agent SDK next to Claude Code? #
Claude Code is the product you sit in. The Agent SDK is how you build your own agents on Claude Code's tools. The overview says that SDK gives you control over the run, tool access, and permissions. I use the CLI when an operator is in the repo. I look at the Agent SDK when my product has to drive those tools without a person typing in the terminal. I do not install the CLI and tell a customer that was the embed.
Should I install Claude Code with Homebrew or the native script? #
The overview calls the native install script the recommended path on macOS, Linux, and WSL, and says native installs update in the background. The Homebrew cask claude-code tracks the stable channel, which that page says is typically about a week behind and skips releases with major regressions. Homebrew does not auto-update. I use native when I want the background update. I use the stable cask when I would rather lag a week on purpose.
Book an AI automation strategy call #
Bring the job, not the logo. I am William Spurlock. I build this class of work at Spurlock Studios LLC. On an AI automation strategy call, tell me who the agent serves, what it is allowed to write, and whether that write is a product record or a git diff. I will tell you whether the seat is the OpenAI Agents SDK, Claude Code, or Anthropic's Agent SDK, and which single check has to pass before it touches a queue.
I will not invent a pass rate to make the page feel finished. Across 600+ automations built and 500+ still live, the useful argument is the seat plus the planted failure. The 35,000+ hours saved for clients showed up after that argument was boring. If you only have a model name, say that. I will pick the seat with you before anyone picks a slogan.
Related Posts

Evaluating AI Agents by Outcome: The Eval Harness That Measures Business Impact, Not Token Math
Evaluating AI agents by outcome means a second person can check the end state in the system of record. Token totals stay on the bill, off the pass line.

Your First MCP Server Without a Developer: What It Takes and What It Does
An MCP server lets Claude Desktop, ChatGPT, or n8n call your business tools. How a non-developer connects — and why the first server should stay read-only.

AI Agents for Operations: Replacing the Repetitive Tasks That Drain Your Team
Yes — an AI agent can analyze your ops data and recommend next moves, but only if you give it the right access, read-only first, and a human approval layer.


