Zilmac Blog
← Back to Tech Practice

After OpenAI DevDay 2026: What Server Do AI Agents Need? Local, Cloud, and Cost

AI Development ·~12 min read

Laptop and planning sheets, standing in for choosing local, cloud, or a specialty server after DevDay
Three bills, not one box
Loop, runtime, and specialty hardware — do not buy “an AI server”
Hosted sandbox is Linux
Python, Node, and a shell are enough; Xcode and device signing are not
Tokens dominate
Containers bill by the minute; swap the default model and the month moves first

The short version: after OpenAI DevDay 2026, the first thing an AI-agent team needs is not a more expensive GPU box. It needs a clean split — where the loop runs, where code executes, and which steps must land on a Mac or a private network. As of 27 September 2026 the event is still 29 September; the keynote has not started. This article does not treat unannounced products as fact. The stack that actually changes procurement already shipped on 10 September: the Agents API public beta turns the Codex harness into a hosted loop and lets you pick an OpenAI Linux sandbox, a partner sandbox, or your own. Date and livestream: DevDay time and what to watch.

Sources: the Agents API overview, hosted sandboxes, self-hosted sandboxes, and the pricing page (containers, tools, gpt-6-astra / gpt-6-sol). Layers: the 2026 agent stack. Mac vs GPU: dual-track deploy.

Evidence levels
  • OpenAI has confirmed: Agents API is in public beta with no extra API fee. You pay tokens, tools, and hosted containers. OpenAI runs the loop; you choose the sandbox.
  • OpenAI has confirmed: the hosted sandbox is a Linux workspace with Python, Node, and a CLI. Self-host with codex exec-server over outbound WebSocket. Data residency is documented as the United States only; self-hosting does not make the API ZDR-eligible.
  • Do not write as fact: new model IDs, new prices, or Agents API GA from the 29 September keynote. After the show, re-read the docs before changing the default model and the budget. You can pick an environment type now.

What is already confirmed

The 10 September Agents API already split “the server”:

  • Harness / loop: model calls, tools, context compaction, long sessions — OpenAI runs this. Your app creates a session and consumes events.
  • Runtime: openai_hosted, self_hosted, or a partner (Blaxel, Cloudflare, Daytona, E2B, Modal, Vercel, and others). Q&A with no shell can use environment.type = none.
  • Keys: the app key (OPENAI_API_KEY, agents read/write plus responses write) stays outside the sandbox. The executor gets a separate, narrowed environment key as CODEX_API_KEY.

The hosted working directory is /workspace. You can install packages, drop files, and pull artifacts; the network can be off or domain-restricted. It is not macOS and not your VPC. Own image, private net, or GPU → self-host or a partner. Xcode, notarization, devices → the Linux sandbox stops and the step moves to Apple silicon. Checklist: M6 Mac mini for agents.

Pricing docs: models at that model’s API rates; built-in tools at tool rates; hosted sandboxes at standard container rates. “No extra API fee” covers the loop only — not tokens, not the container.

Which layer you are actually buying

“What server does an agent need?” is usually one PO. After DevDay, make it three:

Layer Who runs it What you buy
Loop (harness) OpenAI (Agents API) or your own SDK loop An API key and quota — not a new chassis
Runtime (sandbox) Hosted Linux, a partner, or codex exec-server on your box Per-minute containers, or a Linux/POSIX host that can reach api.openai.com outbound
Specialty steps macOS build, signing, devices; or your GPU / VPC A fixed Apple silicon node or a box in the right residency zone — not on the sandbox acceptance sheet

A laptop is a fine first executor and a bad sole production host: keys, disk, and network share a browser session. If the keynote says “swap the model without rewriting orchestration,” that sentence hits the loop, not Xcode inside the hosted sandbox.

How local works — and when to stop

The Agents SDK still ships Unix-local and Docker clients for a same-day minimum path. Self-host docs want a POSIX workspace, @openai/codex@alpha, and codex exec-server --remote … --environment-id … with the narrowed key. Traffic is outbound: register on https://api.openai.com, take commands on wss://codex-cloud-environments.chatgpt.com. A corp firewall that allows browser HTTPS and blocks WebSockets looks “connected” while the executor never reaches connected.

Local is enough when:

  • One person is editing prompts, inspecting artifacts, and checking tool order.
  • Source and secrets must not enter a hosted container, so you prove network policy in local Docker first.
  • The job stays in one repo and does not need four concurrent sub-agents on four machines.

Stop using only local when:

  • Sessions must survive overnight, reconnect, or be reproduced by a teammate on the same environment.id.
  • The lid closes or the VPN blips; the executor dies while OpenAI still meters the session.
  • iOS / macOS signing, notarization, or a device — Docker Linux will not pass.
  • Compliance says source never leaves the laptop, but Agents API session state still resides in the US. Local execution does not turn the API into ZDR. OpenAI says a self-hosted sandbox does not make the API ZDR-eligible either.
Developers around laptops, standing in for splitting local Docker from a cloud agent sandbox
Use local to prove the loop and the network policy. Production sessions, overnight work, and Mac-only steps do not belong on a laptop that sleeps when the lid shuts.

Four cloud shapes

Write “cloud” as four options so the meeting does not collapse into “buy 64 cores”:

  1. OpenAI-hosted sandbox: the app opens a session; OpenAI provisions Linux. Good first path for scripts and artifacts. Egress is on by default; you can disable it or allow-list hosts. Secrets go in the vault — not OPENAI_API_KEY inside the sandbox.
  2. Partner sandbox: docs name Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle Cloud, Runloop, Vercel. Use this for a ready GPU, preview URL, snapshot, or credentials that should live at the provider. You pay their time/RAM plus OpenAI tokens.
  3. Self-hosted Linux: you bring the image and workspace; each session needs its own executor. The box needs stable outbound access, not inbound from the public internet. Fits a VPC, a private package index, or deps that cannot enter the hosted image.
  4. Apple silicon node: only the steps the hosted sandbox cannot do. Keep it off the GPU trainer; see dual-track deploy. Zilmac cloud Mac mini is billed by day / week / month; the entry M4 16 GB plan is about $96/month (live price on the plans page).

No single cloud covers every row. A Q&A agent can skip the sandbox. A new official GPU SKU after the keynote would replace row 2 or 3 — it would not retire the Mac checklist.

How to split the bill

At least four lines. The method is still monthly cost per task. Figures are from the 27 September pricing page; refresh after the keynote.

Line Meter Order of magnitude (as of 2026-09-27)
Model tokens Input / cache / output of the selected model gpt-6-astra short context about $10 / $50 per 1M; gpt-6-sol about $2 / $10. Whole-request uplift above 272K input
Built-in tools Per call plus retrieved content tokens Web search about $10 / 1k calls; file search about $2.50 / 1k, storage $0.10/GB-day after 1 GB free
Hosted container Memory tier, per minute (5-minute minimum) 1 / 4 / 16 / 64 GB listed at $0.03 / $0.12 / $0.48 / $1.92 per 20-minute session; eligible sessions bill by the minute
Your machines Self-hosted Linux, partner time, or a cloud Mac Independent of tokens. Mac is a rental period, not a token meter

A month you can actually compute (one person, 20 sessions on a workday, ~50k input + 15k output, hosted 1 GB for 15 minutes):

  • Astra: 440 × ($0.50 in + $0.75 out) ≈ $550 tokens.
  • Same work on Sol: 440 × ($0.10 + $0.15) ≈ $110 tokens.
  • 1 GB container: $0.03 / 20 min ⇒ ~$0.0015/min × 15 × 440 ≈ $10.
  • 200 web searches: about $2.

At that density the container is a rounding error and the model choice is the month. Move to a 16 GB sandbox, hour-long sessions, and four sub-agents and the container line rises — it still rarely beats Astra tokens. “No extra API fee” does not rescue a wrong default model.

Self-hosted Linux deletes the container line, not the token line. A cloud Mac should appear only on projects that need Xcode or a device. Using it for pure Linux scripts is paying a monthly lease for work that can die in minutes.

A one-page chooser

The job Loop Runtime Specialty box
Q&A, docs, MCP — no shell Agents API none or hosted with network off None
Edit files, run tests, emit artifacts Agents API Hosted 1–4 GB, or local Docker None
Private index, VPC, secrets that never leave Agents API Self-host or partner, outbound WebSocket Always-on Linux — GPU optional
Local GPU inference or a training sidecar Agents API or your own loop Partner GPU or your GPU host Keep it off the Mac builder
iOS / macOS build, signing, device Agents API only splits the work Linux sandbox stops at scripts Fixed Apple silicon — checklist

Solo: local Docker plus a hosted sandbox; default to Sol (or cheaper); keep Astra for long jobs. Small team: add one always-on Linux executor and rent a Mac node only for Mac steps. Do not lock a 12-month fat box because “the keynote might ship a GPU SKU.”

The two weeks after the keynote

After 29 September, change things in this order — do not order hardware first:

  1. Read the docs, not the slide model name. Capture the new default ID and price, then re-run the month above.
  2. Tag current flows as four kinds: no sandbox / hosted Linux / self-host / Mac. AgentKit Builder and Evals that still run before 30 November get their own migration track, not a new chassis.
  3. Land one hosted minimum: edit a file, emit an artifact, write down latency and container minutes.
  4. Try one self-hosted executor on the corp network. Confirm outbound WebSocket. If it fails, stay hosted — do not blame “not enough server.”
  5. Keep Mac work on a separate list: signing, Xcode, devices, notarization. Missing those four? Do not rent a cloud Mac.
  6. A/B the default one tier down: same 20 tasks on Astra vs Sol; write quality and price before you pick a default.

How to watch and which timezone: still the DevDay preview. You can choose an environment type the same night. Which box to buy can wait for the price page.

The loop can be hosted. Xcode and a device cannot live in a Linux container

Agents API means you buy one fewer “orchestration server.” The hosted sandbox runs scripts. When the job hits signing or a device, park that step on fixed Apple silicon — not on the office laptop or the GPU trainer.

Classify the four workflows first, then decide whether to rent a node. — See cloud Mac plans

FAQ

Do I have to buy a GPU server for agents after DevDay?

No. The loop sits on the Agents API. The official runtime is a Linux container. A GPU appears only if you want local inference or a training sidecar. Most apps pay tokens and container minutes before they pay for a card.

Is a laptop enough?

Enough to prove a prompt and a minimum example. Not enough for production: the lid drops the executor, keys share a browser profile, and iOS signing will not fit. Overnight sessions belong on hosted or always-on self-host.

Which is cheaper, hosted sandbox or self-host?

Hosted bills by RAM tier and minute; 1 GB is often a few dollars to low tens per month. Self-host drops that line and adds a machine plus ops. Both pay the same tokens. Get the default model right before you argue about the container.

When do I need a cloud Mac?

When the job includes Xcode, notarization, a device, or a bug that only reproduces on macOS. Pure Python / Node scripts stay on hosted or self-hosted Linux. US residency and “self-host ≠ ZDR” apply to both executors.

Limited Offer

Zilmac

The loop can be hosted. Xcode and a device cannot live in a Linux container.

Back to Home
Limited Offer View Plans