By 2026, the hard part of shipping an AI agent is no longer “GPT versus Gemini versus Claude.” Reliability comes from how models emit structured Function Calling requests, how JSON Schema constrains arguments, and how MCP plugs tools and context into the host runtime. Those layers together form a stack you can audit, version, and swap models in.
This guide is for teams turning chatbots into agents that actually execute work. It splits responsibilities, walks one full tool loop, and shows where build tools should run on a cloud Mac. For cost, see Claude / GPT / Gemini pricing via OpenRouter; for workspace isolation, agent virtual filesystems; for cross-session state, agent memory comparison.
Split the stack: the model is not the product
Treating an agent as “a talkative model” hides the three layers that actually carry reliability:
- Reasoning engine: GPT, Gemini, Claude—plan, pick tools, interpret observations.
- Call contract: Function Calling (tool use) plus JSON Schema—what can be called and how arguments look.
- Runtime slots: Model Context Protocol (MCP) attaches filesystems, repos, browsers, and internal APIs to the agent host.
The 2025–2026 consensus: models may change; contracts and slots should stay stable. Do not freeze business code to one vendor’s private function names. Design around a tool list, schemas, and an execution sandbox.
How GPT, Gemini, and Claude fit agent roles
All three support tool calling. Differences are mostly default style, multimodal entry points, long context, and how conservative they are about risky calls—not whether function calling exists.
- GPT: mature tool-loop docs; OpenAI Function Calling plus a large Responses / Chat Completions ecosystem—often the default orchestrator or gateway backend.
- Gemini: strong multimodal and long context; Gemini function calling fits agents that ingest docs, screenshots, and repo trees together.
- Claude: steady tool use in computer/code settings; Anthropic tool use fits high-risk code edits where extra refusal of overreach matters.
Production usually means routing, not a single-model creed: strong reasoning for planning, cheaper models for bulk structured extraction, and a conservative tier plus human confirm for sensitive writes. Routing and bills: multi-model cost comparison.
Function Calling: the legal interface to the world
Function Calling does not mean “the model runs the function.” It means the model emits a structured call intent, and the host decides whether to execute:
- Model: pick a tool name, fill arguments, keep planning from observations.
- Host: auth, rate limits, validation, real I/O, append results to the message list.
Field names differ (tools / function / input_schema), but the loop is the same: declare tools → model returns a tool call → execute → tool result → reason again. Do not let the model paste shell strings as a “tool”; that bypasses Schema and hands injection risk to the prompt.
Engineering baseline
- Stable names, one side-effect class per tool.
- Split reads and writes; writes default to confirm or sandbox.
- Keep
call_idon results for audit and retry.
JSON Schema: the type system for tools
In an agent stack, JSON Schema is compile-time types plus runtime validation. Even a strong model should fail closed: missing required fields, bad enums, extra properties must fail before business logic, then the model retries on the validation error.
Version each tool’s parameters / input_schema as a real schema (JSON Schema), not a prompt line that says “please pass repo and branch.”
Constraints that belong in Schema
type,required,additionalProperties: falseto cut hallucinated fields.- Paths, URLs, enums via
pattern/enum—not free text posing as IDs. - Huge payloads (logs, diffs) should not be arguments; write a workspace file and return a path, matching on-demand VFS reads.
- Version in the schema or tool suffix (
deploy_v2) so old and new agents do not share a fuzzy contract.
Validate after the model output and before real side effects. “Please emit JSON” in a system prompt is not production control.
MCP: a universal slot for tools and context
MCP sits above Function Calling and solves discovery and transport: tools need not be hard-coded; an MCP server can list tools, resources, and prompts. Hosts (Claude Code, IDE agents, custom orchestrators) connect many servers over one protocol.
Compared with a plugin SDK per model vendor, MCP gives you:
- Pluggable tools: Git, browser, and ticket systems as separate servers; the host keeps connections and policy.
- Addressable context: resource URIs avoid stuffing whole files into the prompt.
- Model decoupling: the same MCP catalog can feed GPT, Gemini, or Claude function-calling layers.
You still map MCP tools to Function Calling declarations with JSON Schema, plus path allowlists on the host. MCP is a standard plug, not the security boundary; sandbox, secrets, and audit stay on the host.
One full loop: from intent to tool result
Example: “open a fix PR from a failed iOS build log.” Typical order:
- User intent; host injects policy (no certs, no production secrets).
- Model reads the tool list (static registry or MCP
list_tools). - Emits
fetch_ci_log; arguments pass JSON Schema. - Executor fetches logs in a sandbox or on a cloud Mac; oversized output spills to VFS.
- Result returns; model may call
apply_patchwith confirm or a read-only diff preview. - Facts such as “profile expired last time” go to memory, not schema args. See memory options.
Failures should be structured (validation, permission, timeout), not a blob of prose, or retries become guesswork.
Comparison and selection
| Layer | Owns | Does not own | Typical 2026 choice |
|---|---|---|---|
| GPT / Gemini / Claude | Planning, tool choice, explaining observations | Real I/O, authorization | Route by task; keep swap ability |
| Function Calling | Intent as a tool request with an id | Business permission model | Vendor tool use, one host adapter |
| JSON Schema | Argument shape, required, enums | Runtime side effects | Versioned schema, validate before execute |
| MCP | Discover tools/resources, standard transport | Replacing sandboxes and secret stores | One server per system + host policy |
Production: validation, auth, fallbacks
Before go-live, pass at least four gates:
- Schema adversarial tests: missing fields, extra fields, bad enums—host rejects, model repairs from the error.
- Unauthorized tools: unknown names, stale MCP servers, forged
call_idmust fail. - Side-effect isolation: build, signing, and deploy stay off the chat process; agents do not hold production credentials by default.
- Observability: log tool name, schema version, latency, success / validation fail / auth fail.
OWASP lists over-authorization and prompt injection as core LLM risks, so “the model asked, so we ran it” is not a policy. Use OWASP LLM Top 10 as an audit anchor.
Running the tool chain on a cloud Mac
For iOS / cross-platform teams, MCP catalogs often include xcodebuild, simulators, and cert-adjacent tools. Those bind to macOS and Apple’s toolchain and do not belong in a generic Linux function sandbox.
A practical split:
- Chat and planning: GPT / Gemini / Claude anywhere.
- Repo I/O: VFS or MCP filesystem with path allowlists.
- Real builds and devices: a controlled MCP server or CI job on a Zilmac cloud Mac, secrets staying on the build host.
The model stack can change; schemas and MCP catalogs stay put; Apple side effects land in recyclable cloud Mac sessions instead of a laptop.
FAQ
Are Function Calling and MCP duplicates?
No. Function Calling is how the model emits a tool request. MCP is how the host discovers and connects tools and resources. Each MCP tool still maps to a Function Calling declaration.
Do we have to wire all three model families?
No. Stabilize contracts and MCP on one vendor, then route a second model through the same schemas. Connecting three vendors “for eval” only widens an unvalidated tool surface.
Can JSON Schema live only in the system prompt?
Not as the only control. Prompts help the model fill args; rejecting illegal args must happen in a host validator. Otherwise injection or drift hits the backend directly.
How do we avoid locking to one vendor API?
Keep an internal shape of tool name + JSON Schema + execution result. GPT / Gemini / Claude are adapters. MCP servers expose one catalog to the host; billing and failover sit on a gateway.
Models can swap; the build host cannot be vague
Once GPT, Gemini, and Claude share Function Calling and MCP contracts, iOS delivery still stalls on Apple’s toolchain. A Zilmac cloud Mac fits as a controlled executor: the agent dispatches under Schema; signing and xcodebuild run on isolated macOS.
Give the agent an auditable build slot without owning a physical Mac. — View cloud Mac plans