Comparisons 8 min read

MCP vs Direct API Calls: Does Your Agent Need MCP? (2026)

MCP vs direct API: MCP tool definitions can eat 55K tokens before your agent reads a request. Here is when MCP is worth it and when direct API wins.

Shabnam Katoch

Shabnam Katoch

Growth Head

MCP vs Direct API Calls: does your agent need MCP?
Free forever

Your agent. Working. Not broken.

One AI agent that just works.

No silent failures. Free forever, not a trial.

Start free

No credit card · No Docker · No config files

Five MCP servers can burn 55,000 tokens before your agent reads a single word you typed. That doesn't make MCP a bad idea. It makes it an expensive default. Here's how to tell when it pays for itself.

On 20 September, a post titled "MCP was always a bad idea?" climbed to 335 points on Hacker News. The argument: models are now smart enough to call APIs and CLIs directly, so most MCP servers should be deleted.

Nine days later, Earendil, whose Pi agent had proudly refused to support MCP, shipped MCP in Pi's core. That post passed 680 points.

Both teams are smart. Both are partly right. So instead of picking a side, we measured where the cost actually comes from in the MCP vs direct API debate, and when each one earns its place in an agent.

MCP vs direct API: the 30-second answer

MCP (Model Context Protocol) is a standard way to hand an agent a set of tools, each described by a name, a description and a JSON schema. A direct API call skips that layer: the agent, or code you wrote, calls the service's HTTP API itself. If MCP is new to you, our plain-English guide to the Model Context Protocol covers the basics.

MCP serverDirect API call
Who writes the integrationUsually the vendor or communityYou, or the agent writes it on the fly
AuthHandled once by the server, often OAuthYou manage keys and tokens
Token costTool definitions sit in context every turn, unless your client defers themOnly what you send and get back
PortabilitySame server works in Claude, ChatGPT, Cursor and other clientsTied to your code
Best forMany tools, many clients, non-developersA few services, called often, at volume

MCP buys you portability and someone else's integration work. You pay for it in context.

What MCP actually costs in tokens

This is the part most opinion pieces skip, so here are real numbers. Anthropic published measurements of common MCP servers and how many tokens their tool definitions take up:

MCP serverToolsApprox. tokens
GitHub35~26,000
Slack11~21,000
JiraNot stated~17,000
Sentry5~3,000
Grafana5~3,000
Splunk2~2,000

GitHub, Slack, Sentry, Grafana and Splunk together come to 58 tools and about 55,000 tokens before the conversation even starts. Anthropic says it has seen tool definitions consume 134,000 tokens internally before optimization.

Here's the weird part. Slack has a third as many tools as GitHub, but nearly as many tokens. Tool count tells you almost nothing. Schema verbosity is what you pay for.

Why it compounds in an agent loop

An agent doesn't call the model once. It calls it every step: read the request, call a tool, read the result, call another tool. Unless your client defers tool loading, those 55,000 tokens ride along on every one of those calls.

Take an illustrative price of $3 per million input tokens. That's about $0.17 per model call just for tool definitions. A 10-step task pays it ten times: $1.65. At 300 tasks a month, that's roughly $500 in tool descriptions alone, before any actual work. Prompt caching cuts this a lot, since most providers bill cached reads at a fraction of the normal price, but it doesn't make it zero, and it doesn't give you the context window back.

That last point matters more than the bill. Tokens spent on tool schemas are tokens your agent can't spend on your documents, your history or its own reasoning. We saw the same pattern in self-hosted agents, where heartbeats and hidden token overhead quietly ate budgets.

Results are a second leak. In Anthropic's example, piping a two-hour meeting transcript from one tool to another through the model added about 50,000 tokens, because the full transcript passed through the context twice.

Where 55K tokens go: five MCP servers' tool definitions (GitHub 26K, Slack 21K, Sentry 3K, Grafana 3K, Splunk 2K) sent every turn, filling the context window before the first message

The case against MCP

The anti-MCP argument, at its strongest, goes like this.

MCP arrived in November 2024, when models were much weaker. It solved a real problem: getting agents to use external services at all. But models now write and run code well. They can read an API's docs, call it, and compose several services in a short script.

And most remote MCP servers are thin wrappers around APIs that already exist. If your agent has a sandbox and can run code, wrapping that API in a protocol adds a layer, adds tokens, and adds one more thing to monitor.

There's also a whole industry of fixes now: gateways, tool routers and search-and-execute layers that exist mainly to stop MCP from flooding the context. The critics' point is simple. If you need that much scaffolding to make a protocol usable, maybe you don't need the protocol.

For a single agent that hits two APIs a thousand times a day, they're right. A direct API call in tested code is cheaper, faster and easier to debug than a 26,000-token tool catalogue.

Why MCP isn't dead

Now the other side, and it's stronger than the hot takes admit.

First, the token problem has largely been fixed at the client level. Anthropic's Tool Search Tool loads one small search tool upfront (about 500 tokens) and pulls in only the 3 to 5 tools a task needs. In its tests, that took a 50-plus-tool setup from about 77,000 tokens of context down to about 8,700, an 85% cut. Accuracy went up too: Opus 4.5 improved from 79.5% to 88.1% on Anthropic's MCP evaluations with tool search on.

Second, code mode. Instead of the model calling tools one by one and reading every result, it writes a short script that calls the tools in a sandbox and returns only the answer. Anthropic reported one workflow dropping from 150,000 tokens to 2,000 with this approach. That's what Earendil built into Pi alongside MCP.

Third, MCP does things raw APIs don't. It handles auth once, so every client doesn't need its own OAuth flow. The same server works across Claude, ChatGPT, Cursor and other clients. And since December 2025 it's been governed by the Agentic AI Foundation under the Linux Foundation, not by one company.

The old MCP dumped every tool into context. Modern MCP clients load tools on demand. Most of the "token bloat" criticism is aimed at the first one.

Load everything vs load on demand: all tools upfront costs about 77K tokens, while tool search loads only the tools a task needs for about 8.7K, leaving the rest of the context free

If you'd rather spend your time on what the agent does than on which protocol it speaks, that's what we built BetterClaw for. Plans start at $19 a month, you bring your own model key with zero markup, and there's a 7-day money-back guarantee. Start with BetterClaw and connect your first service in a couple of clicks.

When to use MCP vs direct API calls

Stay with me here, because the answer isn't one or the other. It's which tool, for which job.

Use MCP when:

  • You need many services and don't want to write and maintain each integration yourself.
  • The same tools should work across several AI clients, or for people who don't write code.
  • Your client supports deferred loading or tool search, so unused tools don't sit in context.
  • The service ships an official MCP server and has no clean API or CLI.

Use a direct API call when:

  • One or two services do most of the work, and you call them constantly.
  • The agent runs at volume, where every repeated token shows up on the bill.
  • You need exact control over requests, retries, rate limits and error handling.
  • Your agent can run code in a sandbox, so it can call the API itself.

The hybrid most production teams land on: MCP for breadth, direct calls for the hot path, and code mode in between to keep results out of the context. We compared a related choice, Agent Skills vs MCP, if you're deciding how to package what an agent knows rather than what it can reach.

MCP or direct API? If a service is called constantly, use a direct API call; if not and you need many services or clients, use MCP with tool search; otherwise either works

How we handle this at BetterClaw

We think about this at the level of each agent, not each workspace.

In BetterClaw, connectors handle sign-in to services like Gmail, Zendesk or PostHog, and access is granted per account, per agent. Your support agent gets the helpdesk connector. Your reporting agent gets analytics. Neither carries the other's tools around. The cheapest tool definition is the one that isn't there.

For anything a connector doesn't cover, skills and the secrets vault let an agent call a service's API directly with a key it reads by reference, without ever seeing the raw value. So you can mix both approaches on one agent without writing glue code.

The protocol was never the problem

Strip away the hot takes and the MCP vs direct API debate is really about one thing: what you let into your agent's context window.

An agent with 58 tools it doesn't need is slower, pricier and more likely to pick the wrong one, whatever protocol delivers them. An agent with three well-described tools does fine over MCP or plain HTTP.

Don't ask "MCP or API?" Ask "does this agent need to see this tool right now?"

If any of this resonated, give BetterClaw a try. Plans start at $19 a month for one agent, Pro is $49 for five, and every plan has a 7-day money-back guarantee. Bring your own key from 28+ providers, give each agent only the tools it needs, and pay your provider directly with zero markup. See full pricing. We handle the infrastructure. You handle the interesting part.

Frequently Asked Questions

What is the difference between MCP and a direct API call?

MCP (Model Context Protocol) is a standard way to give an agent tools, each described by a name, description and JSON schema, usually served by a vendor or community server. A direct API call skips that layer: your code, or the agent itself, calls the service's HTTP API. MCP saves integration work and works across AI clients; direct calls are leaner and give you more control.

How does MCP compare to direct API calls on token cost?

MCP tool definitions can be large. Anthropic measured GitHub's MCP server at about 26,000 tokens and a five-server setup at about 55,000, sent on every model call unless your client defers loading. Direct API calls cost only the request and response you send, though tool search and code mode now cut MCP overhead by 85% or more.

How do I reduce MCP token bloat in my agent?

Give each agent only the servers and tools it actually uses. Use a client that supports tool search or deferred loading, so definitions load on demand. And where possible, use code mode, so the agent filters large results in a sandbox instead of reading them all.

Is MCP worth it for AI agents?

It's worth it when you need many services, want the same tools across several AI clients, or don't want to write and maintain integrations. It's usually not worth it for one or two high-volume services, where a direct API call in tested code is cheaper and easier to debug. Most production agents use both.

Is MCP safe to use in production?

It can be, but every MCP server is code with access to your accounts, so treat it like any dependency. Use official or well-maintained servers, scope OAuth permissions narrowly, and give each agent only the tools it needs. The protocol is now governed by the Agentic AI Foundation under the Linux Foundation, but the security of each server is still up to whoever runs it.

Every model above, one platform.

All models compared work on BetterClaw via BYOK. Switch between them in settings. No config changes.

Try it free
Tags:mcp vs apimcp vs direct apido agents need mcpmcp token bloatmcp overheadis mcp worth itmodel context protocol problems
Share this article
Was this helpful?