Five MCP servers can burn 55,000 tokens before your agent reads a single word you typed. That doesn't make MCP a bad idea. It makes it an expensive default. Here's how to tell when it pays for itself.
On 20 September, a post titled "MCP was always a bad idea?" climbed to 335 points on Hacker News. The argument: models are now smart enough to call APIs and CLIs directly, so most MCP servers should be deleted.
Nine days later, Earendil, whose Pi agent had proudly refused to support MCP, shipped MCP in Pi's core. That post passed 680 points.
Both teams are smart. Both are partly right. So instead of picking a side, we measured where the cost actually comes from in the MCP vs direct API debate, and when each one earns its place in an agent.
MCP vs direct API: the 30-second answer
MCP (Model Context Protocol) is a standard way to hand an agent a set of tools, each described by a name, a description and a JSON schema. A direct API call skips that layer: the agent, or code you wrote, calls the service's HTTP API itself. If MCP is new to you, our plain-English guide to the Model Context Protocol covers the basics.
| MCP server | Direct API call | |
|---|---|---|
| Who writes the integration | Usually the vendor or community | You, or the agent writes it on the fly |
| Auth | Handled once by the server, often OAuth | You manage keys and tokens |
| Token cost | Tool definitions sit in context every turn, unless your client defers them | Only what you send and get back |
| Portability | Same server works in Claude, ChatGPT, Cursor and other clients | Tied to your code |
| Best for | Many tools, many clients, non-developers | A few services, called often, at volume |
MCP buys you portability and someone else's integration work. You pay for it in context.
What MCP actually costs in tokens
This is the part most opinion pieces skip, so here are real numbers. Anthropic published measurements of common MCP servers and how many tokens their tool definitions take up:
| MCP server | Tools | Approx. tokens |
|---|---|---|
| GitHub | 35 | ~26,000 |
| Slack | 11 | ~21,000 |
| Jira | Not stated | ~17,000 |
| Sentry | 5 | ~3,000 |
| Grafana | 5 | ~3,000 |
| Splunk | 2 | ~2,000 |
GitHub, Slack, Sentry, Grafana and Splunk together come to 58 tools and about 55,000 tokens before the conversation even starts. Anthropic says it has seen tool definitions consume 134,000 tokens internally before optimization.
Here's the weird part. Slack has a third as many tools as GitHub, but nearly as many tokens. Tool count tells you almost nothing. Schema verbosity is what you pay for.
Why it compounds in an agent loop
An agent doesn't call the model once. It calls it every step: read the request, call a tool, read the result, call another tool. Unless your client defers tool loading, those 55,000 tokens ride along on every one of those calls.
Take an illustrative price of $3 per million input tokens. That's about $0.17 per model call just for tool definitions. A 10-step task pays it ten times: $1.65. At 300 tasks a month, that's roughly $500 in tool descriptions alone, before any actual work. Prompt caching cuts this a lot, since most providers bill cached reads at a fraction of the normal price, but it doesn't make it zero, and it doesn't give you the context window back.
That last point matters more than the bill. Tokens spent on tool schemas are tokens your agent can't spend on your documents, your history or its own reasoning. We saw the same pattern in self-hosted agents, where heartbeats and hidden token overhead quietly ate budgets.
Results are a second leak. In Anthropic's example, piping a two-hour meeting transcript from one tool to another through the model added about 50,000 tokens, because the full transcript passed through the context twice.

The case against MCP
The anti-MCP argument, at its strongest, goes like this.
MCP arrived in November 2024, when models were much weaker. It solved a real problem: getting agents to use external services at all. But models now write and run code well. They can read an API's docs, call it, and compose several services in a short script.
And most remote MCP servers are thin wrappers around APIs that already exist. If your agent has a sandbox and can run code, wrapping that API in a protocol adds a layer, adds tokens, and adds one more thing to monitor.
There's also a whole industry of fixes now: gateways, tool routers and search-and-execute layers that exist mainly to stop MCP from flooding the context. The critics' point is simple. If you need that much scaffolding to make a protocol usable, maybe you don't need the protocol.
For a single agent that hits two APIs a thousand times a day, they're right. A direct API call in tested code is cheaper, faster and easier to debug than a 26,000-token tool catalogue.
Why MCP isn't dead
Now the other side, and it's stronger than the hot takes admit.
First, the token problem has largely been fixed at the client level. Anthropic's Tool Search Tool loads one small search tool upfront (about 500 tokens) and pulls in only the 3 to 5 tools a task needs. In its tests, that took a 50-plus-tool setup from about 77,000 tokens of context down to about 8,700, an 85% cut. Accuracy went up too: Opus 4.5 improved from 79.5% to 88.1% on Anthropic's MCP evaluations with tool search on.
Second, code mode. Instead of the model calling tools one by one and reading every result, it writes a short script that calls the tools in a sandbox and returns only the answer. Anthropic reported one workflow dropping from 150,000 tokens to 2,000 with this approach. That's what Earendil built into Pi alongside MCP.
Third, MCP does things raw APIs don't. It handles auth once, so every client doesn't need its own OAuth flow. The same server works across Claude, ChatGPT, Cursor and other clients. And since December 2025 it's been governed by the Agentic AI Foundation under the Linux Foundation, not by one company.
The old MCP dumped every tool into context. Modern MCP clients load tools on demand. Most of the "token bloat" criticism is aimed at the first one.

If you'd rather spend your time on what the agent does than on which protocol it speaks, that's what we built BetterClaw for. Plans start at $19 a month, you bring your own model key with zero markup, and there's a 7-day money-back guarantee. Start with BetterClaw and connect your first service in a couple of clicks.
When to use MCP vs direct API calls
Stay with me here, because the answer isn't one or the other. It's which tool, for which job.
Use MCP when:
- You need many services and don't want to write and maintain each integration yourself.
- The same tools should work across several AI clients, or for people who don't write code.
- Your client supports deferred loading or tool search, so unused tools don't sit in context.
- The service ships an official MCP server and has no clean API or CLI.
Use a direct API call when:
- One or two services do most of the work, and you call them constantly.
- The agent runs at volume, where every repeated token shows up on the bill.
- You need exact control over requests, retries, rate limits and error handling.
- Your agent can run code in a sandbox, so it can call the API itself.
The hybrid most production teams land on: MCP for breadth, direct calls for the hot path, and code mode in between to keep results out of the context. We compared a related choice, Agent Skills vs MCP, if you're deciding how to package what an agent knows rather than what it can reach.

How we handle this at BetterClaw
We think about this at the level of each agent, not each workspace.
In BetterClaw, connectors handle sign-in to services like Gmail, Zendesk or PostHog, and access is granted per account, per agent. Your support agent gets the helpdesk connector. Your reporting agent gets analytics. Neither carries the other's tools around. The cheapest tool definition is the one that isn't there.
For anything a connector doesn't cover, skills and the secrets vault let an agent call a service's API directly with a key it reads by reference, without ever seeing the raw value. So you can mix both approaches on one agent without writing glue code.
The protocol was never the problem
Strip away the hot takes and the MCP vs direct API debate is really about one thing: what you let into your agent's context window.
An agent with 58 tools it doesn't need is slower, pricier and more likely to pick the wrong one, whatever protocol delivers them. An agent with three well-described tools does fine over MCP or plain HTTP.
Don't ask "MCP or API?" Ask "does this agent need to see this tool right now?"
If any of this resonated, give BetterClaw a try. Plans start at $19 a month for one agent, Pro is $49 for five, and every plan has a 7-day money-back guarantee. Bring your own key from 28+ providers, give each agent only the tools it needs, and pay your provider directly with zero markup. See full pricing. We handle the infrastructure. You handle the interesting part.
Frequently Asked Questions
What is the difference between MCP and a direct API call?
MCP (Model Context Protocol) is a standard way to give an agent tools, each described by a name, description and JSON schema, usually served by a vendor or community server. A direct API call skips that layer: your code, or the agent itself, calls the service's HTTP API. MCP saves integration work and works across AI clients; direct calls are leaner and give you more control.
How does MCP compare to direct API calls on token cost?
MCP tool definitions can be large. Anthropic measured GitHub's MCP server at about 26,000 tokens and a five-server setup at about 55,000, sent on every model call unless your client defers loading. Direct API calls cost only the request and response you send, though tool search and code mode now cut MCP overhead by 85% or more.
How do I reduce MCP token bloat in my agent?
Give each agent only the servers and tools it actually uses. Use a client that supports tool search or deferred loading, so definitions load on demand. And where possible, use code mode, so the agent filters large results in a sandbox instead of reading them all.
Is MCP worth it for AI agents?
It's worth it when you need many services, want the same tools across several AI clients, or don't want to write and maintain integrations. It's usually not worth it for one or two high-volume services, where a direct API call in tested code is cheaper and easier to debug. Most production agents use both.
Is MCP safe to use in production?
It can be, but every MCP server is code with access to your accounts, so treat it like any dependency. Use official or well-maintained servers, scope OAuth permissions narrowly, and give each agent only the tools it needs. The protocol is now governed by the Agentic AI Foundation under the Linux Foundation, but the security of each server is still up to whoever runs it.




