Updated September 28, 2026. Prices checked against Google Cloud's Agent Platform, Generative AI, Agent Search and CX Agent Studio pricing pages on September 28, 2026.
How much does Vertex AI cost? Vertex AI, now called Gemini Enterprise Agent Platform, has no monthly fee. You pay per use across four meters. Agent Compute is $0.085 per vCPU-hour (50 free each month) and Agent Memory is $0.009 per GiB-hour (100 free). Sessions and Memory Bank storage is $0.30 per GiB-month (1 free), billed on this structure since September 1, 2026. Agent Search is $1.50-$4.00 per 1,000 queries. Model tokens are extra: Gemini 3.8 Flash is $0.75 in and $3.75 out per million until December 31, 2026. A small support bot costs about $90 a month. A busy agent on Gemini 3.1 Pro can pass $10,000.
Vertex AI pricing at a glance (September 2026)
| Component | Price | Free tier |
|---|---|---|
| Agent Runtime compute (Agent Compute) | $0.085 per vCPU-hour | First 50 vCPU-hours / month |
| Agent Runtime memory (Agent Memory) | $0.009 per GiB-hour | First 100 GiB-hours / month |
| Sessions and Memory Bank storage (Agent Storage) | $0.30 per GiB-month | First 1 GiB-month |
| Sessions and Memory Bank operations | 1 vCPU-hour ($0.085) per 3M reads, per 1M writes | Covered by the 50 free vCPU-hours |
| Agent Search, Standard | $1.50 per 1,000 queries | 10,000 queries / month |
| Agent Search, Enterprise | $4.00 per 1,000 queries | 10,000 queries / month |
| Grounding with Google Search | $14 per 1,000 grounding queries | 5,000 / month across Gemini 3 models |
| Gemini 3.8 Flash | $0.75 input / $3.75 output per 1M tokens (promo until Dec 31, 2026) | None on Agent Platform |
| New Google Cloud customers | $300 credit |
Model tokens usually dominate the bill, and the runtime is usually the smallest line. Worked examples below. Training, prediction endpoints, Vector Search, RAG Engine and CX Agent Studio are in the quick reference further down.
Vertex AI is now Gemini Enterprise Agent Platform: what got renamed
Google renamed Vertex AI to Gemini Enterprise Agent Platform around Cloud Next 2026 (late April), folding Agentspace in at the same time. The services are the same and existing customers didn't have to migrate. Only the names on the pricing page and the console changed, so both sets of names still show up in docs, invoices and search results:
| Old name (Vertex AI) | New name (Agent Platform) |
|---|---|
| Vertex AI / Vertex AI Agent Builder | Gemini Enterprise Agent Platform |
| Agent Engine | Agent Runtime |
| Vertex AI Search | Agent Search |
| Agent Engine compute and memory SKUs | Agent Compute (vCPU-hours) and Agent Memory (GiB-hours) |
| Sessions / Memory Bank per-event SKU | Agent Storage (GiB-months) plus operations billed in vCPU-hours |
If you're searching for "Vertex AI pricing" or "Google Agent Platform pricing", it's the same price list.
Agent Runtime pricing (formerly Agent Engine)
Agent Runtime is the managed environment your agent runs in. It bills on two resources, rounded to the nearest second:
- Agent Compute: $0.085 per vCPU-hour. The first 50 vCPU-hours each month are free.
- Agent Memory: $0.009 per GiB-hour. The first 100 GiB-hours each month are free.
Idle time between turns, while the agent waits for the next prompt, is not billed. Code Execution and Computer Use are billed on the same two resources.
If you run agents all month, Google's Gemini Enterprise Flexible Savings Plans cut all three resources by 10% on a 1-year commitment and 20% on a 3-year one:
| Resource | On demand | 1-year plan | 3-year plan |
|---|---|---|---|
| Agent Compute | $0.085 per vCPU-hour | $0.0765 | $0.068 |
| Agent Memory | $0.009 per GiB-hour | $0.0081 | $0.0072 |
| Agent Storage | $0.30 per GiB-month | $0.27 | $0.24 |
The free allowances apply on every plan. Google bills storage per GiB-hour ($0.000410959 on demand), which works out to the monthly figures above. You may need to opt in to a savings plan before the discount applies.
Agent Compute is also the unit Google uses to bill several request-based services. For example, one vCPU-hour ($0.085) covers 15,000 calls through Agent Gateway, and the same rate covers 15,000 agent-response evaluations.
Sessions and Memory Bank pricing (billed from Sept 1, 2026)
Sessions (conversation history and state) and Memory Bank (long-term memory) moved to a storage-plus-operations model. Billing on this structure started September 1, 2026:
- Storage: $0.30 per GiB-month of data stored, with Memory Bank counting revisions. The first GiB-month is free.
- Reads: 1 vCPU-hour ($0.085) per 3 million read requests.
- Writes: 1 vCPU-hour ($0.085) per 1 million write requests.
- Memory generation and embedding: tokens billed separately under the model's own SKU.
For most agents this meter is tiny. 60,000 conversations with 10 writes and 20 reads each come to about nine cents in operations. The old flat per-1,000-events rate is gone, so if you budgeted against it, rebuild the estimate.
Agent Search pricing (formerly Vertex AI Search)
Agent Search is the retrieval service that grounds your agent in your own documents. On the default (General) pricing model:
| Edition | Price per 1,000 queries | What you get |
|---|---|---|
| Standard Edition | $1.50 | Unstructured and structured search |
| Enterprise Edition | $4.00 | Adds website search and core Generative Answers (answers, summaries, follow-ups) |
| Advanced Generative Answers (AI Mode) | +$4.00, on either edition | Suggested follow-ups, complex and multi-turn queries, multimodal input |
- Free usage: 10,000 queries per account, every month. It doesn't cover Advanced Generative Answers.
- Indexed data: $5.00 per GiB-month (billed hourly), with the first 10 GiB free each month.
- Over 15 million queries a month: Google steers you to Configurable pricing, a monthly subscription for query capacity and storage with add-ons billed per 1,000 queries.
Agent Search is billed separately from Agent Runtime. If your agent grounds answers in your documents, these costs sit on top of compute and tokens.
If you ground on the public web instead, Grounding with Google Search includes 5,000 grounding queries a month across Gemini 3 models, then $14 per 1,000. One user prompt can fire several grounding queries, and each one is billed. Retrieval is the main reason to pick this platform, so expect this meter to run on almost every interaction.
Gemini 3.8 and other Gemini model pricing on Vertex AI
Gemini 3.8 Flash is Google's current Flash model, and on Vertex AI (Agent Platform) it carries an introductory price: $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, rising to $1.50 / $7.50 on January 1, 2027. Gemini 3.7 Flash and 3.6 Flash get the same introductory price. Read the footnote before you budget on it: Google delivers the promotion as "50% credits back on net spend", so the invoice shows the full rate and the discount arrives as a billing credit.
Standard prices per 1 million tokens on the global endpoint, from Google's Generative AI pricing page:
| Model | Input | Cached input | Output | Notes |
|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | Introductory until Dec 31, 2026; $1.50 / $0.15 / $7.50 from Jan 1, 2027 |
| Gemini 3.7 Flash and 3.6 Flash | $0.75 | $0.075 | $3.75 | Same introductory offer and 2027 rates as 3.8 Flash |
| Gemini 3.5 Flash | $1.50 | $0.15 | $9.00 | Now the legacy Flash |
| Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | |
| Gemini 3.1 Flash-Lite | $0.25 | $0.025 | $1.50 | Audio input $0.50 |
| Gemini 3.1 Pro Preview (up to 200K context) | $2.00 | $0.20 | $12.00 | |
| Gemini 3.1 Pro Preview (over 200K context) | $4.00 | $0.40 | $18.00 | |
| Gemini 2.5 Pro (up to 200K / over 200K) | $1.25 / $2.50 | $0.125 / $0.25 | $10.00 / $15.00 | |
| Gemini 2.5 Flash | $0.30 | $0.03 | $2.50 |
How the rate changes with how you call it:
- Regional (non-global) endpoints cost 10% more: Gemini 3.8 Flash is $0.825 / $4.125 on a regional endpoint during the promotion.
- Flex and Batch halve the standard rate: Gemini 3.8 Flash drops to $0.375 in and $1.875 out.
- Priority costs 1.8x: Gemini 3.8 Flash is $1.35 in, and Gemini 3.1 Pro Preview is $3.60 in and $21.60 out.
- Context caching bills cached input at a tenth of the normal input price (a 90% discount).
- Provisioned Throughput buyers get a 50% billing credit on Gemini 3.8, 3.7 and 3.6 Flash spend from August 13 to December 31, 2026.
There is no "Gemini 3.5 Pro" on the price list. The current Pro model on Vertex AI is Gemini 3.1 Pro Preview. Model prices move almost monthly, so confirm the rate in the console before you commit.
Model Garden pricing (Claude, Llama, DeepSeek and more)
Model Garden is the catalogue of non-Gemini models on Agent Platform. Partner and open models offered as managed APIs bill per token, like Gemini, and each provider sets its own rate. Standard prices per 1 million tokens:
| Model | Provider | Input | Output |
|---|---|---|---|
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 |
| Claude Opus 5.5 | Anthropic | $4.00 | $20.00 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 |
| GLM-5.2 | Z.ai | $1.40 | $4.40 |
| Llama 3.3 70B | Meta | $0.72 | $0.72 |
| DeepSeek-V3.2 | DeepSeek | $0.56 | $1.68 |
| Mistral Medium 3 | Mistral AI | $0.40 | $2.00 |
| Llama 4 Maverick | Meta | $0.35 | $1.15 |
| Qwen3-Coder-480B-A35B | Qwen | $0.22 | $1.80 |
| gpt-oss-120b | OpenAI | $0.09 | $0.36 |
Claude prices are for the global endpoint. Regional endpoints cost about 10% more (Opus 5.5 is $4.40 in on a regional endpoint). Note that the newer Opus 5.5 is cheaper than Opus 5. Most models also offer batch pricing at roughly half the rate.
The other way to use Model Garden is to deploy an open model to your own endpoint. That bills per node-hour for the machine and GPU, not per token, and it keeps billing while the endpoint sits idle. Either way, Model Garden charges sit on top of the Agent Runtime meters. For the full catalogue, see Google's Generative AI pricing page.
What a real agent costs per month: 3 worked examples

All three examples assume 1 vCPU and 2 GiB of memory, busy for about one minute per conversation, and one Agent Search query per conversation.
Small: a support bot on Gemini 3.8 Flash. 300 conversations a day, 9,000 a month, averaging 6,000 input and 800 output tokens each.
| Line | Maths | Cost |
|---|---|---|
| Model input | 54M tokens × $0.75 | $40.50 |
| Model output | 7.2M tokens × $3.75 | $27.00 |
| Agent Search (Standard) | 9,000 queries × $1.50 / 1,000 | $13.50 |
| Agent Compute | 150 vCPU-hours − 50 free = 100 × $0.085 | $8.50 |
| Agent Memory | 300 GiB-hours − 100 free = 200 × $0.009 | $1.80 |
| Storage and operations | under the free GiB, about 1 cent of operations | $0.01 |
| Total | about $91 / month |
Medium: an internal agent on Gemini 3.5 Flash. 2,000 conversations a day, 60,000 a month, 8,000 input and 1,000 output tokens each, Enterprise search, 5 GiB of memories.
| Line | Maths | Cost |
|---|---|---|
| Model input | 480M tokens × $1.50 | $720.00 |
| Model output | 60M tokens × $9.00 | $540.00 |
| Agent Search (Enterprise) | 60,000 queries × $4.00 / 1,000 | $240.00 |
| Agent Compute | 1,000 − 50 free = 950 vCPU-hours × $0.085 | $80.75 |
| Agent Memory | 2,000 − 100 free = 1,900 GiB-hours × $0.009 | $17.10 |
| Storage and operations | 4 billable GiB × $0.30, plus operations | $1.29 |
| Total | about $1,600 / month |
Heavy: a customer-facing agent on Gemini 3.1 Pro. 10,000 conversations a day, 300,000 a month, 10,000 input and 1,200 output tokens each, Enterprise search with Advanced Generative Answers, 50,000 Google Search grounding queries, 20 GiB of memories.
| Line | Maths | Cost |
|---|---|---|
| Model input | 3B tokens × $2.00 | $6,000.00 |
| Model output | 360M tokens × $12.00 | $4,320.00 |
| Agent Search (Enterprise + Advanced) | 300,000 queries × $8.00 / 1,000 | $2,400.00 |
| Grounding with Google Search | 50,000 − 5,000 free = 45,000 × $14 / 1,000 | $630.00 |
| Agent Compute | 5,000 − 50 free = 4,950 vCPU-hours × $0.085 | $420.75 |
| Agent Memory | 10,000 − 100 free = 9,900 GiB-hours × $0.009 | $89.10 |
| Storage and operations | 19 billable GiB × $0.30, plus operations | $6.21 |
| Total | about $13,870 / month |
Two things jump out. The runtime is never the problem: even the heavy agent spends only about 4% of its bill there. And the model choice matters more than anything else. Moving the heavy agent from 3.1 Pro to 3.8 Flash would cut about $6,700 a month in tokens at the introductory rate, or about $3,100 once 3.8 Flash goes to $1.50 / $7.50 in January 2027.
Budget for that January step. The small support bot's token lines double from $67.50 to $135, taking it from about $91 to about $159 a month.
The numbers above are the expected path. The danger is the unexpected one. Models deployed to your own endpoint bill per node-hour even when idle, because Agent Platform Inference doesn't scale to zero (Google's pricing page says so directly). One developer on Google's own AI forum reported a £925 bill for March 2026, most of it Vertex AI charges from services that kept running. There is no built-in hard spending cap, so a misconfigured agent bills quietly until someone catches it. Set budget alerts on day one.
Rather have one flat monthly price?
If predictable cost matters more to you than per-token flexibility, a flat-rate platform is a different trade. See flat-rate agent pricing on BetterClaw: you bring your own model keys with no markup, and the platform fee doesn't move with traffic.
Is Vertex AI free? Free tiers and the $300 credit
There is no free plan, but there is more free usage than most write-ups admit:
- Standing monthly allowance: the first 50 vCPU-hours of Agent Compute, 100 GiB-hours of Agent Memory, and 1 GiB-month of Agent Storage, on every account.
- Express mode: a developer with an ordinary Google account can try the platform inside fixed quotas for up to 90 days without adding billing information.
- Agent Search: 10,000 free queries every month (not counting Advanced Generative Answers).
- Grounding with Google Search: 5,000 free grounding queries a month across Gemini 3 models.
- $300 credit for new Google Cloud customers.
Two caveats on the credit, per Google's Free Trial terms. The $300 can't pay for Gemini API usage through AI Studio, but it can pay for the same Gemini models called through Agent Platform. And it can't be used for partner models offered as a managed API, which rules out Claude and the other Model Garden APIs. If you're prototyping on credit, call Gemini through Agent Platform rather than an AI Studio key.
How far does the standing allowance go? Using the same assumptions as the worked examples (1 vCPU and 2 GiB, busy for about a minute per conversation), 50 free vCPU-hours and 100 free GiB-hours each cover about 3,000 conversations a month. Past that you pay the rates above, and model tokens are never free on Agent Platform. The small support bot above handles 9,000 conversations for about $91.
With four meters running at once, $300 goes fast during real development. The free tier is built for prototyping, not production. For platforms with a genuine free plan and no usage-based surprises, our free AI agent builder guide covers the options.
Other Vertex AI pricing (quick reference)
This page goes deep on agent pricing. Here's a one-line summary of the rest of the Vertex AI price list, checked September 22, 2026:
| Service | How it's billed | Starting price | Google's page |
|---|---|---|---|
| Custom training | Per node-hour for the machine type plus any GPU, plus a management fee | Varies by machine, GPU and region | Agent Platform pricing |
| Online and batch prediction | Per node-hour while a model is deployed. No scale-to-zero, so idle endpoints bill | Varies by machine and region | Agent Platform pricing |
| AutoML (image) | Per node-hour for training and for deployed prediction | $3.465 / node-hour to train, $1.375 / node-hour to serve (classification) | Agent Platform pricing |
| Vector Search | Node-hours for the deployed index, plus data processed | $3.00 per GiB to build or update an index | Agent Platform pricing |
| Agent Retrieval | Capacity units, plus payload storage and operations | $0.065 / hour per Performance-Optimized CU, $0.06 per 100,000 reads | Agent Platform pricing |
| RAG Engine | The engine is free. You pay for embedding and generation tokens, plus the Spanner database behind it in Spanner mode (or Vector Search storage in the preview Serverless mode) | Spanner pricing (Basic tier is 100 processing units) | RAG Engine billing |
| Model Garden | Per token for managed models, per node-hour for self-deployed ones | $0.09 / $0.36 per 1M tokens (gpt-oss-120b) | Generative AI pricing |
| Agent Search | Per 1,000 queries | $1.50 Standard, $4.00 Enterprise | Agent Search pricing |
| CX Agent Studio | Per chat or voice session | $0.50 per session | CX Agent Studio pricing |
| Gen AI Evals | Model-based metrics bill as autorater model tokens | The judge model's token rate | Agent Platform pricing |
Two of these catch people out. CX Agent Studio (Google's customer-service agent builder) starts a new billable chat session every 50 turns or after 30 minutes of inactivity. Voice sessions switch to $0.0025 per second after five minutes. RAG Engine looks free, but in the default Spanner mode the managed Spanner instance behind it bills every hour it exists, whether or not anyone queries it. The newer Serverless mode (preview, us-central1 only) drops the Spanner instance and has no additional charge, but you still pay for models, reranking and vector storage, and an old Spanner-mode instance keeps billing until you remove it.
Vertex AI vs AWS AgentCore pricing
AWS Bedrock AgentCore, Amazon's equivalent, prices its runtime slightly higher:
| Meter | Agent Platform (Vertex AI) | AWS AgentCore |
|---|---|---|
| Compute | $0.085 per vCPU-hour (50 free / month) | $0.0895 per vCPU-hour |
| Memory | $0.009 per GiB-hour (100 free / month) | $0.00945 per GB-hour |
| Idle time between turns | Not billed | Not billed (active consumption only) |
On runtime alone the two are within about 5% of each other. Pick the one where your data already lives. The full AgentCore meter list, including memory events, gateway and web search, is in our AWS AgentCore pricing breakdown. For token costs across every provider, see the LLM pricing guide.
Vertex AI vs Microsoft Copilot Studio pricing
Microsoft's agent builder bills on a completely different unit, so the two don't compare line by line:
| Google Agent Platform (Vertex AI) | Microsoft Copilot Studio | |
|---|---|---|
| Billing unit | vCPU-hours, GiB-hours, search queries and model tokens | Copilot Credits, used each time an agent responds or takes an action |
| Entry price | No monthly fee, pay per use | $200 per month for a 25,000-credit pack, or a pay-as-you-go meter |
| Model tokens | Billed separately per model | Covered by the credits the agent consumes |
| Free usage | 50 vCPU-hours, 100 GiB-hours and 1 GiB-month every month, plus $300 credit for new customers | Free trial, plus a $200 Azure credit for 30 days on a new Azure account |
| Commitment discount | 10% (1-year) or 20% (3-year) savings plans | Up to 20% on prepaid Copilot Credit Commit Units |
The short version: Google's bill tracks compute, retrieval and tokens, so cost scales with how heavy each conversation is. Microsoft's bill tracks agent responses and actions, which makes a busy, simple bot easier to forecast. Microsoft 365 Copilot ($30 per user per month, paid yearly) also lets you build internal agents inside Microsoft 365, but publishing an agent to outside channels needs a standalone Copilot Studio plan. Prices are from Microsoft's Copilot Studio pricing page, checked September 22, 2026.
What is Vertex AI Agent Builder?
Vertex AI Agent Builder was Google Cloud's platform for building, deploying and governing AI agents. Since Cloud Next 2026 it has been part of the Gemini Enterprise Agent Platform. Everything priced on this page is what it runs on: Agent Studio and the ADK to build, Agent Runtime to host, Agent Search to ground and Model Garden for models.
If you're weighing it against other platforms, see our BetterClaw vs Vertex AI head-to-head or the eight Vertex AI Agent Builder alternatives we compared. For the wider market, see other AI agent platforms compared.
Frequently Asked Questions
How much does Vertex AI cost?
There is no monthly fee. Vertex AI, now Gemini Enterprise Agent Platform, bills across four meters:
- Agent Runtime at $0.085 per vCPU-hour and $0.009 per GiB-hour, after 50 free vCPU-hours a month
- Sessions and Memory Bank storage at $0.30 per GiB-month
- Agent Search at $1.50 to $4.00 per 1,000 queries
- Model tokens
A small support bot on Gemini 3.8 Flash costs about $90 a month at the introductory token price. A heavy Gemini 3.1 Pro agent can pass $13,000.
How much does Gemini 3.8 Flash cost on Vertex AI?
$0.75 per million input tokens and $3.75 per million output tokens on the global endpoint, through December 31, 2026. From January 1, 2027 it's $1.50 and $7.50. Cached input is a tenth of the input price, Flex and Batch halve the rate, regional endpoints add 10%, and Priority costs 1.8x. The introductory price is delivered as a 50% billing credit rather than a lower list price.
Is Vertex AI free to use?
Partly. Every account gets 50 free vCPU-hours of Agent Compute, 100 GiB-hours of Agent Memory and 1 GiB-month of storage each month. Agent Search includes 10,000 free queries a month, and Google Search grounding includes 5,000 free queries a month. New customers also get $300 in credit, which covers Gemini through Agent Platform but not through AI Studio, and not partner models such as Claude.
What is Gemini Enterprise Agent Platform?
It's the new name for Vertex AI, adopted around Google Cloud Next 2026 in late April. Vertex AI Agent Builder, Agent Engine and Vertex AI Search became Agent Platform, Agent Runtime and Agent Search. The services and prices carried over, and existing customers didn't need to migrate. Agentspace was folded in at the same time.
How is Vertex AI Agent Engine priced?
Agent Engine, now Agent Runtime, bills $0.085 per vCPU-hour of compute and $0.009 per GiB-hour of memory, rounded to the second. The first 50 vCPU-hours and 100 GiB-hours each month are free, and idle time between turns isn't billed. The 1-year and 3-year Flexible Savings Plans lower compute to $0.0765 and $0.068 per vCPU-hour, and memory to $0.0081 and $0.0072 per GiB-hour.
How much does Vertex AI Search cost?
Vertex AI Search, now Agent Search, costs $1.50 per 1,000 queries on Standard Edition and $4.00 per 1,000 on Enterprise Edition, which includes core Generative Answers. Advanced Generative Answers add $4.00 per 1,000 queries on either edition. The first 10,000 queries each month are free. Grounding with Google Search is separate: $14 per 1,000 after 5,000 free a month.
How much does Model Garden cost on Vertex AI?
Managed models in Model Garden bill per token at each provider's rate. On the global endpoint, Claude Sonnet 5 is $2 in and $10 out per million tokens, Claude Opus 5.5 is $4 and $20 (Opus 5 is $5 and $25), and open models run from $0.09 and $0.36 (gpt-oss-120b) to $0.72 and $0.72 (Llama 3.3 70B). Open models you deploy to your own endpoint bill per node-hour instead, including idle time.
How much does CX Agent Studio cost?
$0.50 per chat or voice session. A chat session rolls over into a new billable session every 50 turns or after 30 minutes of inactivity. A voice session switches to $0.0025 per second after five minutes.
Is Vertex AI cheaper than Microsoft Copilot Studio?
It depends on the workload. Vertex AI bills compute, search and model tokens separately, so a light agent on a cheap model can cost very little. Copilot Studio bills Copilot Credits, $200 a month for a 25,000-credit pack, with model usage included. That makes it simpler to forecast but not always cheaper.
How does Vertex AI Agent Builder compare to BetterClaw?
Vertex AI is built for GCP-native enterprises that need RAG grounding on Google Cloud data with enterprise governance. BetterClaw is built for teams who want agents running in 60 seconds without cloud infrastructure, with a free plan, $49 per month Pro ($39 billed annually), 200+ verified skills, and 28+ model providers. The trade-off: BetterClaw does not offer GCP-native grounding.




