Updated September 2026. Every price below was re-checked against the provider's own page on 23 September 2026. This post now covers the RTX 5090, renting a Spark by the hour, and NVIDIA's upcoming RTX Spark.
A DGX Spark costs $4,699. An RTX 5090 lists at $1,999 and actually sells for more than twice that. A cloud API costs $0 upfront. Here is what each one really costs over 36 months of running AI agents, with output tokens counted properly, and why the honest answer is still "rent something first".
Local and cloud on one dashboard.
BetterClaw routes cloud APIs and local Ollama endpoints from a single agent config via BYOK, with zero inference markup. Free forever, not a trial. Start free No credit card. BYOK. No hardware to manage.
The NVIDIA DGX Spark is a GB10 Grace Blackwell superchip with 128 GB of LPDDR5x unified memory in a box the size of a hardback book. Linux only. The promise: run 200B+ parameter models locally without renting a single GPU hour. (If you are pricing the rest of NVIDIA's line too, start with all NVIDIA DGX alternatives.)
Loaded with GLM 5.2 at Q4 (753B MoE, 40B active), it runs. Inference is smooth. No per-token billing, no Ollama fetch failed debugging.
Then you do the maths, and the maths is where most of these comparisons quietly cheat. They price cloud APIs on input tokens only. Agents generate output too, and output costs four to five times more per token on every major provider. Fix that one assumption and the break-even moves by months.

The four options (specs and prices, September 2026)

DGX Spark
Price: $4,699 MSRP, raised from $3,999 in February 2026 when memory supply tightened. Street prices are now climbing above MSRP: NVIDIA's own marketplace has been showing the Founders Edition out of stock, and the cheapest unit actually in stock has been listing near $5,000 on Amazon through Micro Center. Budget $4,699 if you can wait, $5,000 if you cannot.
Specs: NVIDIA GB10 Grace Blackwell superchip. 128 GB LPDDR5x coherent unified memory shared between CPU and GPU. 273 GB/s memory bandwidth. 240 W power supply, 140 W chip TDP. DGX OS, which is Linux.
What it runs: GLM 5.2 at Q4 (753B MoE, 40B active). Qwen 3.8 27B at full precision. Gemma 4 12B comfortably. Most open-weight models under 200B.
What it does not run: Full-precision 400B+ dense models. Multiple large models at once. Anything that needs bandwidth more than capacity.
RTX 4090 build
Price: the 4090 went out of production in early 2024. There are no new cards coming. Remaining stock and the used market put a 24 GB card between $2,500 and $3,700 depending on condition and warranty, well above its old $1,599 MSRP. Add roughly $800 for the rest of the machine and you are at about $3,300 to $4,500. We use $3,400 below.
Specs: 24 GB GDDR6X. 1,008 GB/s memory bandwidth. 450 W card TDP. Windows, Linux or a macOS eGPU setup.
What it runs: Qwen 3.8 27B at Q8. Gemma 4 12B at FP16. Up to roughly 27B dense or 70B MoE at Q4.
What it does not run: anything needing more than 24 GB without heavy quantization.
RTX 5090 build (the card people are actually cross-shopping)
Price: this is the number that has moved most. NVIDIA's MSRP is $1,999. Nobody is paying $1,999. Through September 2026 the tracked new-card price has run between about $4,200 and $7,400, with $6,795 the lowest in-stock listing on 22 September, driven by a GDDR7 shortage and by AI shops buying 5090s in bulk for inference servers. A full build is therefore either about $2,800 (MSRP, if you find one) or about $5,000 to $7,600 (reality). We show both.
Specs: 32 GB GDDR7 on a 512-bit bus, roughly 1.8 TB/s of memory bandwidth, which is about 6.5 times the Spark. 21,760 CUDA cores. 575 W card TDP. Windows or Linux.
What it runs: everything the 4090 runs, faster, plus a bit more headroom at 32 GB. On LMSYS's published gpt-oss-20b (MXFP4) benchmark a single 5090 hit 8,519 tokens/sec prefill and 205 tokens/sec decode, against the Spark's 2,053 and 49.7. That is roughly 4x on generation, on a model both machines can load, because at that size bandwidth decides the outcome and capacity is irrelevant.
What it does not run: the 128 GB models. A 753B MoE at Q4 does not fit in 32 GB at any quantization worth using. This is the whole trade: the 5090 is faster at what fits, the Spark fits more.
Cloud API (BYOK)
Price: $0 upfront, then per token. Current list prices, input and output per million tokens, checked 23 September 2026:
| Model | Input | Output |
|---|---|---|
| DeepSeek Flash (V4.1) | $0.30 | $1.20 |
| MiniMax M3 (up to 512K context) | $0.30 | $1.20 |
| Gemini 3.8 Flash (introductory, through 31 Dec 2026) | $0.75 | $3.75 |
| GLM 5.2 | $1.40 | $4.40 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| GPT-6 Sol | $2.00 | $10.00 |
| Claude Opus 5 | $5.00 | $25.00 |
What it runs: every model, including the proprietary ones. No hardware limits.
What it does not run: nothing. If it has an API, you can point an agent at it.
The 36-month cost table (this is the one that matters)
Assume your agent processes 500 tasks a day at 5K tokens a task: 2.5M tokens a day, 75M tokens a month.
The assumption that fixes the old version of this table: 80% input, 20% output. That is typical for agent work, where long tool outputs and file contents go in and short decisions come out. Blended price per million = (0.8 x input) + (0.2 x output). For Sonnet 5 that is (0.8 x $2) + (0.2 x $10) = $3.60 per million, so 75M tokens costs $270 a month, not the $225 you get by pricing input only. If your split is closer to 50/50, raise every cloud number below by roughly 40%.
Power: 8 hours a day of real inference, 30 days a month, at $0.15/kWh, measured on whole-system draw. The Spark is the efficient one here. A 5090 under sustained load draws more than twice what the whole Spark does.
| DGX Spark | RTX 4090 build | RTX 5090 (MSRP) | RTX 5090 (street) | Flash / M3 | Gemini 3.8 Flash | GLM 5.2 | Sonnet 5 | Opus 5 | |
|---|---|---|---|---|---|---|---|---|---|
| Upfront | $4,699 | $3,400 | $2,800 | $5,000 | $0 | $0 | $0 | $0 | $0 |
| Blended $/M | n/a | n/a | n/a | n/a | $0.48 | $1.35 | $2.00 | $3.60 | $9.00 |
| Monthly tokens | $0 | $0 | $0 | $0 | $36 | $101 | $150 | $270 | $675 |
| Monthly power | $9 | $22 | $26 | $26 | $0 | $0 | $0 | $0 | $0 |
| 12 months | $4,807 | $3,664 | $3,112 | $5,312 | $432 | $1,215 | $1,800 | $3,240 | $8,100 |
| 24 months | $4,915 | $3,928 | $3,424 | $5,624 | $864 | $2,430 | $3,600 | $6,480 | $16,200 |
| 36 months | $5,023 | $4,192 | $3,736 | $5,936 | $1,296 | $3,645 | $5,400 | $9,720 | $24,300 |

Break-even in one formula
Months to break even = hardware price divided by (monthly cloud bill minus monthly power cost)
At 75M tokens a month:
- DGX Spark vs Sonnet 5: $4,699 / ($270 - $9) = about 18 months
- RTX 5090 build at MSRP vs Sonnet 5: $2,800 / ($270 - $26) = about 12 months
- RTX 5090 build at September street price vs Sonnet 5: $5,000 / ($270 - $26) = about 21 months
- DGX Spark vs Opus 5: $4,699 / ($675 - $9) = about 7 months
- DGX Spark vs GLM 5.2: $4,699 / ($150 - $9) = about 33 months
- Either machine vs DeepSeek Flash or MiniMax M3: never. $4,699 / ($36 - $9) is 174 months, roughly 14 years. The hardware will be landfill first. On a budget model the cloud bill is barely bigger than the power bill.

Two things that formula hides, and you should hold in your head. First, it assumes a local open-weight model does the job as well as the cloud model you are replacing, which is the entire bet. Second, hardware has residual value and rental spend does not, so if you resell the Spark after two years at half price the real payback is faster than 18 months.
One more wrinkle if you are estimating against Claude: the tokenizer introduced with Opus 4.7 counts up to about 35% more tokens for the same text than earlier Claude models. If your 75M figure came from an older Claude bill, re-baseline it with a token count before you trust the break-even.
Should you rent a DGX Spark instead?
If you are not certain you will use it every day, do not buy one yet.
Hosted DGX Sparks rent for about $0.60 to $0.75 an hour with root SSH access. Enverge lists a single 128 GB node at $0.75/hr and a two-node NVLinked setup with 256 GB unified memory at $1.65/hr; there are cheaper community hosts on NVIDIA's own developer forum renting spare units nearer $0.60.
At $0.75 an hour, buying only wins after roughly 6,300 hours of use, and nearer 6,600 once you add the owner's power bill. That is more than two years of eight-hour days, every working day, without a gap. For bursty or project-based work, renting is not close.

One thing that confuses people: NVIDIA DGX Cloud is not a DGX Spark. DGX Cloud is NVIDIA's enterprise cloud for multi-GPU training clusters, sold per node per month and aimed at teams with a procurement department. If you want Spark-class hardware by the hour, you want a host that rents the actual Spark unit.
What about RTX Spark?
NVIDIA's RTX Spark is the consumer version of the same idea: a Grace Blackwell-class Arm chip with up to 128 GB of unified memory, running Windows, in laptops and mini desktops from ASUS, Dell, HP, Lenovo, MSI and Microsoft Surface, with Acer and Gigabyte following later. NVIDIA has confirmed an October 2026 launch window for the N1X. No vendor has published a price. Morgan Stanley's estimate puts 128 GB N1X machines near $2,899 and the smaller N1 models near $1,799, but that is an analyst note, not a price list.
Why it matters here: "the 4090 is your only local option if you need Windows" stops being true the week RTX Spark ships. If Windows matters to you and you can wait a few months, wait. We will put real prices in this post the week they are announced.
When DGX Spark actually makes sense
Replacing a premium model at high volume. Against Sonnet 5 at 75M tokens a month it pays back in about 18 months. Against Opus 5 it pays back in seven. Against anything cheap, never. That is the whole rule.
Models that need the capacity. 128 GB runs models a 5090 cannot load at all. If your workload genuinely needs a 200B+ open-weight model rather than a 27B one, the Spark is not competing with the 5090, it is the only one of the two in the race. Just know that the 273 GB/s memory bandwidth is what you pay for that capacity, and on big models it is the bottleneck.
Data sovereignty. Your data never leaves the building. For healthcare, legal, financial or government workloads, the premium buys privacy, not performance. If you are running OpenClaw on NVIDIA hardware, our NemoClaw vs OpenClaw breakdown explains what NVIDIA's security wrapper actually changes.
Air-gapped environments. No internet required. Cloud APIs are physically impossible in some buildings, and then price stops being the question.
Experimentation. Zero marginal cost means thousands of test prompts without watching a billing dashboard. For a team iterating on prompts, a fixed cost is easier to budget than a variable one.
When a GPU build is the better choice

Speed on models that fit in 32 GB. Four times the decode throughput on gpt-oss-20b is not a rounding error. If your agents run 27B-and-under models, a 5090 is the faster machine by a wide margin, and at MSRP it would also be the cheaper one.
You need more than inference. A GPU trains, fine-tunes, generates images and video, and plays games. The Spark is an inference box.
Windows, for now. Until RTX Spark ships in October, a GeForce card is still the only local option on Windows. After that, this reason expires.
The catch, in September 2026: neither card is cheap any more. The 4090 is out of production and the 5090 sells at two to three times MSRP. At $5,000 for a 5090 build, you are paying more than a Spark for less memory. Check the actual street price on the day you buy, then redo the formula above. If the shortage eases and 5090s land near $2,000 again, the 5090 build becomes the best-value option on this page by a distance.
When cloud API wins (and it is most of the time)
For most agent builders, cloud is still the right answer.
$0 upfront. No purchase, no depreciation, no maintenance, and no exposure to a memory shortage that has doubled GPU prices in six months.
Access to proprietary models. Claude Sonnet 5 and Opus 5, the GPT-6 line, Gemini. These do not run locally at any price.
Scales to zero. Quiet month? Pay nothing. Hardware costs the same whether you run one task or ten thousand.
The budget models are absurdly cheap. DeepSeek Flash and MiniMax M3 both blend to about $0.48 a million at an 80/20 split. That is $36 a month for the workload in this post, which no hardware purchase will ever undercut. DeepSeek also halves its rates off peak, which pushes it lower still.
If you are building agents on cloud APIs, BetterClaw is a no-code AI agent platform supporting 28+ providers via BYOK with zero inference markup, plus 200+ verified skills. Free plan with 1 agent and 100 credits a month. $49/month on Pro for 5 agents and 12,000 credits. Per-agent cost caps, so a runaway loop cannot run up a bill.
The hybrid setup (what production teams actually run)

The teams shipping the best agents in 2026 do not pick one path.
Development and testing: local hardware, whichever you have. Zero marginal cost for iterating on prompts, testing tool configs and debugging agent behaviour.
Privacy-sensitive production: local inference through Ollama. Customer PII, medical records, financial data. It never leaves the building.
Everything else: cloud via BYOK, routed per task. Classification to DeepSeek Flash, coding to GLM 5.2, the hard reasoning calls to Sonnet 5. Best model for each job, not one model for every job.
Monthly cost of the hybrid: $0 for dev and test, $9 to $26 in power for the privacy work, $50 to $300 for production API. Call it $60 to $330 a month plus whatever the hardware cost you. Compare that to $270 to $675 a month all-cloud on a premium model, or all-local at $0 a month but $2,800 to $5,000 upfront and no access to Sonnet or Opus at all.
The question is not "Spark or cloud?" It is "which tasks need local, and which need cloud?" The answer is nearly always both, and model routing handles the split for you.
The honest recommendation
Run on cloud APIs first. Measure your actual token volume and your actual GPU-hours for a month. Then:
- Under about 20M tokens a month: stay on cloud. Nothing pays back.
- 20M to 75M on a budget model: stay on cloud. Nothing pays back.
- 75M+ on a premium model, and a local open-weight model genuinely does the job: rent a Spark by the hour for a month to prove the quality holds, then buy if it does.
- You need 128 GB: the Spark, or rent one.
- You need speed on small models and can find a 5090 near MSRP: the 5090 build.
- You need Windows: wait for October.
Buying first and measuring second is how $4,699 machines end up idle.
Give BetterClaw a look if you want cloud APIs and local model endpoints on one dashboard. Free plan with 1 agent and 100 credits a month. $49/month for Pro. BYOK with zero markup. Connect your Ollama instance, your cloud keys, or both, and we handle the routing.
Frequently Asked Questions
Is DGX Spark worth $4,699 for running AI agents?
It depends entirely on which cloud model you are replacing. At 500 tasks a day (75M tokens a month, 80% input and 20% output), a DGX Spark breaks even against Claude Sonnet 5 in about 18 months, and against Opus 5 in about 7. Against DeepSeek Flash or MiniMax M3 at roughly $36 a month, it never breaks even: the maths comes out at 174 months, which is longer than the hardware will last. Buy it to replace a premium model at high volume, for data sovereignty, or for an air-gapped environment. For most agent builders, cloud at $0 upfront is still cheaper.
Should I buy an RTX 4090 or a DGX Spark?
In September 2026, neither, without checking prices first. The 4090 has been out of production since early 2024 and now trades at $2,500 to $3,700 for a used or old-stock 24 GB card, so the real comparison is now the RTX 5090: 32 GB of GDDR7, about 1.8 TB/s of bandwidth, $1,999 MSRP but $4,200 to $7,400 on the street during the current GDDR7 shortage. Choose the 5090 for speed on models that fit in 32 GB, where it generates tokens roughly 4x faster than the Spark. Choose the Spark for models that need 128 GB, which the 5090 cannot load at all. For local model setup, see our Qwen 3.8 on Ollama guide.
Can I rent a DGX Spark?
Yes. Hosted DGX Sparks rent for about $0.60 to $0.75 an hour with root SSH access, and about $1.65 an hour for a two-node NVLinked setup with 256 GB of unified memory. At $0.75 an hour, buying a $4,699 unit only beats renting after roughly 6,300 hours of use, which is over two years of eight-hour working days. Note that NVIDIA DGX Cloud is a different product: it is an enterprise multi-GPU training cloud sold per node per month, not Spark hardware by the hour.
Does DGX Spark have a license cost?
No, not to use the machine. The hardware price includes DGX OS and the preinstalled NVIDIA AI software stack, so you can buy a Spark, plug it in and run models with nothing further to pay. The one thing sold separately is NVIDIA AI Enterprise for DGX Spark, the enterprise support and production software entitlement. NVIDIA's own DGX Spark documentation states that this entitlement exists only if you purchased it, requested an evaluation, or hold an entitlement certificate that names it, and a 90-day trial is available. Individual developers do not need it.
DGX Spark vs Claude: which is cheaper?
At 75M tokens a month with an 80/20 input-output split, Claude Sonnet 5 costs about $270 a month and Opus 5 about $675. A DGX Spark costs $4,699 up front plus about $9 a month in power, so it pays back against Sonnet 5 in roughly 18 months and against Opus 5 in roughly 7, provided a local open-weight model does the job as well. Below about 20M tokens a month, Claude is cheaper for years. Note that Claude models from Opus 4.7 onward use a tokenizer that counts up to about 35% more tokens for the same text, so re-baseline any estimate you carried over from an older Claude bill.
What is RTX Spark and should I wait for it?
RTX Spark is the consumer, Windows-based version of the Spark idea: a Grace Blackwell-class Arm chip with up to 128 GB of unified memory, shipping in mini desktops and laptops from ASUS, Dell, HP, Lenovo, MSI and Microsoft Surface. NVIDIA has confirmed an October 2026 launch window for the N1X chip but no vendor has announced a price; Morgan Stanley estimates around $2,899 for 128 GB N1X machines and around $1,799 for the smaller N1. If you need Windows and can wait a few months, wait, because RTX Spark removes the only remaining reason to buy a GeForce card for local inference on Windows.
How does cloud API compare to local hardware for agent costs?
At 75M tokens a month with an 80/20 split: DeepSeek Flash or MiniMax M3 cost about $432 a year, Gemini 3.8 Flash about $1,215, GLM 5.2 about $1,800, Claude Sonnet 5 about $3,240 and Opus 5 about $8,100. An RTX 5090 build costs $2,800 to $5,000 in year one depending on what you pay for the card, and a DGX Spark costs about $4,807. Cloud is cheaper than hardware for the first one to three years on every budget model. Hardware only wins on premium models, at high volume, sustained over two years or more.
What is the best setup for production AI agents in 2026?
A hybrid. Local hardware for development, testing and privacy-sensitive tasks, cloud APIs via BYOK for everything else, routed per task to the cheapest model that can do the job. On BetterClaw ($0 free with 1 agent and 100 credits, $49/month Pro with 5 agents and 12,000 credits), you connect local Ollama endpoints and cloud provider keys to the same dashboard and route between them automatically. Expect $60 to $330 a month for a mixed workload, plus whatever hardware you already own.
Do not buy hardware to find out.
Start on cloud via BYOK with zero markup, add a local Ollama endpoint when you need one, all from one BetterClaw dashboard. Free forever, not a trial. Start free



