GPU prices for AI, ranked by $/GB VRAM.
For local LLMs, VRAM decides what you can run at all. We rank GPUs by dollars per GB of VRAM so you can see real value through the shortage — lower is better. Live Amazon prices, refreshed daily.
GPUs by $/GB VRAM
Lower $/GB VRAM is better. Tap Check → for live listings.
| Product | VRAM | Lowest $/GB VRAM | Amazon |
|---|---|---|---|
| Intel Arc B580 12GB · Battlemage | 12 | $25.83 | |
| NVIDIA RTX 5060 Ti 16GB 16GB GDDR7 · Blackwell | 16 | $36.87 | |
| AMD RX 9070 XT 16GB GDDR6 · RDNA4 | 16 | $43.12 | |
| NVIDIA RTX 5060 8GB GDDR7 · Blackwell | 8 | $43.75 | |
| NVIDIA RTX 5070 12GB GDDR7 · Blackwell | 12 | $53.00 | |
| AMD RX 7900 XTX 24GB GDDR6 · RDNA3 | 24 | $58.29 | |
| NVIDIA RTX 5070 Ti 16GB GDDR7 · Blackwell | 16 | $60.56 | |
| NVIDIA RTX 4070 12GB GDDR6X · Ada | 12 | $60.75 | |
| NVIDIA RTX 4070 Super 12GB GDDR6X · Ada | 12 | $69.92 | |
| NVIDIA RTX 5080 16GB GDDR7 · Blackwell | 16 | $82.23 | |
| NVIDIA RTX 4070 Ti Super 16GB GDDR6X · Ada | 16 | $84.37 | |
| NVIDIA RTX 4080 Super 16GB GDDR6X · Ada | 16 | $87.50 | |
| NVIDIA RTX PRO 6000 96GB GDDR7 ECC · Blackwell workstation | 96 | $128.96 | |
| NVIDIA RTX 5090 32GB GDDR7 · Blackwell | 32 | $132.78 | |
| NVIDIA RTX 4090 24GB GDDR6X · Ada | 24 | $141.66 | |
| NVIDIA RTX 4060 Ti 16GB 16GB GDDR6 · Ada | 16 | — | see live price Check → |
Snapshot aggregated Jul 20, 2026, 6:08 AM. Prices older than 24 hours are hidden — tap Check → for the live price.
As an Amazon Associate I earn from qualifying purchases. The Check → buttons are paid affiliate links, so we may earn a commission on a qualifying purchase at no extra cost to you. Prices and availability are accurate only as of the date/time shown and are subject to change; the price and availability on the retailer's own site at the time of purchase is the one that applies.
Flagship VRAM (RTX 4090 / 5090) is scalped well above MSRP right now, so it ranks poorly on $/GB — the value is in mid-range cards. AMD (RX 9070 XT, 7900 XTX) and Intel (Arc B580) often win on raw $/GB VRAM; just note NVIDIA/CUDA still leads on AI software maturity, while AMD (ROCm) and Intel are usable but need more setup. Prices are the lowest current new listing per model.
How much VRAM do you need?
VRAM is the gate: if the model fits, it runs fast; if it doesn't, it spills to system RAM and crawls. Rough fit at 4-bit (Q4) quantization:
| VRAM | Runs (Q4) | Typical cards |
|---|---|---|
| 8 GB | ~7–8B models | RTX 4060 Ti 8GB |
| 12 GB | ~12–14B models | RTX 5070 |
| 16 GB | ~14B models | RTX 5080, 5070 Ti, 5060 Ti 16GB |
| 24 GB | ~32B models | RTX 4090, used 3090 |
| 32 GB | 32B with headroom / faster | RTX 5090 |
| 48 GB (2×24) | ~70B models | Dual 3090 / 4090 |
| 128 GB unified | up to ~200B (slower) | DGX Spark |
FP16 needs roughly 4× the VRAM of Q4. Numbers are approximate and depend on context length and quantization.
💸 Best value: used RTX 3090 (24GB)
The 24GB RTX 3090 remains the community value king for local AI — same VRAM as a 4090 at roughly half the used price. We don't price-track the used market, but if you're chasing $/GB VRAM, this is the card to watch on the second-hand market.
🧠 DGX Spark — 128GB unified memory
A desktop AI appliance (128GB unified LPDDR5x · GB10 Grace Blackwell) that runs very large models locally — up to ~200B parameters. The catch: unified LPDDR5x is high-capacity but far lower bandwidth than GPU VRAM, so it's about fitting big models, not peak speed. MSRP is ~$4,699 (raised from $3,999 amid the memory shortage).
Check DGX Spark on Amazon →Why $/GB VRAM?
A 32GB card isn't twice as useful as a 16GB one if it costs three times as much. For local AI, the first question is always "does the model fit in VRAM?" — so the honest value metric is dollars per GB of VRAM. We show the lowest current new listing per model and link straight to Amazon. Prices older than 24 hours are hidden rather than shown as current.
See also: Best GPU for local LLMs guide · DDR5 by $/GB · SSDs by $/TB · all guides.
GPU for AI FAQ
What is the best GPU for running local LLMs?
VRAM decides what you can run; everything else decides how fast. For most people the value sits in 16GB cards (RTX 5080 / 5070 Ti / 5060 Ti 16GB) and the 24GB RTX 4090 or used RTX 3090. The RTX 5090 (32GB) is the fastest single card but is heavily price-inflated right now. Rank by $/GB VRAM and buy the most VRAM you can afford at a sane price.
How much VRAM do I need for a given model size?
Roughly, at 4-bit (Q4) quantization: 8GB runs ~7-8B models, 16GB ~14B, 24GB ~32B, and 48GB (two 24GB cards) ~70B. FP16 needs about 4× that. If a model does not fit in VRAM it spills to system RAM and slows 5-20×, so fitting in VRAM matters more than raw speed.
Is a used RTX 3090 still worth it for AI?
Yes — the 24GB RTX 3090 is the long-standing value pick for local AI, offering the same VRAM as a 4090 at roughly half the price on the used market. It is slower, but for development and inference where the model just needs to fit, it is hard to beat on $/GB VRAM.
What is the NVIDIA DGX Spark?
A desktop AI appliance with 128GB of unified LPDDR5x memory (GB10 Grace Blackwell), able to run very large models (up to ~200B parameters) locally. Its memory is high-capacity but far lower bandwidth than GPU VRAM, so it is about fitting big models, not maximum speed.