Run big LLMs on cheap, mismatched GPUs

Updated June 2026

In a GPU shortage, the smart way to a 70B is not one $12,000 card — it is several cheap cards pooled together. Here is how to combine mismatched GPUs into a single endpoint, and which cards give you the most VRAM per dollar to do it.

Pool cheap VRAM instead of buying one big card

A 70B model needs roughly 48GB of VRAM at Q4 — normally that means a single workstation card costing five figures, or two matched 24GB cards. But there is a cheaper way: pool the VRAM of several mismatched cards you already have (or can buy cheaply), across one or more machines, so they act as one big pool.

The open-source tool for this is Tightwad (MIT-licensed, free). Its pitch: "pool your mismatched CUDA + ROCm + Metal cards into one OpenAI-compatible endpoint, so a model that fits on no single machine runs across all of them." It uses speculative decoding for speed (~1.86× in its own benchmarks), supports Mixture-of-Experts models, and runs entirely on your hardware — no cloud bill.

Tightwad — pool mismatched GPUs into one local LLM endpoint. Open-source, runs in your homelab.

Why this changes how you buy GPUs

Pooling flips the goal from "one card big enough" to "the most cheap VRAM you can combine." That favors high-value cards on $/GB VRAM — even across brands, since Tightwad mixes NVIDIA, AMD and Apple. Two cheaper 16–24GB cards can out-VRAM a single flagship for less money. Rank candidates by $/GB on the GPU tracker and pair with the NVIDIA vs AMD vs Intel guide for the software trade-offs.

The build, end to end

Live GPU prices ($/GB VRAM)

Product VRAM Lowest $/GB VRAM Amazon
Intel Arc B580 12GB · Battlemage 12 $25.83
$309.99
seen 10m ago
Check →
NVIDIA RTX 5060 Ti 16GB 16GB GDDR7 · Blackwell 16 $38.12
$609.99
seen 10m ago
Check →
AMD RX 9070 XT 16GB GDDR6 · RDNA4 16 $43.12
$689.99
seen 10m ago
Check →
NVIDIA RTX 5060 8GB GDDR7 · Blackwell 8 $43.75
$349.99
seen 10m ago
Check →
NVIDIA RTX 5070 12GB GDDR7 · Blackwell 12 $53.00
$635.99
seen 10m ago
Check →
AMD RX 7900 XTX 24GB GDDR6 · RDNA3 24 $58.29
$1399.00
seen 10m ago
Check →
NVIDIA RTX 4060 Ti 16GB 16GB GDDR6 · Ada 16 $59.37
$949.99
seen 10m ago
Check →
NVIDIA RTX 5070 Ti 16GB GDDR7 · Blackwell 16 $60.56
$969.00
seen 10m ago
Check →
NVIDIA RTX 4070 12GB GDDR6X · Ada 12 $60.75
$729.00
seen 10m ago
Check →
NVIDIA RTX 4070 Super 12GB GDDR6X · Ada 12 $69.92
$838.99
seen 10m ago
Check →
NVIDIA RTX 5080 16GB GDDR7 · Blackwell 16 $78.12
$1249.99
seen 10m ago
Check →
NVIDIA RTX 4070 Ti Super 16GB GDDR6X · Ada 16 $84.37
$1349.99
seen 10m ago
Check →
NVIDIA RTX 4080 Super 16GB GDDR6X · Ada 16 $99.94
$1599.00
seen 10m ago
Check →
NVIDIA RTX PRO 6000 96GB GDDR7 ECC · Blackwell workstation 96 $128.96
$12379.99
seen 10m ago
Check →
NVIDIA RTX 5090 32GB GDDR7 · Blackwell 32 $132.81
$4249.95
seen 10m ago
Check →
NVIDIA RTX 4090 24GB GDDR6X · Ada 24 $144.38
$3465.00
seen 10m ago
Check →

Snapshot aggregated Jul 21, 2026, 6:08 AM. Prices older than 24 hours are hidden — tap Check → for the live price.

Pooling GPUs for local AI — FAQ

Can you run a 70B model on cheap GPUs instead of one expensive card?

Yes — by pooling several cards across your machines so their VRAM combines. A 70B that needs ~48GB at Q4 can run across, say, a couple of 24GB cards (even mismatched brands) instead of buying a single 96GB workstation card. Free open-source tools like Tightwad pool mismatched GPUs into one endpoint to do exactly this.

What is Tightwad?

Tightwad is open-source (MIT) software that pools mismatched CUDA, ROCm and Apple Metal GPUs across machines into a single OpenAI-compatible endpoint, so a model that fits on no single machine runs across all of them. It uses speculative decoding for speed (~1.86x measured in its benchmarks) and runs entirely on your own hardware with no cloud cost.

Do the GPUs have to match?

No — that is the point. Tightwad is built for mismatched hardware: you can combine NVIDIA, AMD and Apple cards of different sizes over your network. That makes the cheapest-$/GB cards from any brand useful together, rather than forcing one matched set.

How does this change what GPU I should buy?

It shifts the goal from "one card big enough for the model" to "the most cheap VRAM you can pool." Instead of paying a flagship or workstation premium, you can buy several high-$/GB-value cards (see the live tracker) and pool them. Watch $/GB VRAM and buy on dips.