GPU VRAM vs models: what fits
Heuristic VRAM tables for 7B/32B/70B/MoE and GPUs from 4060 to H100/H200.

GPU VRAM vs models: what fits
Heuristic VRAM tables for 7B/32B/70B/MoE classes and common GPUs (4090, A100, H100, Apple Silicon).
Three rules
VRAM decides fit; FLOPs decide speed; quantization trades quality for memory.
Model classes
3B–8B Q4 on 8GB; 32B Q4 on 24GB; 70B Q4 often needs 48GB or multi-GPU; 100B+ prefers datacenter cards.
Practical picks
Personal coding: 24GB. Team private API: 48GB+. Burst training: rent H100 by the hour.
Always re-test
KV cache, concurrency, and vision resolution change the math—benchmark your prompts.
---
Disclaimer: GrokCode aggregates information only. Not purchase or legal advice. Verify hardware vendors, cloud ToS, and model licenses yourself.
适用于 GrokCode 倍率榜。信息仅供参考,不构成购买、投资或法律意见。