Fit an LLM on one A100 80GB. Move the sliders, watch the bar, see whether your hypothetical model fits and how much room is left for KV cache. Starts loaded with meta-llama/Llama-3.3-70B-Instruct on TP=2 as the reference case.
Model architecture
Parameters (billions) 70.55
Layers 80
KV heads 8
Head dim 128
Precision (dtype)
Deployment
max_model_len (tokens) 32768
Concurrent requests (target) 1
Tensor parallelism (per-GPU sharding)
VRAM composition (per GPU)
A100 80GB · 1 GPU74.5 GiB usable
↓ zoom into the KV cache budget
KV cache budget · what live requests fillGiB budget