Will it fit? Napkin math for LLM memory

Fit an LLM on one A100 80GB. Move the sliders, watch the bar, see whether your hypothetical model fits and how much room is left for KV cache. Starts loaded with meta-llama/Llama-3.3-70B-Instruct on TP=2 as the reference case.

Model architecture

Parameters (billions) 70.55
Layers 80
KV heads 8
Head dim 128

Precision (dtype)

Deployment

max_model_len (tokens) 32768
Concurrent requests (target) 1

Tensor parallelism (per-GPU sharding)

VRAM composition (per GPU)

A100 80GB · 1 GPU 74.5 GiB usable
zoom into the KV cache budget
KV cache budget · what live requests fill GiB budget