Self-hosting

GPU memory planner

Beam is a Mixture-of-Experts model: only 23 billion parameters work on each token, but all 501 billion must sit in memory. Pick a precision, context and hardware to see what it takes.

beam-501b-a23b
Weight precision
KV cache precision
Attention assumptions

Reflection describes interleaved local and global attention but has not published config.json yet. These placeholders drive the KV cache estimate; change them when the real values are out.

Total memoryKV cache per token: 26.0 KB
Weights
501 GB
KV cache
16.6 GB
Overhead
51.8 GB
Total memory
569 GB
WeightsKV cacheOverhead
NVIDIA H200 141GB needed
5 minimum, 8 in practice
Fits on one 8-GPU server

GPUs needed for the weights alone

Minimum devices per precision before KV cache, with 92% of each device's memory usable.

Accelerator BF16 / FP16
1,002 GB
FP8
501 GB
GGUF Q4_K_M
304 GB
INT4 / FP4
251 GB
NVIDIA H100 80GB × 14× 7× 5× 4
NVIDIA H200 141GB × 8× 4× 3× 2
NVIDIA B200 180GB × 7× 4× 2× 2
NVIDIA B300 288GB × 4× 2× 2× 1
NVIDIA A100 80GB × 14× 7× 5× 4
NVIDIA RTX PRO 6000 96GB × 12× 6× 4× 3
NVIDIA L40S 48GB × 23× 12× 7× 6
NVIDIA RTX 5090 32GB × 35× 18× 11× 9
NVIDIA RTX 4090 24GB × 46× 23× 14× 12
AMD Instinct MI300X 192GB × 6× 3× 2× 2
AMD Instinct MI325X 256GB × 5× 3× 2× 2
AMD Instinct MI355X 288GB × 4× 2× 2× 1
Apple M3 Ultra 512GB (unified) × 3× 2× 1× 1

How the estimate works

  • Weights: 501 billion parameters × bits per weight ÷ 8. GGUF formats include their block scales, so Q4_K_M counts as about 4.85 bits.
  • KV cache: 2 × KV heads × head dimension × bytes per value for every cached token in every layer. Global layers cache the whole context; local layers cache only their window.
  • Overhead covers activations, CUDA graphs and fragmentation. 10% is a reasonable start for vLLM or SGLang; measure your own setup.
  • Device count assumes 92% of memory is usable and rounds to tensor-parallel group sizes of 1, 2, 4 or 8 per server.

Questions

Can I run Beam on one GPU?

No consumer or data-center GPU has enough memory. Even at 4-bit the weights take about 250 GB, so you need several large accelerators or a machine with very large unified memory.

Why does a 23B-active model need so much memory?

Each token is routed to a few experts, but which ones changes from token to token, so every expert must be ready. Active parameters set speed; total parameters set memory.

Can experts live in system RAM?

Some engines, such as llama.cpp, can keep expert weights in CPU memory and stream them to the GPU. It works but is much slower; plan for it only for experiments.

When will these numbers be exact?

Weights memory is already exact for a given precision. The KV cache estimate becomes exact once Reflection publishes the model's configuration.