Self-hosting
GPU memory planner
Beam is a Mixture-of-Experts model: only 23 billion parameters work on each token, but all 501 billion must sit in memory. Pick a precision, context and hardware to see what it takes.
Attention assumptions
Reflection describes interleaved local and global attention but has not published config.json yet. These placeholders drive the KV cache estimate; change them when the real values are out.
GPUs needed for the weights alone
Minimum devices per precision before KV cache, with 92% of each device's memory usable.
| Accelerator | BF16 / FP16 1,002 GB | FP8 501 GB | GGUF Q4_K_M 304 GB | INT4 / FP4 251 GB |
|---|---|---|---|---|
| NVIDIA H100 80GB | × 14 | × 7 | × 5 | × 4 |
| NVIDIA H200 141GB | × 8 | × 4 | × 3 | × 2 |
| NVIDIA B200 180GB | × 7 | × 4 | × 2 | × 2 |
| NVIDIA B300 288GB | × 4 | × 2 | × 2 | × 1 |
| NVIDIA A100 80GB | × 14 | × 7 | × 5 | × 4 |
| NVIDIA RTX PRO 6000 96GB | × 12 | × 6 | × 4 | × 3 |
| NVIDIA L40S 48GB | × 23 | × 12 | × 7 | × 6 |
| NVIDIA RTX 5090 32GB | × 35 | × 18 | × 11 | × 9 |
| NVIDIA RTX 4090 24GB | × 46 | × 23 | × 14 | × 12 |
| AMD Instinct MI300X 192GB | × 6 | × 3 | × 2 | × 2 |
| AMD Instinct MI325X 256GB | × 5 | × 3 | × 2 | × 2 |
| AMD Instinct MI355X 288GB | × 4 | × 2 | × 2 | × 1 |
| Apple M3 Ultra 512GB (unified) | × 3 | × 2 | × 1 | × 1 |
How the estimate works
- Weights: 501 billion parameters × bits per weight ÷ 8. GGUF formats include their block scales, so Q4_K_M counts as about 4.85 bits.
- KV cache: 2 × KV heads × head dimension × bytes per value for every cached token in every layer. Global layers cache the whole context; local layers cache only their window.
- Overhead covers activations, CUDA graphs and fragmentation. 10% is a reasonable start for vLLM or SGLang; measure your own setup.
- Device count assumes 92% of memory is usable and rounds to tensor-parallel group sizes of 1, 2, 4 or 8 per server.
Questions
Can I run Beam on one GPU?
No consumer or data-center GPU has enough memory. Even at 4-bit the weights take about 250 GB, so you need several large accelerators or a machine with very large unified memory.
Why does a 23B-active model need so much memory?
Each token is routed to a few experts, but which ones changes from token to token, so every expert must be ready. Active parameters set speed; total parameters set memory.
Can experts live in system RAM?
Some engines, such as llama.cpp, can keep expert weights in CPU memory and stream them to the GPU. It works but is much slower; plan for it only for experiments.
When will these numbers be exact?
Weights memory is already exact for a given precision. The KV cache estimate becomes exact once Reflection publishes the model's configuration.