Workbench / Online
BUILD YOUR COMPUTE STACK
Start with the model. MuseBoard maps the memory and hardware. Settings are mirrored to the URL, so any configuration can be shared.
Compute workbench
B / Estimated VRAM
Inference369.6GB
Serving: weights + KV cache + runtime overhead. Estimate — not a guarantee.
- Model weights335.5 GB
- KV cache0.58 GB
- Runtime overhead33.6 GB
Memory headroom
Does not fit this configuration
Utilization
Target ≤ 90% of 512.0 GB
Decode ceiling
Theoretical, batch 1, bandwidth-bound
VRAM usage
369.6 GB / 512.0 GB
Pinned configuration
16 × NVIDIA GeForce RTX 5090
- Total VRAM
- 512.0 GB
- Topology
- 2 nodes × 8
- Interconnect
- PCIe 5.0
- Board power
- 9,200 W
Compatibility / 16 × RTX 5090
- Ready
Inference
369.6 GB · 72% of 16 × RTX 5090
- Ready
Fine-tuning
442.8 GB · 86% of 16 × RTX 5090
- Cluster
Training
11,824 GB · Needs 416 GPUs — exceeds one consumer node
Alternative configurations
- Needs 16 GPUs — exceeds one consumer node
- Mixture-of-experts: all 671B parameters stay resident in memory; only ~37B are active per token.
- 01SERVE / READY
Inference
Estimate memory requirements for serving models.
weights + KV cache + overhead
Open in workbench - 02ADAPT / LORA
Fine-tuning
Explore memory requirements for adapting existing models.
frozen base + adapters + activations
Open in workbench - 03TRAIN / ADAM
Training
Understand large-scale compute requirements.
weights + grads + optimizer + activations
Open in workbench
MuseBoard gives first-order estimates for planning conversations. It does not replace profiling, capacity testing or detailed ML infrastructure planning.