Skip to content
MUSEBOARD

Workbench / Online

BUILD YOUR COMPUTE STACK

Start with the model. MuseBoard maps the memory and hardware. Settings are mirrored to the URL, so any configuration can be shared.

Compute workbench

Workbench / Online
Model / Llama 70BUnits / GB (10⁹ B)

A / Model configuration

Meta
B

Editing the count switches to a custom model with an estimated architecture.

Precision0.5 B / param
Context length8,192 tokens
Batch sizeConcurrent sequences
Workload
Best fit

Architecture / published config

Layers
80
Hidden
8,192
KV heads
8 × 128
KV / token
320 KB

B / Estimated VRAM

Training

1,262GB

Full training: BF16 mixed precision + Adam. Estimate — not a guarantee.

  • Model weights141.2 GB
  • Gradients141.2 GB
  • Optimizer states847.2 GB
  • Activations17.2 GB
  • Runtime overhead114.7 GB

Memory headroom

274.5 GB

Free after estimated load

Utilization

82%

Target ≤ 90% of 1,536 GB

Aggregate bandwidth

64 TB/s

8 × 8 TB/s

VRAM usage

1,262 GB / 1,536 GB

0 GB8 × 192 GB

Recommended configuration

8 × NVIDIA B200 192GB

Total VRAM
1,536 GB
Topology
Single node · TP
Interconnect
NVLink 5 · 1.8 TB/s
Board power
8,000 W

Compatibility / 8 × B200

  • Inference

    41.5 GB · 3% of 8 × B200

    Ready
  • Fine-tuning

    64.0 GB · 4% of 8 × B200

    Ready
  • Training

    1,262 GB · 82% of 8 × B200

    Ready

Alternative configurations

  • Full training in INT4 is not a standard setup — the estimate uses BF16 mixed precision.
  • Assumes no ZeRO / FSDP sharding of optimizer states across data-parallel ranks.