—
—Empirical single-GPU training planner Snapshot 01
Rent the result,
not the spec sheet.
Compare measured OLM training throughput against current RunPod rates. Find the cheapest route, the fastest route, and every non-dominated option between them.
01 / Configure the workload
Cost frontier
Cost and time use measured steady-state tokens/second at the selected batch-size objective.
Price assumptions RunPod Secure Cloud
Edit any rate to model a different host or a future RunPod price. Overrides stay in this browser only.
Price snapshot: RunPod Pods,
—
——
Neither slower nor more expensive—
—Wall time × compute cost
Training frontier
02 / Compare every board
Measured options
Median metrics are computed at a shared batch size across the available seeds. Click a row to inspect its raw sweep.
| GPU | Batch | Tokens/s | MFU | Peak VRAM | Replicates | Rate | Time | Compute cost |
|---|
03 / Inspect the evidence
Batch saturation
Every dot below is a measured, stabilized training step—not a theoretical throughput estimate.
Select a GPU
—
Per-seed measurements
—| Seed | Batch | Status | Step ms | Tokens/s | Nominal MFU | Configured MFU | Allocated VRAM | Jitter |
|---|
04 / Know what is missing
Coverage matrix
Coverage for the selected model across every tested context. “OOM” means the attempted configuration did not fit.
05 / Read the fine print
Method, not magic
This is an empirical compute-cost comparison, not a promise about end-to-end job billing.
Pick one shared batch
For each GPU, batch metrics are aggregated by median across available seeds. The selected batch maximizes either TPS or configured-clock MFU.
Convert throughput to time
hours = training tokens ÷ measured TPS ÷ 3,600
Apply the rental rate
compute cost = hours × USD per GPU-hour
Keep the boundary honest
Estimates exclude initialization, data loading, storage, checkpoint I/O, evaluation, failures, taxes, and multi-GPU communication. RunPod states that Pods are billed by the minute.