Financial Grounding
The Substrate Measured: What Rented Compute Actually Costs
Every number on this page is derived from the repository's append-only cost ledger. In a marketplace where 62% of machines fail to deliver working GPUs, fail-fast verification is what makes the economics work.
| Substrate Measure | Observed Value | Accounting Definition & Grounding |
|---|---|---|
| Distinct Instance Creates | 121 | Counted by distinct provider instance IDs across all campaigns. |
| Total Create Events | 127 | Includes immediate retries following transient API timeouts. |
| Never-Delivered Instances | 75 (62.0%) | Instances that froze during pull, failed SSH, or died on bf16 preflight. |
| Time to Running (Median / p95) | 2.4 / 8.6 min | Measured only when a host successfully delivered a verified GPU. |
| Mean Alive per Bad Host | ~12.3 min | Short lifetime enforced by 6-min frozen pull watchdog & 10s preflight. |
| Spend on Bad Hosts | $4.96 | 11.9% of total campaign spend. Absorbed without stalled engineering. |
| Total Campaign Marketplace Spend | $41.62 | Includes all 121 instances, all 75 failures, and local verification. |
| Gated Model Versions | 16 | Model candidate iterations evaluated through the framework gate. |
| Deploy Gate Refusals | 6 (38%) | Refused due to statistical shortfall, canary failure, or format clash. |
Why Two in Three Rented Hosts Fail to Deliver
Roughly two in three creates died before any training occurred, and on documented platform-wide bad days the failure rate reached three in four. That is not the pipeline failing: that is the consumer GPU marketplace being what it is.
Marketplace providers aggregate rigs located in residential basements, student apartments, and makeshift mining frames. Rigs report modern GPUs while running six-year-old NVIDIA drivers, sit behind throttled 20 Mbps residential uplinks that freeze on a 15GB container pull, or crash the moment PyTorch allocates a tensor.
The number that proves the engineering is $4.96: the total cost of all 75 bad host attempts combined. Because tuneharness detects frozen downloads in six minutes and verifies compute in ten seconds, dead hosts are terminated before they can idle-burn dollars.
Interactive Spend Comparison Calculator
Use the interactive calculator below to model your expected fleet spend. Adjust training volume and failure rates to see how renter-side guardrails protect your budget.
The Five Money Traps We Closed With Code
Every defensive guard in tuneharness exists because a real incident cost money:
- The 105-Minute Stale-Driver Idle Burn: A host passed
nvidia-smi, accepted the container, and sat idle while training failed to initialize due to a driver incompatibility. An operator noticed after 105 minutes. We wrote the post-boot bf16 matmul preflight the same evening; now bad drivers are caught and destroyed within ten seconds. - The 30-Minute Frozen Container Pull: Consumer hosts with saturated uplinks frequently freeze on multi-gigabyte layers. The boot progress parser now enforces a hard six-minute progress deadline.
- The Duplicate Instance Leak: When an API call timed out, retrying naively left an unrecorded instance running on the provider. The lease manager now reconciles by deterministic label before issuing any retry.
- The Ctrl-C Abandoned Rental: An operator terminating a local CLI with Ctrl-C once left a GPU instance running overnight. Exit traps now intercept all abort signals, execute an automatic STOP command, and preserve the disk volume.
- The Silent Filter That Matched Nothing: A search filter query emitted
gpu_ram>=24576(MB) while vast.ai expected GB, causing the search to return zero offers and appear as a dead market. We added live-market assertion tests that verify filters actually match real inventory.