share · copy link · LinkedIn · X
Rebuilding an eighty-four component design system by hand takes months. I did it in days. That number is the reason this essay exists. Without it, what follows is a list of things that went wrong, and any reasonable person would ask why I kept going. With it, the question becomes the interesting one: what does supervision have to mean when the thing you are supervising works that fast?
I have an answer now. It came from seven failures across two projects, and they split cleanly into two kinds.
The chart that was correct until someone used it
The agent built a set of chart components. On the component page they were fine. Every one of them overflowed the moment it was placed inside a card, because none had scale constraints.


Neither defect was visible where the components lived. Both were obvious within seconds of putting a chart into a real dashboard.
Two more from the same rebuild
It deleted work without knowing what it was deleting. Rebuilding a page, it removed a documentation frame. Inside that frame was a sixteen-variant avatar illustration set from a remote library with no recoverable keys, so I re-pasted every variant by hand.
It broke adjacent things while fixing one. Correcting a single artwork type, it stripped the first child from every variant in the set. Thirty components lost their initials and icon artwork. The fix was correct. The blast radius was not.
What those four have in common
A design file exists twice. It is a node tree, and it is a rendered thing. The agent operates fluently on the node tree. It reads structure, computes combinations, applies tokens, restructures variant axes, and does all of it faster and more completely than I can. Every one of those operations was technically correct. Design judgement happens on the rendered thing. Whether a chart fits a card. Whether a donut is round. Whether the artwork is still there.
Every one of those four defects lived in the gap between the two, and I found every one of them the same way: I rendered the change and I looked at it.
Not once did I catch anything by reading the code the agent wrote. That is not a statement about this particular agent. It is a statement about where the two representations diverge, which is a property of the medium.
What supervision turned out to mean
I was accountable for the rendered result of every change. That meant rendering every change, looking at each one in the context it would actually be used in rather than on the page it was built on, and treating the component page as insufficient evidence.
It also meant holding the judgement about what the work meant. I had the agent decode the library into its node tree, over fourteen thousand nodes, and compute for every component its axes, their values, how many combinations were possible and how many were built. Reading that inventory, I noticed that Mobile header carried a Device axis whose two values changed the frame width from 360 to 375 and nothing else. Empty state had one. Dialog had one. Modal header was the same fault in a different form. Four components encoding a breakpoint as a property of the component rather than a context the component sits in. No query detected that. There was no query for it.
The completeness was the agent’s. The finding was mine.
The three that had nothing to render
Rendering everything works until the failure has no artifact. The remaining three did not, and one of them came from a different project entirely: a video meeting product I specified end to end and handed to an agent to build.
It told me the right answer was impossible. I asked for a content slot. It built an instance-swap workaround and said the direct approach could not be done in Figma. Figma slots exist. They were the right answer. The agent did not fail to do the thing. It declined to, confidently. An admission of uncertainty would have sent me straight to the documentation. An assertion nearly closed the question, in the direction of a worse implementation that would have shipped and been maintained.
It built a capability I had specified and I did not notice for a version.

The meeting product’s waiting room had a default and an enforcer and no control of any kind. A scheduled meeting was gated permanently. An instant one could not be gated at all. Changing either meant going into the database by hand. I had written that capability into the specification, the build did not have it, and nothing in the document said so.
It broke a rule I had written, and the check that catches it never ran.

A rule in that project says the LiveKit client ships on the room route and nowhere else. A dev page imported a participant row, which reaches a panel, which imports the LiveKit React package for a quality subscription. An import is all or nothing, so the client landed in a second route’s first load. The fix was a file move rather than a redesign. The check existed, worked the moment it was asked, and was not asked.
These three are the same shape. A claim about what is possible, a document claiming a capability, a rule about what ships where. None of them produce something to look at. Each one needs a different kind of verification, and none of that verification happens unless you go and ask for it.
The method
Build programmatically. Render every change. Verify in the context of use. Treat any claim about what is impossible as unverified. Use the capability rather than reading the specification back. Run the checks you wrote.
That is the whole thing. It is not sophisticated and it did not come from anywhere except being wrong repeatedly.
What I cannot claim
I caught all seven, which means this entire taxonomy is built from near-misses. I do not have an example of something that got through, which is either evidence the method works or evidence that I have not run it long enough. Two projects now, but one maintainer on both, and on both I was the person who wrote the specification and the person who reviewed the build. I suspect it gets harder when those are two different people, because most of what I caught, I caught because I already knew what the work was supposed to do. And the speed claim cuts both ways: months to days is real, and it is also what made eight or nine publishes in a few days possible, which is why the system now has a separate work-in-progress space.
Four failures that were visible the moment I rendered them, three that produced nothing to render at all, and a method that follows from the difference. So I render everything, and I check that what I specified actually exists. That is the entire discipline, and it came from being wrong seven times in ways that never announced themselves.
