share · copy link · LinkedIn · X
The automation tested well. Buyers liked it in every session, right up to the moment they had to commit to something.
That is where it fell over. A commodity buyer would move through a sourcing flow, reach a recommendation the system had made for them, and stop. Not because the recommendation was wrong. Because there was no way to tell whether it was right, and the next action was worth real money.
Recommendations without reasoning did not make the flow faster. They made it slower, and they moved the hesitation to the most expensive point in it.
The three things that fixed it

What held was a pattern rather than a feature. The system proposes a default. A confidence score sits on it. A plain-language explanation says why this and not something else. And an override takes one action. Each one does a different job, and the pattern stops working if any is missing.
The confidence score sets the stakes. It tells the buyer how much of their own attention this decision deserves. High confidence on a routine call means move on; low confidence on a large volume means look closer. Without it, every recommendation demands the same scrutiny, which is the same as demanding none.
The explanation has to be in the domain, not about the model: this supplier was recommended because the volume was confirmed against a harvest window and the price had held for three weeks. The explanation is useful when it is a sentence the buyer could have written themselves if they had the time.
And the override has to be cheap. If leaving the default costs more than accepting it, the default is not a suggestion. It is a decision somebody else made for you with a confirmation dialog attached.
Buyers accepted automation once they could see its reasoning and leave it. Decision time from selection to quote request fell by roughly a third, and the clarification requests that used to land on the operations team dropped with it.
What it cost
Speed that nobody completes is not speed.
A system that simply decides is faster than one that proposes, explains itself and waits. We gave that up on purpose, because the faster version was the one buyers abandoned. That is the whole trade, and it is worth stating plainly rather than pretending the trustworthy version was also the quickest.
Where the pattern does not port
I also built a reading app with AI summarising in it. A student uploads a textbook and the app summarises a page of it. None of the three affordances fit.
A confidence score tells her nothing she can act on, because she has no way to calibrate it against a page she has not read. An explanation of why the summary kept what it kept is a second claim from the same system that produced the summary, which is one more thing to take on faith rather than a way out of taking things on faith. And an override is meaningless. Override it with what?
The assumption underneath all three
An override is only meaningful if the person can judge the alternative.
A commodity buyer overriding a supplier recommendation knows the market. They have bought that crop before, they know what a reasonable price looks like, and they can tell when a confidence score is misplaced. A student reading a summary of a page she has not read has none of that. Transparency assumes a reader who can evaluate what is being disclosed, and in a learning product that assumption is exactly inverted: the user is there because they do not yet know the material.
What worked instead, by accident

The app ships three summary lengths: a tl;dr, a standard one and a detailed one. I checked all three for correctness against the source text across more than ten PDFs. I built them because people want different amounts of detail. They turn out to be doing something else entirely.
The student cannot evaluate a summary against a page she has not read. She can evaluate a summary against a longer summary, and that comparison is one she is fully equipped to make. The difference between the tl;dr and the detailed version is a readable account of what the short one dropped. That is not transparency, because nothing is being explained. It is a second measurement of the same thing at a different resolution, and she locates the difference herself.
What that leaves
For professional tools where the user knows the domain, I would build the three every time. They are cheap, they are learnable, and they converted automation from something people tested well and abandoned into something they used.
For products where the user cannot check the output, explanation is close to useless, because it adds a claim from the same source as the thing being questioned. What appears to help instead is a second view of the same content at a different resolution, where the gap between the two views carries the disclosure. I did not design that on purpose. The interface still presents three lengths as a menu, which reads as a convenience. Saying out loud what they are for is the next version.
