You open a card and it says 7 out of 10. Seven what. The quality of the draft, the completeness of the supplier's quote, the model's confidence in its own output, the fit between your offer and the requirement. These are four different measurements with four different consequences, and the interface has flattened them into one number.
So the number gets ignored, which is the rational response. An unlabelled score is not information, it is decoration that looks like information.
If a score does not say what it measures, nobody can act on it, and everybody eventually stops looking.
The two we actually show
On a deal there are two things worth scoring, and conflating them is the most expensive mistake available.
| Score | What it measures | What it decides |
|---|---|---|
| Our reply 9 out of 10 | Whether our outbound draft answers the actual requirement, supported by the record | Whether it can send at all |
| Quote 6 out of 10 | Whether an incoming supplier quote is arithmetically sound, complete, and for the specified item | Whether we can price from it |
Two measurements, two consequences. A bare 6 out of 10 tells you neither.
Consider what happens when they are merged. A deal shows 7 out of 10. Does that mean your reply is nearly ready to send but the supplier pricing is solid, or that the pricing is shaky but the draft is fine. The first is a writing task. The second is a sourcing task. They go to different people and take different amounts of time.
What makes a score honest
- It names its subject. Our reply, the quote, this requirement line. Always attached, never assumed from context.
- It has a threshold with a consequence. A score that does not gate anything is a comment. Ours gates sending: below 8, nothing goes.
- It comes with the reason. A blocked draft that does not say what to fix has converted work into waiting. The card lists the specific problems.
- It is graded by something independent. A model scoring its own output is reporting its own confidence, which is the least useful number in the system. The reviewer is a separate pass against the requirement itself.
The point about self-assessment
This is worth dwelling on, because most confidence scores in AI products are exactly the thing described above. The same process that produced the answer also reports how good the answer is.
That number tracks fluency more than correctness. A confidently wrong draft, the kind written from half the record, will score itself well, because nothing in its own context contradicted it. The whole failure mode is invisible from the inside.
So the grade has to come from a pass that goes back to the requirement and checks the claims against what the record actually contains. Not a second opinion on the prose. A different question entirely.
How this reads day to day
In practice it means the language on the board is slightly more verbose than a designer would like, and considerably more useful. Quote 8 out of 10 qualified. Our reply 6 out of 10, two gaps. Nobody has to ask what the number refers to, which means nobody has to ask twice.