Now onboarding a limited number of teams Request access
This field note is not published yet. See what is live on the blog.
FIELD NOTES

A number without its noun is not a score

Confidence scores are the standard way AI tools report on themselves, and most of them are unreadable. Not because the number is wrong, but because nobody says what it is a number about.

You open a card and it says 7 out of 10. Seven what. The quality of the draft, the completeness of the supplier's quote, the model's confidence in its own output, the fit between your offer and the requirement. These are four different measurements with four different consequences, and the interface has flattened them into one number.

So the number gets ignored, which is the rational response. An unlabelled score is not information, it is decoration that looks like information.

If a score does not say what it measures, nobody can act on it, and everybody eventually stops looking.

The two we actually show

On a deal there are two things worth scoring, and conflating them is the most expensive mistake available.

ScoreWhat it measuresWhat it decides
Our reply 9 out of 10Whether our outbound draft answers the actual requirement, supported by the recordWhether it can send at all
Quote 6 out of 10Whether an incoming supplier quote is arithmetically sound, complete, and for the specified itemWhether we can price from it

Two measurements, two consequences. A bare 6 out of 10 tells you neither.

Consider what happens when they are merged. A deal shows 7 out of 10. Does that mean your reply is nearly ready to send but the supplier pricing is solid, or that the pricing is shaky but the draft is fine. The first is a writing task. The second is a sourcing task. They go to different people and take different amounts of time.

What makes a score honest

  1. It names its subject. Our reply, the quote, this requirement line. Always attached, never assumed from context.
  2. It has a threshold with a consequence. A score that does not gate anything is a comment. Ours gates sending: below 8, nothing goes.
  3. It comes with the reason. A blocked draft that does not say what to fix has converted work into waiting. The card lists the specific problems.
  4. It is graded by something independent. A model scoring its own output is reporting its own confidence, which is the least useful number in the system. The reviewer is a separate pass against the requirement itself.

The point about self-assessment

This is worth dwelling on, because most confidence scores in AI products are exactly the thing described above. The same process that produced the answer also reports how good the answer is.

That number tracks fluency more than correctness. A confidently wrong draft, the kind written from half the record, will score itself well, because nothing in its own context contradicted it. The whole failure mode is invisible from the inside.

So the grade has to come from a pass that goes back to the requirement and checks the claims against what the record actually contains. Not a second opinion on the prose. A different question entirely.

How this reads day to day

In practice it means the language on the board is slightly more verbose than a designer would like, and considerably more useful. Quote 8 out of 10 qualified. Our reply 6 out of 10, two gaps. Nobody has to ask what the number refers to, which means nobody has to ask twice.

Every score carries its noun

Celestix AI grades outbound drafts against the requirement and incoming quotes against completeness, always labelled, always with the reason, always graded by an independent pass. Below 8 out of 10, an outbound draft cannot send.

Field notes are written from work we do on live deals. All figures and documents shown are sample data. No customer, supplier, or buyer is identified.