Now onboarding a limited number of teams Request access
This field note is not published yet. See what is live on the blog.
FIELD NOTES

A fluent reply is the dangerous kind of wrong

Every model writes well now. That was the hard part five years ago and it is table stakes today. The problem is that good writing is no longer evidence of a good answer, and on a technical bid the difference is expensive.

There is a specific failure that only shows up once AI is genuinely good at writing. A draft comes back clean. The tone is right, the structure is right, it answers the question that was asked. Everyone reads it, nobody objects, it goes out.

Three weeks later a buyer holds you to a delivery commitment nobody remembers making.

The draft was not badly written. It was confidently written about something it had not checked. And confident prose is the one thing a reviewer is least likely to interrogate, because reading it produces the feeling of having verified it.

Bad writing gets caught. Good writing about the wrong facts is what actually reaches the customer.

Where the wrong facts come from

Not from the model inventing things, usually. On a bid the more common cause is duller and harder to spot: the draft was written from a fraction of the record, and the missing fraction was where the constraint lived.

  • The lead time in the reply came from the supplier's first email. The revised one, sent eleven days later as a PDF, said something different.
  • The technical answer was correct for revision A of the specification. Revision B arrived with the amendment and changed the operating temperature.
  • The commercial position contradicts what a colleague already committed to in a separate thread with the same buyer.
  • The scope answer describes what you usually supply, not what this particular tender excludes.

In every one of those cases the reply reads perfectly. It is internally consistent, professionally worded, and wrong in a way that only the original document would reveal.

Why review does not catch it

Because review is being asked to do something it is bad at. A human reviewing a polished draft is checking whether it reads correctly, not re-deriving every factual claim from source documents. Nobody has time to reopen nine attachments to verify a sentence that already sounds right.

This gets worse as volume rises. A reviewer who sees four drafts a week reads them carefully. A reviewer who sees forty starts trusting the ones that look like the ones that were fine last time. Fluency becomes a proxy for correctness, and the proxy is wrong.

The two things that fix it

First, the draft has to be written from the complete record rather than the visible message. That removes most of the wrong facts at the source, because the constraint that would have been missed is in the context that produced the sentence.

Second, and this is the part most tools skip, something other than a human reading prose has to check the draft against the requirement before it can go. We grade every outbound draft on a 10-point scale against the actual requirement it answers. Below 8, the send does not fire, and the card lists what to fix.

The grade is not a quality score for the writing. It is a check on whether the claims in the draft are supported by what the record actually contains. A beautifully written reply that asserts a delivery date no supplier has confirmed scores badly, which is the entire point.

What changes in practice

The obvious change is that fewer wrong commitments leave the building. The less obvious change matters more: review stops being a bottleneck staffed by the most senior person available.

When the gate is doing the factual check, a human review is answering a different and much better question. Not is this correct, but is this the position we want to take. That is a judgement call, it is the one people are actually good at, and it takes a fraction of the time.

Fluency is not a control

Celestix AI writes every draft from the complete deal record and grades it against the real requirement before it can send. Below 8 out of 10, nothing goes, and the fix list is on the card. The same bar applies to drafts written by people.

Field notes are written from work we do on live deals. All figures and documents shown are sample data. No customer, supplier, or buyer is identified.