We have now built several production assistants, and replaced two that clients had withdrawn after they gave customers confidently incorrect information. The pattern in both failures was the same, and it was not the model's fault in any interesting sense.
Answering from memory rather than from sources
A language model asked a question about your refund policy will produce a plausible refund policy. It will be well written, appropriately worded, and possibly wrong, because it is generating text consistent with refund policies in general rather than reading yours.
The fix is retrieval. Search your approved content first, and construct the answer only from what was retrieved. If nothing relevant is found, the correct response is that the assistant does not know, and here is a human who does.
No citation, no accountability
An answer without a source cannot be checked by the customer, by your support team, or by you. Attaching the source document to every response changes the dynamic entirely: customers can verify, agents can correct the underlying document rather than the individual answer, and you can audit what the assistant has been telling people.
Citations also create a useful feedback loop. When an answer is wrong, the citation shows whether the retrieval failed or the source document itself was out of date. In our experience it is usually the document.
No boundary on what it may discuss
Both withdrawn assistants had been permitted to answer questions about billing. One told a customer their subscription would be refunded. It would not have been. The customer had it in writing, and the company chose to honour it, which was the right call and an expensive one.
Some topics should escalate immediately regardless of confidence: refunds, cancellations, account changes, anything with a financial or legal consequence. The assistant hands over with the full transcript. The customer is not made to repeat themselves and no commitment is made that the business did not intend.
Measuring rather than assuming
Before launch, assemble a few hundred real historical questions with the answers your team actually gave, and measure the assistant against them. This produces a number you can improve and a regression test for every subsequent change. Without it, quality assessment reduces to whoever tried it most recently having an opinion.
The trade that works
An assistant that declines more often is trusted more, and an assistant that is trusted is used. The instinct to maximise coverage is exactly the instinct that produces the withdrawal.