Best Questions to Ask an AI Medical Coding Vendor
A demo shows you a curated encounter processed by a tuned system in front of an audience. It is not evidence. These questions are.
Ask them in this order. The early ones are easy and the later ones are not, and vendors who cannot answer the early ones rarely survive to the late ones.
On accuracy
- What corpus were your published numbers computed on, and how was it sampled?
- What is your top-5 recall, and separately, your micro F1?
- What share of your recommendations lack supporting text in the source document?
- Can you break accuracy out by specialty rather than blended?
- How was the answer key constructed — originally billed codes, or post-rework codes?
On workflow
- What is your measured average review time per encounter, and on what case mix?
- What happens when the coder disagrees — is the override recorded and reviewable?
- Can we operate without any EHR integration on day one?
- What formats can you ingest? Name the ones you cannot.
- What is the failure mode when a document is illegible or incomplete?
On compliance
- How long is evidence retained, and does retention survive contract termination?
- Can we export a complete audit package for an arbitrary date range without vendor assistance?
- Is there any configuration in which a code is submitted without human acceptance?
- Where does processing occur, and is an on-premise deployment available?
On commercials
- Is pricing per seat, per encounter, or per claim — and what happens at 3x volume?
- What is the total first-year cost including implementation, not the license line?
- What is the contractual definition of the value you are promising, and who measures it?
- What is the exit process and what do we take with us?
On the company
- Which customers of our size and specialty mix can we speak to unsupervised?
- What is on your roadmap that is not shipped, stated plainly?
- What does your product do badly today?
The last question is the tell
Any vendor who cannot name a weakness has either not deployed at scale or is not being straight with you. Both are disqualifying.
For our part: deep in-workflow EHR embedding is roadmap, not shipped, and specialty performance is not uniform across our corpus. Ask us for the breakdown.