Medical coding is one of those occupations that most people encounter only as a line item on a bill they don't understand. But behind every Explanation of Benefits, behind every approved claim and denied procedure and risk-adjusted premium, sits a translation act so consequential that getting it wrong can trigger a federal investigation, and so invisible that the person most affected by it will never know it happened.
The coder reads a clinical encounter and converts it into standardized codes drawn from more than 70,000 diagnostic categories.1 Those codes determine what gets paid, what gets covered, what shows up in a patient's permanent record, and what conditions become visible at the population level. The translation is always lossy. Clinical reality is continuous and contextual; the code set is discrete and categorical. The coder lives in the gap.
That gap is filling up with software. Sixty-three percent of U.S. healthcare organizations already use AI and automation in revenue cycle operations.2 The tools are fast, confident, and often plausible. They are also, by independent measure, far less accurate than vendor marketing suggests. One benchmark found GPT-4's exact match rate for ICD-10-CM codes was just 33.9%.3 Specialized systems do better on common codes and worse on rare ones, which is exactly backwards from where human judgment matters most.
We spoke with Dolores "Lola" Cifuentes, a fictional inpatient coder with three decades of experience at a large academic medical center, about what it means to translate suffering into categories and what happens when the translation starts running on autopilot. Lola does not exist, but the structural pressures she describes are documented, measurable, and affecting millions of patients whose codes they will never see.
You've been coding for thirty years. What does the job actually feel like day to day?
Lola: You're reading someone's worst week and turning it into a string of alphanumerics. That's the whole job. A guy comes in with chest pain, turns out it's a STEMI, he's got diabetes, chronic kidney disease, he's on dialysis, and oh, the surgeon notes "possible peripheral neuropathy, left foot" in a throwaway line on page nine. Every one of those things is a code. Some are obvious. Some require you to read the whole record and figure out what the physician actually meant versus what they technically wrote.
And some, like that neuropathy, you look at and think: is this a confirmed diagnosis, or is the surgeon just thinking out loud? That distinction determines the code. And the code determines what follows the patient.
What follows?
Lola: Everything. Whether their insurer considers it a preexisting condition. Whether a future procedure gets prior authorization. Their risk score, which affects what Medicare pays for their care going forward.4 None of that is visible to the patient. They get a bill. They don't get the codes. They definitely don't get a note that says, "Your coder spent eleven minutes deciding whether 'possible peripheral neuropathy' warranted a query to your surgeon."
Tell me about the query process.
Lola: This is the thing nobody outside coding understands, and it's the thing I'm most worried about losing. When I can't determine the right code from the documentation, when it's ambiguous or incomplete or contradictory, I'm supposed to query the physician. Go back and say: did you mean this, or this? Is this a confirmed diagnosis or a differential? Left or right?5
That conversation is where ambiguity gets resolved correctly. Not by pattern-matching. By the person who was in the room with the patient.
And AI tools don't query.
Lola: They don't. They pick the statistically most likely code and move on. Confidently. And look, for a straightforward outpatient visit, that's fine. But for the cases where I would have queried? The AI just resolves it. Picks a code. The ambiguity disappears. The record looks clean. And nobody knows that a question should have been asked.
How often would you say you query in a typical week?
Lola: On complex inpatient records? Multiple times a day. Some weeks I'm sending fifteen, twenty queries. That's fifteen times the documentation wasn't clear enough to code accurately, and fifteen times a conversation with the physician changed what went into the patient's record.
Now imagine an AI handling those same records. Zero queries. Fifteen codes assigned anyway.
You're now reviewing AI-generated code suggestions. How is that different from coding?
Lola: Completely different, and I don't think the people designing these workflows understand why.
When I code from scratch, I'm reading the record with a question in mind: what happened to this patient, and what codes capture it? I'm building the answer. When I'm reviewing AI suggestions, I'm reading the record with an answer already in front of me, asking: is this right?
Those are not the same cognitive task. The second one is dramatically easier to say yes to. Confirmation bias isn't a side effect here. It's the architecture of the workflow.
I review maybe 200 codes a day now. I am not rereading every record at the depth I would if I were coding it myself. I can't. Nobody can. So I'm catching obvious errors and approving plausible suggestions. And plausible is doing a lot of work in that sentence.
The FY2025 Medicare improper payment rate was 6.55%, roughly $28.8 billion, with incorrect coding among the leading causes.6 Does AI make that better or worse?
Lola: Yes.
That's not an answer.
Lola: It is, though. AI will reduce certain kinds of errors: typos, missed codes on routine visits, simple sequencing mistakes. And it will introduce other kinds of errors that are harder to detect because they look right. A code that's plausible but not accurate. A code that's technically defensible but doesn't reflect what actually happened.
The error rate might go down on paper while the character of the errors gets worse. Fewer wrong answers, more almost-right answers. Which, in coding, can be worse than obviously wrong ones, because nobody flags them.
There's a legal dimension here. False Claims Act settlements exceeded $6.8 billion in fiscal year 2025, with over $5.7 billion tied to healthcare.7 Who's accountable when an AI suggests a code and a coder approves it?
Lola: Me. My credential is on that code. That hasn't changed.
What's changed is how much of the underlying judgment is actually mine.
I'm signing off on work I didn't produce, at a volume that makes deep review impossible, in a system where both overcoding and undercoding can constitute fraud.8
Meanwhile, payers are running their own algorithms to automatically downcode claims without reviewing the medical record.9 So I'm caught between an AI suggesting codes on one side and an AI contesting them on the other, and my name is the one on the line. It has a certain dark comedy to it, if you're into that sort of thing.
What about new coders coming into the field? If AI handles the routine cases, where do they learn?
Lola: That's the question that keeps me up at night. I learned complex coding by doing simple coding first. Thousands and thousands of routine encounters that taught me how documentation works, how physicians think, what the codes actually mean in practice.
If new coders start by reviewing AI suggestions on routine cases, which is what's happening now, they're not building the same foundation.
They're learning to approve, not to derive.
And in five years, when the AI gets a complex case wrong and nobody catches it, we'll wonder where all the experienced coders went. We automated their training ground.
What keeps you doing this work?
Lola: The hard cases still need me. The twelve-page inpatient record with six comorbidities and contradictory notes from three specialists. No AI is touching that with confidence. Those cases are why I got into this work. The puzzle of it. Reading a record and understanding not just what the physician wrote but what they meant.
That's craft. I don't want to be dramatic about it, but it's a form of care. The patient will never know I was there, but I was there, and I got it right, and that mattered.
The question is whether anyone will be able to do that in ten years. Or whether anyone will notice when they can't.
Dolores "Lola" Cifuentes is a fictional character. Her concerns are not.
Footnotes
-
CMS/NCHS, FY 2026 ICD-10-CM Coding Guidelines, updated October 1, 2025. https://stacks.cdc.gov/view/cdc/250974 ↩
-
HFMA survey data, cited in Helpsquad, "The 2026 Guide to AI in Medical Coding," April 2026. https://helpsquad.com/blog/the-2026-guide-to-ai-in-medical-coding/ ↩
-
Yang et al., "Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying," NEJM AI. https://ai.nejm.org/doi/full/10.1056/AIdbp2300040 ↩
-
RapidClaims, "ICD-10 HCC: A Complete Guide for Healthcare Professionals." https://www.rapidclaims.ai/blogs/icd-10-hcc-coding ↩
-
AHIMA, Standards of Ethical Coding. https://journal.ahima.org/Portals/0/archives/AHIMA%20files/Standards%20of%20Ethical%20Coding.pdf ↩
-
CMS FY2025 improper payment data, cited in Healthcare Compliance Pros, "What Is Upcoding and Downcoding?" https://www.healthcarecompliancepros.com/what-is-upcoding-and-downcoding-how-coding-errors-are-costing-healthcare-practices-today ↩
-
Department of Justice, FY2025 False Claims Act statistics, cited in Healthcare Compliance Pros. https://www.healthcarecompliancepros.com/what-is-upcoding-and-downcoding-how-coding-errors-are-costing-healthcare-practices-today ↩
-
The AMA considers systematic downcoding a form of fraud and abuse, parallel to upcoding. See Healthcare Compliance Pros, ibid. ↩
-
Payer-side algorithmic downcoding practices documented across major insurers. See Healthcare Compliance Pros, ibid. ↩
