← Back to Case Studies

Billing Rules Automated At 98%+ Expert Agreement

ML LABS replaced hand-applied billing rules at HeartSciences with a rule engine that ran in shadow mode against live claims data until it reached 98%+ agreement with expert-adjudicated determinations, and only then became the primary processor, with human reviewers moving from processing every claim to reviewing exceptions. The automation was delivered under the ongoing engineering retainer: the same practice that built their AI-ECG platform's cloud backend still owns their engineering, and this system is one result of that ownership. A retainer that never ships is a subscription. This is what the other kind looks like.

What it replaced is common. A team reads a case, checks it against the eligibility rules, coding requirements and payer-specific formatting, and keys the result into a billing system. Errors made that way do not announce themselves — they leave as unpaid claims, late cash, and a number that nobody reconciles back to a cause. The Medicare improper payment rate ran at roughly 7% in 2025, close to $29 billion, according to the Fiscal Year 2025 Improper Payments Fact Sheet (CMS, 2025). Most of it is not fraud. It is documentation and coding rules, applied inconsistently, by people, at volume, making it a determinism problem, and software is unusually good at determinism.

A medical billing specialist working through claim folders at her desk
Reviewers moved to the exceptions.

Where Deterministic Rules Pay Most

  • The density of rules is rising. Eligibility logic, coding logic, and payer logic interact too tightly for reliable execution inside a spreadsheet, and a queue of humans performs worse the more the rules interact.
  • Turnaround time is material. Delay affects cash flow, patient operations, and service quality, and human queueing adds latency no headcount removes cleanly.
  • Auditability matters. Program integrity guidance (CMS) is explicit: if a workflow affects payment or compliance, the reasoning path must be recoverable. A spreadsheet records an answer, not the reason it was reached.

The moment a workflow needs deterministic rule application and durable auditability at once, manual processing stops being cheap even when the labor rate still looks low.

Every Claim Gets The Same Rules

The instinctive move, automating the existing process, is the wrong one. RPA bots replicate human keystrokes faster, and so replicate process defects at machine speed. What replaces manual processing is not a faster clerk. It is process re-architecture: define the rules, exception paths, and validation layers explicitly, then automate the rest.

Pipeline: data extraction from FHIR, HL7 and adapters feeds an eligibility rule engine, then CPT code generation, validation and submission, and an audit trail with an exception queue.

Data extraction normalizes records from heterogeneous EMRs into one canonical record: FHIR R4 APIs where they exist, adapters for legacy interfaces (HL7 v2, flat files) where they do not, and validation at the point of extraction, not three stages downstream.

The eligibility rule engine is the core. It encodes business logic as explicit, testable, version-controlled rules that produce identical output for identical input. Medicare eligibility for chronic care management alone means checking Part B enrollment, hospice status, institutional stays, patient presence, and service time thresholds, conditions that interact across billing periods. The engine produces a complete audit trail, correct rule interaction, and atomic updates when CMS changes a rule, so a rule change becomes one deployment rather than a retraining exercise for everyone who touches a claim.

Every Failure Routed With Its Reason

CPT code generation aggregates clinical staff time across the billing period and maps it to the correct billing code, with modifier application and correct-coding validation (CMS) running before submission rather than after denial. Validation then runs in tiers, and each failure is classified by severity and routed rather than dumped in a queue:

  • Schema validation: required fields and data types
  • Referential validation: active NPIs, valid facility codes
  • Temporal validation: service dates inside the billing period
  • Business rule validation: diagnosis codes that support medical necessity

With the engine's audit trail beside each tier, the reasoning path behind each claim stays recoverable when a payer or an auditor asks for it.

Cutover Earned Against Expert Judgment

Replacing a live billing process is a migration, not a product launch. The automated system ran in shadow mode at first: processing live claims data, submitting nothing to payers, and having its outputs compared against those of the manual team every day, with every discrepancy investigated. That phase surfaces errors in both systems at once, which is the point — the manual process is not a gold standard; it is only the incumbent.

The gate to supervised automation was the 98%+ agreement bar, and it is meaningful because the comparison set was expert adjudication, not the manual team's own output. Agreeing with the incumbent's mistakes is not accuracy. After the gate comes full automation, where human involvement narrows to exception-queue management and periodic accuracy audits, and each new facility onboards through a compressed shadow cycle rather than a fresh build. Audit-trail continuity is kept across every phase, and a site showing elevated rejections goes back a phase, not forward. The rules that make this survivable across many sites belong to the multi-site clinical operations architecture, the second system delivered under the same retainer.

Source Data Checked Before The Build

This approach assumes upstream data (EMR records, eligibility feeds, encounter logs) exists in extractable form. Where it does not, automation makes the mess faster, not smaller.

The signal is measurable before anyone commits to the build: if extraction finds more than 15–20% of records with missing critical fields, or one clinical event is coded differently across sites, source-data instability is the constraint and no rule engine will out-argue it. Normalize first, then automate. Automation inherits a data problem. It does not fix one.

Three Steps To A Safe Cutover

  1. Audit one facility's workflow end to end. Every data source, decision point, and handoff, measuring hidden costs (rework, delay, audit exposure) beyond the labor line.
  2. Validate the rules against history before touching live claims. Hold the engine to the 98%+ agreement bar against expert-adjudicated outcomes before it submits anything.
  3. Shadow, then cut over. Connect to live feeds, submit nothing, and investigate every discrepancy. The discrepancies are the specification you did not know you were missing.

Reviewers Spend Time On Exceptions

If billing is still run by people applying interacting rules from a spreadsheet, the constraint is not effort or talent. The workflow has outgrown the control point. The path out is a deterministic rule engine with explicit validation tiers, a shadow-mode migration that earns its cutover against expert adjudication, and an exception queue treated as a feature.

That is one contained workflow, made into production software against written targets by an engineering partnership and kept running by the person who built it, because a rule engine meets the payers' real behavior only once it submits. If several billing and clinical workflows are moving at once and each new one lands on the same person, a standing engineering owner beats a sequence of projects: the compounding case for ownership.

References

  1. CMS. Fiscal Year 2025 Improper Payments Fact Sheet. Centers for Medicare & Medicaid Services, 2025.
  2. CMS. National Correct Coding Initiative (NCCI). Centers for Medicare & Medicaid Services.
  3. CMS. Center for Program Integrity. Centers for Medicare & Medicaid Services.

NEXT · TO PRODUCTION

Could this work for you?

Two minutes. Find out where you stand.

15 minutes · no charge · with Omar