← Back to Case Studies

EHR Integration Built Without Waiting On The Hospital

HeartSciences' AI-ECG platform was integrated with a major EHR vendor without its development waiting on anyone else's calendar. ML LABS built an EHR API simulator replicating the targeted workflows at the protocol level, so HL7 message handling, FHIR session management, and per-organization field mappings could be developed and tested before sandbox access existed. The simulator then became the regression harness for the integration it de-risked.

The constraint it removed is a common one. The sandbox is shared, the windows are scheduled, and every test cycle waits on a coordination call with someone else's team. The engineering is not the constraint. The calendar is, and it belongs to a vendor whose incentives do not include your delivery date. This integration required bidirectional HL7 messaging alongside FHIR OAuth flows, each with per-hospital configuration differences that only surface during live testing.

A simulator is a spike with a long half-life. It is the cheapest instrument for learning what you do not yet know about a system you must integrate with — and unlike a spike, it keeps paying after the answer arrives. That is the same argument the cloud backend it was built for makes about seams: they are cheap to design and expensive to retrofit.

An integration engineer and clinical trainer check an unoccupied bedside training station.
Work never waited on a sandbox window.

Work Proceeds Before Sandbox Access

This was not a generic FHIR mock server. It replicated the behaviors the platform depended on — the ones that would have required vendor sandbox access to discover.

Flowchart: the EHR simulator sends HL7 orders and FHIR OAuth to the AI-ECG platform, which returns HL7 results to the simulator. The platform feeds per-organization configuration profiles, which feed the integration test suite.

On the HL7 side, the EHR sends an order message when a physician requests an ECG interpretation and the platform responds with result messages carrying the AI analysis. The simulator replicated that bidirectional flow including per-organization message behavior, placer order number tracking that links each result to its order, and diagnosis segments carrying the billing codes that revenue cycle systems depend on. The messaging layer that had to satisfy all of it is the bidirectional HL7 result-delivery system, and the standard gives it very little help: an analysis of HL7 v2 optionality (JAMIA, 2009) counted 4,132 data elements in the standard result message, with 85% of them designated optional, which is precisely why two hospitals' "compliant" messages disagree with each other.

Isolation Tested Before Production

On the FHIR side, each hospital authenticates through its own endpoint. The SMART App Launch framework (HL7 FHIR, 2024) defines the launch handshake; what it does not define is which credentials, certificates, and enabled features a given hospital binds to it — the part a simulator has to carry. The simulator also replicated those security boundaries, including adversarial scenarios where requests try to cross organizational boundaries, verifying tenant isolation under conditions that would otherwise only surface in multi-organization production. All of it became the foundation of the integration test suite, from order receipt through AI processing to result delivery, catching parsing regressions on every commit.

The simulator's value was never in replicating the vendor's API surface. It was in replicating the per-organization behavioral differences that only surface during live integration testing.

Site Differences Caught In Testing

Three patterns account for most of the rework the simulator absorbed on this integration, and each one is a single field that looks like a footnote during setup.

Patient identifier formatting. Two hospitals on the same vendor sent the identifier in the same HL7 field, but one zero-padded the medical record number to a fixed width and the other left it variable. A parser that worked end to end for the first site would mismatch results to the wrong patient at the second. The per-org profiles exposed that mismatch on the first integration test rather than later, during clinical validation.

Ordering provider representation. One site populated the field with an internal directory ID, another with a national provider identifier, and a third with a free-text initial-and-surname string. Without per-org profiles, the platform would have inherited whichever convention the first hospital used and produced unattributable results everywhere else.

The unsolicited result workflow. Some sites expect results only in response to an explicit order message; others expect results as soon as an interpretation is available, with no preceding order. Invert that assumption and the receiving system either drops the result or queues it for manual reconciliation, and the symptom, a missing result in the EHR view, is indistinguishable from a transport failure until someone reads the message logs. Encoding both behaviors as named configuration profiles meant the routing logic was tested against the variation up front instead of patched after a hospital reported missing studies.

The simulator changes what scarce sandbox time is for: sessions stop being first-pass parsing tests and become validation of work already known-good, each divergence feeding the simulator profile so the next iteration catches it without more vendor time.

Where A Simulator Saves Time

A simulator is not a substitute for validation against the real target system. The question that decides it is whether the external dependency is a scheduling bottleneck or a knowledge bottleneck. Simulators solve the first and can disguise the second: if nobody on the project can describe what the target system does when a required field is absent, building a simulator will not tell you — it will only make the wrong answer run faster.

Three Steps To Unblock The Work

  1. Simulate the blocking surface, not the API. Start with the workflows that block development and that carry the most variation between organizations.
  2. Calibrate on every real interaction. Each sandbox session should update the simulator's per-organization profiles for any behavior difference that session revealed.
  3. Build the test suite on top from day one. The durable value is regression coverage across every per-organization profile, long after the sandbox stops being scarce.

Risk Named Before Budget Moves

On this integration the binding constraint was never engineering complexity — it was iteration speed imposed by a dependency ML LABS did not control, and the simulator converted an external constraint into an internal, schedulable one.

This is what de-risking a build is supposed to produce before anyone commits the budget: not a slide deck, but a written plan that names the risk. That is what comes out of the first call — the architecture, a build plan, and a price, settled before any budget moves. Whether the risky part of your integration is a scarce sandbox, an undecided data path, or a workflow nobody has written down, the answer is the same: name it in writing before the budget is committed, and let the build be where it gets proven running rather than argued.

References

  1. Sujansky WV, Overhage JM, Chang S, Frohlich J, Faus SA. The Development of a Highly Constrained Health Level 7 Implementation Guide to Facilitate Electronic Laboratory Reporting to Ambulatory Electronic Health Record Systems. Journal of the American Medical Informatics Association, 2009.
  2. HL7 International. SMART App Launch Implementation Guide. HL7 FHIR, 2024.

NEXT · TO PRODUCTION

Could this work for you?

Two minutes. Find out where you stand.

15 minutes · no charge · with Omar