Clinical experts on HeartSciences' MyoVista Insights platform now find records by typing the question the way they would say it, such as "abnormal AI impressions for Site 12 this month", and get sub-second results at production data volumes, with desktop and mobile served unmodified. ML LABS built that natural language search layer on the cloud ECG backend behind the platform, the multi-tenant AI platform whose backend ML LABS engineered, which has reached clinical production in two countries and keeps expanding. Clinical experts needed to search a worklist of records annotated with AI-generated outputs, and the phone was a first-class surface in the requirement.
The cost it removes is a familiar one. The system holds the answer, and getting it out means a screen of dozens of fields, a stack of dropdowns, and a combination rule that only two people in the building actually understand. So experts ask one of those two people, export to a spreadsheet, or work from memory. Every one of their questions can be answered from data that already exists; the gap is the distance between how an expert thinks and how the system insists that it be asked, and that distance is a tax charged to your most expensive people on their busiest days.

Experts Ask In Their Own Words
The people who need the data most urgently think in the language of their own work: unconfirmed AI findings from last week, abnormal impressions for Site 12. The core of the search is a constrained intent parser that maps a query like that onto a known set of domain entities and structured filters, such as a classification status, a site identifier, a confirmation state or a date range, and translates them into an executable query against fields that exist in the schema. Type "abnormal AI impressions for Site 12 this month" and it resolves three entities and one date range, then runs the exact query those four things define.
The parser's reliability comes from constraint, not sophistication. It recognizes every entity that maps to a filterable field, and nothing else. Bounded vocabulary means predictable results.
Misreads Surface Instead Of Hiding
A general LLM-driven parser demos better and operates worse. A free-form parser fails silently: it returns a plausible result set that misses what the user asked for, and nothing in the interface distinguishes that from a correct answer. A constrained parser fails loudly: when a phrase does not map to a known entity, the system says so and offers to refine.
Take "borderline impressions Site 12 not yet confirmed last week." A free-form parser might read "borderline" as a confidence modifier, a severity class, or free text to match anywhere: three interpretations, and no way to see which was taken. The constrained parser resolves "borderline" against a lookup. If it is a registered classification value, the filter pins and the query runs; if not, the parser asks whether the user meant the equivocal-finding class.
That matters because trust in these systems is asymmetric. People abandon an algorithm faster after seeing it err than they adopt one after seeing it succeed (Journal of Experimental Psychology: General, 2015), and a search tool that silently drops the one record an expert was looking for has spent its credibility with that expert permanently. An assistant that is confidently approximate is worse than a plain filter UI that is honestly rigid.
Sub-Second At Production Volumes
The worklist holds two layers of searchable data: the raw operational records, and the AI-generated outputs the inference pipeline attaches to them. Search runs across both as one surface. "Abnormal AI impression" searches the outputs, "Site 12 records from last week" searches the primary records, and a query that mentions both returns the intersection.
Hitting the latency budget is an indexing problem, not a query-engine problem. Primary records are indexed on the fields that appear in nearly every query: site, date, assignment status. AI outputs are indexed on classification value and confirmation state, with the impression text in a secondary inverted index. Because the parser already knows which index each entity belongs to, a query naming a site and a classification touches two indexes and intersects the keysets instead of scanning the worklist and filtering afterward.
The engine serves desktop and mobile unmodified, because mobile set the bar. A query that takes 800ms on a wired desktop reads as fine, and the same 800ms on a phone over cellular reads as broken. Tuning for the harsher surface gave the desktop its performance for free.
Fit Checked Before Any Build
Two conditions put a system outside this approach's envelope. Poor upstream data quality, meaning inconsistent naming, missing fields and duplicate records, produces results no parser can rescue. Unstructured content, such as freeform narrative notes, is a full-text ranking problem, and a constrained parser has no schema to map against.
Three Steps To Search Experts Trust
- Collect real queries before writing any parser. What people actually ask determines the entities and index shapes. Guessing yields a parser tuned for questions nobody has.
- Ship the filter API before the language layer. The API is a stable, testable foundation; the parser translates into it, and translations are easier to fix than foundations.
- Set the latency budget on the worst surface you support. Measure the 95th percentile on mobile at production volumes. Your experts remember the queries that time out.
Experts' Questions Set The Scope
The pattern that holds up avoids the two failure modes that kill operational search projects: over-engineering the language layer to handle questions users never ask, and under-engineering the data layer so that correctly parsed queries still time out.
Search in a live operational platform touches the record store, the AI output pipeline, the tenancy model, the auth boundary and the mobile client, and it has to be correct in all of them at once. That is the shape of the build, kept running by the person who built it, because the queries that expose a missing entity are the ones real experts type once the system is live. If your experts already describe what they want in plain language and then translate it into dropdown filters by hand, they have specified the parser for you.
References
- Dietvorst, B. J., Simmons, J. P., and Massey, C. Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err. Journal of Experimental Psychology: General, 2015.



