Digital Health: From Telemedicine to AI Diagnostics | LeverVenture
Model accuracy is table stakes. AI diagnostic winners are decided by prospective validation, workflow integration, and who bears liability for a miss.
In this note06 · 6 min
An algorithm that reads a chest X-ray as well as a board-certified radiologist is not, by itself, a business. The peer-reviewed literature is full of models that match or beat human readers on a held-out test set and never reach a single hospital's imaging queue. What separates a published result from a deployed product is not accuracy. It is evidence generated the way clinicians and regulators actually require it, and a workflow a radiologist will use on a Tuesday afternoon with forty studies left in the queue.
Growth equity investors underwriting AI diagnostics should treat model performance as necessary and close to table stakes. The decision that determines winners happens downstream of the paper: prospective validation across sites the model has never seen, integration into the reading workflow without adding friction, and a clear answer to who is accountable when the software is wrong. Distribution and evidence separate durable companies from demonstrations. Model quality, past a reasonable bar, does not.
This matters for how a diligence process should be built. A demo that performs well on a vendor's own test set answers almost none of the questions that determine whether a health system will keep paying for the tool in year three.
01Retrospective Accuracy Is Not Clinical Evidence
Most AI diagnostic models are first validated retrospectively, on a curated dataset assembled after the fact, often from a small number of academic centers with high-quality imaging equipment and consistent protocols. That kind of validation is a reasonable first gate, but it systematically overstates real-world performance. Image acquisition varies by scanner vendor and site protocol. Patient populations vary by referral pattern and disease prevalence. A model tuned to one health system's case mix degrades when it meets a community hospital's actual distribution of studies, and the size of that degradation is exactly the number a retrospective study cannot show.
Prospective, multi-site validation, run on consecutive real-world cases rather than a curated set, is a materially harder and more expensive thing to generate. It is also the thing that predicts whether a health system's own radiologists will trust the output enough to change how they read a study. Diligence should ask a specific question: how many of the sites in the validation data were also sites where the company later sold the product, versus genuinely independent sites added after the model was locked.
02The Workflow Is the Product
A tool that requires a radiologist to leave the PACS viewer, open a separate application, and re-enter patient context will lose to a materially worse model that surfaces its output inside the existing reading environment. Radiologists read hundreds of studies a week under time pressure; anything that adds clicks gets turned off within a month, regardless of what the accuracy statistics say. The companies that have built durable radiology relationships did it by shipping DICOM-native integration, worklist prioritization that does not require a new login, and outputs that look like an overlay on the existing image rather than a separate report to reconcile.
Alert fatigue is the underappreciated failure mode. A model tuned for sensitivity flags more studies as abnormal, which looks good in a marketing deck and terrible in six months of actual use, once radiologists learn to discount its flags. The right diligence question is not "what is the sensitivity and specificity," but "what is the flag rate radiologists actually act on, twelve months after go-live, at a site that is not a reference customer."
03Comparing Deployment Models
Not every AI diagnostic tool asks the same thing of the clinician or carries the same evidentiary bar. The three broad deployment postures in use today require different validation depth and carry different liability exposure.
| Deployment model | Clinician role | Evidentiary bar | Liability exposure |
|---|---|---|---|
| Concurrent second read | Radiologist reads first; software flags discrepancies for review | Moderate — output is advisory, not determinative | Lower — physician remains sole author of the report |
| Worklist triage/prioritization | Radiologist still reads every study, but order changes | Moderate — errors delay rather than replace judgment | Lower, with exposure if a deprioritized case is a genuine miss |
| Autonomous interpretation | No human review of the software's primary output | Highest — treated closer to a standalone diagnostic device | Highest — accountability shifts toward the manufacturer |
04Liability Follows Labeling, Not the Marketing Deck
Under the current U.S. regulatory and legal framework, a tool cleared as an assistive, physician-facing device generally does not relieve the reading physician of the standard malpractice duty of care. The radiologist who signs the report remains the accountable clinical decision-maker in most jurisdictions, regardless of how the software performed. That allocation of liability is precisely why hospital systems have been willing to adopt assistive tools faster than autonomous ones: the legal exposure profile for the institution barely changes.
What has changed is contract language. Vendor agreements increasingly carry indemnification provisions addressing what happens if a marketed accuracy claim proves wrong in a specific case, and malpractice insurers are starting to ask hospital risk committees whether an AI-assisted miss changes coverage terms. For an investor, the liability question worth asking in diligence is not abstract: it is whether the company's standard contract shifts any risk back onto the health system, and whether that term has ever actually been negotiated away by a sophisticated buyer.
05The Change Control Plan Changes What "Cleared" Means
In December 2024, the FDA published final guidance, "Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions," allowing manufacturers to pre-specify how an AI-enabled device may be modified after clearance without filing a new submission for every update. That is a structural change to the economics of maintaining a diagnostic AI product. Historically, retraining a model on new data risked triggering a new regulatory filing; a well-constructed change control plan lets a company improve the model on a defined cadence, inside pre-agreed guardrails, without resetting the clock.
The company that wins is not the one with the best model on the day of clearance. It is the one that can prove, five years later, that the model deployed today still performs the way its FDA-reviewed change control plan said it would.
This raises the bar for diligence in a useful way. A company with a filed, FDA-accepted change control plan has effectively pre-negotiated its path to model improvement; a company without one is exposed to a slower, more expensive resubmission cycle every time the underlying data drifts, which it eventually will.
06What This Means for Diligence
Model quality is a screening filter, not a differentiator. The diligence questions that actually separate durable companies from well-funded demonstrations are about evidence and distribution: how the validation data was assembled and by whom, what the twelve-month utilization rate looks like at a non-reference site, how liability is allocated in the standard contract, and whether the company has a change control plan on file rather than a promise to build one. A related discipline applies to any AI-enabled device entering U.S. clinical workflows, not diagnostics alone; see the broader regulatory posture in the FDA's evolving digital health guidance. The infrastructure layer supporting these tools — compute, data pipelines, and model governance — is covered separately in our view on AI infrastructure for life sciences founders, and the adjacent question of how AI changes precision oncology specifically is addressed in our analysis of precision medicine in cancer care. None of that is a substitute for reading the validation data and the contract terms directly. Both are where the real answer lives.
Nothing in this piece is investment, legal, tax or accounting advice, and nothing in it is an offer to sell or a solicitation of an offer to buy any security.

