We closed an earlier piece on continuous monitoring with a hypothesis and an admission.[1] The hypothesis: the active ingredient in remote monitoring is not the sampling rate but the response loop, whether the signal reaches someone, or something, positioned to act before the event. The admission: the field had not run the experiment that would settle it.
That piece left a question hanging, because in July there was nothing much to answer it with. The one heart-failure telemonitoring trial that reduced mortality, TIM-HF2, achieved its result with a clinical team monitoring the data around the clock.[2] That is the response loop working. It is also a staffing model that no health system in the United States is going to reproduce across a population. So the question underneath the hypothesis was never whether the loop matters. It is whether the loop can be held by anything other than a physician who does not sleep.
A preprint posted to arXiv in March tests precisely that substitution, and it is the first piece of evidence we have seen that speaks to the automated case rather than the staffed one.[3] It deserves a careful reading, which is what follows, including the parts that are less flattering than the headline.
What the study did
The authors built an agent, Sentinel, that triages remote monitoring vital signs by pulling patient context through twenty-one clinical tools and reasoning over it in multiple steps, rather than comparing a reading against a fixed threshold. They evaluated it three ways: self-consistency across a hundred readings run five times; a comparison against rule-based thresholds; and validation against six clinicians, three physicians and three nurse practitioners, in a connected matrix design. Classification used four severity levels. The reference standard was the clinicians' majority vote, and a leave-one-out analysis compared the agent against each individual reviewer. Cases where the agent escalated well beyond the human consensus were sent to independent physician adjudication.
That design is more serious than most of what gets published about clinical AI. It has a control comparison, a human reference standard, and a deliberate attempt to interrogate its own errors.
What it found
Against the human majority-vote standard across 467 readings, the agent reached 95.8 percent sensitivity on emergency classifications and 88.5 percent on all actionable alerts, at 85.7 percent specificity. Agreement with the human standard was substantial, with a quadratic-weighted kappa of 0.778, and 95.9 percent of its classifications landed within one severity level of the humans. Self-consistency was high, kappa 0.850. Median cost was thirty-four cents per triage.
Then the number that is not in most summaries of this paper: four-level exact accuracy was 69.4 percent. The agent picked the same severity level as the human consensus in roughly seven cases out of ten. The kappa and the within-one-level figure are the honest way to read that (on an ordinal scale, being one level off is a much smaller error than being wrong), but anyone citing 95.8 percent without also citing 69.4 percent is quoting selectively, and a clinical reader will find the second number in about ninety seconds.
The genuinely important result is the shape of the errors. Disagreements skewed toward overtriage, 22.5 percent, rather than undertriage. Independent physician adjudication of the severe gaps, where the agent escalated two or more levels above the humans, validated the agent's escalation in 88 to 94 percent of cases, and consensus resolution validated it in all of them. That direction matters more than the aggregate. A triage system that errs toward escalation creates work; one that errs toward reassurance creates harm. This one errs toward work, and when experts went back over its worst-looking calls, they mostly agreed with it.
What it does not establish
It is a preprint. Posted March 2026, not yet peer reviewed. Cite it as such.
It is a component test, not an outcomes trial. The study measures whether an agent agrees with clinicians about the severity of a reading. It does not measure whether patients monitored this way are admitted less often or live longer. TELE-HF and BEAT-HF are cautionary precisely here: both delivered data to clinicians competently and neither changed outcomes.[4][5] Agreement with a clinician on a retrospective reading is a necessary condition for an automated response loop. It is nowhere near a sufficient one.
The clinician comparison is doing more rhetorical work than it can bear. The figure everyone repeats is that the agent beat every individual clinician on emergency sensitivity, 97.5 percent against a 60.0 percent aggregate. Two things are worth holding alongside it. First, comparisons of an individual against a consensus structurally disadvantage the individual, because the consensus is partly constituted by the people being measured against it; leave-one-out is the standard mitigation and the authors used it, but it does not fully dissolve the asymmetry. Second, a 60 percent figure says something about the task as much as the clinicians: reviewing an isolated reading retrospectively is not what a clinician does in practice, where they have the patient, the history, and the ability to make a phone call. The right reading is that the agent is competitive at this specific task, not that it is better at clinical judgment than physicians.
The paper does not claim federal program funding, and neither should anyone citing it. Its architecture is thematically aligned with ARPA-H's ADVOCATE program, announced in January 2026 to fund agentic AI for cardiovascular care with nationwide deployment as the stated goal.[6] That alignment is real and it tells you something about where the field is heading. It is not a funding relationship stated in the preprint, and the two should not be conflated.
Where this leaves the hard part
The paper's own conclusion is the most consequential sentence in it, and it is not about the agent. The authors note that the technical barriers to scaling their system are minimal, because it needs only FHIR-compliant access to aggregated health information exchange data, which federal interoperability rules increasingly mandate. In other words: the intelligence is now the tractable part.
If that holds, it relocates the difficulty. The obstacle to running an automated response loop in production stops being can an agent triage and becomes everything that has to be true around the agent before a health system will let it near a patient. Where the signal came from and whether its provenance survives to the record. What the agent is permitted to say, checked before it says it, against a policy someone clinically accountable wrote. What it proposed against what actually went out, recorded in a form that an auditor or a regulator can read a year later. Who holds the protected health information, and under what agreement. None of that is a modelling problem, and none of it is solved by a better agent.
We would add one caution to the paper's optimism. "Only FHIR-compliant HIE access" is a sentence that is technically accurate and operationally enormous, in the same way that "only needs a bank account" understates opening one.
Our position
We think this is the strongest published support so far for the proposition our earlier piece could only assert: that the response loop is the active ingredient and that it may not require a human at every step. We are not going to overstate it. It is one preprint, on a retrospective task, without an outcome. The honest claim it licenses is that automated triage of remote monitoring data is now a reasonable thing to test prospectively, which is a considerable upgrade from where the question stood a year ago and considerably short of proven.
What we take from it is mostly a reallocation of attention. If the agent is becoming the easy part, then the case for building carefully underneath it gets stronger, not weaker. That is the part we work on.
Sources
- AnyBio, "The case for continuous over periodic monitoring in value-based care," July 2026 - anybio.io
- Koehler F, et al. (TIM-HF2), "Efficacy of telemedical interventional management in patients with heart failure," The Lancet, 2018 - thelancet.com
- Kim S, Kung TH, Verma H, et al., "From Days to Minutes: An Autonomous AI Agent Achieves Reliable Clinical Triage in Remote Patient Monitoring," arXiv preprint 2603.09052, March 2026 (not peer reviewed) - arxiv.org
- Chaudhry SI, et al. (TELE-HF), "Telemonitoring in Patients with Heart Failure," NEJM, 2010 - nejm.org
- Ong MK, et al. (BEAT-HF), "Effectiveness of Remote Patient Monitoring After Discharge of Hospitalized Patients With Heart Failure," JAMA Internal Medicine, 2016 - jamanetwork.com
- ARPA-H, "ADVOCATE" program announcement, January 2026 - arpa-h.gov
