We won! 2026 Best Overall Patient Engagement Platform →
← Back to Articles
IndustryJuly 22, 2026

Don't Take the AI's Word for It

By Videra Health

Don't Take the AI's Word for It

AI Summary

When an AI tool flags a patient’s health risk, the clinician should be able to check its work, not just trust a label. The newest version of TDScreen, Videra Health’s tardive dyskinesia screener, now returns a predicted AIMS score, the ranked movements behind it, population benchmarks, and one-click video evidence for every finding. The AI does the watching; the clinician does the deciding.

Key Takeaways:

  • Earlier versions of TDScreen returned only a risk level (low, moderate, or high); the new version returns a predicted AIMS score in the same clinician-rated scale already used to assess these movements.
  • The score is broken down into its evidence: every detected movement is ranked and labeled by body region and severity, rather than delivered as a single black-box number.
  • Findings are benchmarked against a reference population, giving context on severity and frequency by region.
  • Clicking any flagged moment jumps the video to the exact second the model detected it, so a clinician verifies the evidence in one click instead of scrubbing a full session.
  • The underlying detection is peer-reviewed and matched or outperformed trained human raters (Journal of Clinical Psychiatry), and the design keeps every clinical decision with the clinician.

Why should a clinician trust an AI risk score?

They shouldn’t have to take its word for it. A bare risk label, low, moderate, or high, is a verdict with no evidence attached, and asking a clinician to act on a verdict they can’t inspect is exactly how good tools lose the room. The fix isn’t a more confident model. It’s a model that shows its work, so the person accountable for the patient can check it in seconds and decide for themselves.

Why healthcare is where this matters most

Explainability gets treated as a nice-to-have across most of AI. In healthcare it is closer to a precondition. A clinical model’s output is not a movie recommendation; it points a licensed clinician toward a decision about a real person, and when something goes wrong the accountability sits with that clinician, not the model. Handing a professional a conclusion with no way to interrogate it asks them to absorb a risk they can’t evaluate, and they are right to be skeptical.

That skepticism is part of why some clinical tools never earn real use. A clinician who can’t see why a model reached its conclusion has no way to separate a real signal from a spurious one, so the cautious default is to discount the whole thing. What makes a tool easy to set aside is not that it offers a recommendation, but that it offers one with no reasoning attached.

So the bar for AI in healthcare was never accuracy alone. It is whether the person held responsible can see why the model reached its conclusion and check it against reality. A model that is accurate but unexaminable still fails that test, because it cannot be safely acted on. Transparency is what turns a good score into something a clinician can actually use.

What that looks like in a tool

The newest version of TDScreen, Videra Health’s tool for detecting tardive dyskinesia, a movement disorder that can develop as a side effect of antipsychotic medications, is built to that bar. The detection underneath it is already peer-reviewed and performs on par with or better than trained clinical raters. What is new is not the accuracy; it is how much of the model’s reasoning the tool now hands back to the clinician.

From a label to a number clinicians already use

The old output was a risk level: low, moderate, high. The new output is a predicted AIMS score, in the same structured scale clinicians already use to rate these movements.

That small change does real work. Instead of translating an unfamiliar tier into their own frame of reference, a clinician reads a number that already means something, and it slots straight into how they think about severity, documentation, and follow-up.

The moments behind the score

A score on its own is still a claim, so the new version breaks it into the evidence behind it. Every movement the model detects is ranked and labeled by body region and severity, and benchmarked against a reference population, so a clinician sees not just a number but what it is made of and how it compares: which movements drove it, where they occurred, how pronounced they were, and whether that sits within the range they would expect or looks like an outlier worth a closer look. It is an itemized read to scan and question movement by movement, not a single figure to take or leave.

The evidence, one click away

This is the part that matters most. Click any flagged movement, and the video jumps to the exact second the model saw it.

TDScreen evidence view: ranked flagged movements beside the recorded exam, with a callout noting that one click on any flagged movement jumps to the exact second on video.

Picture the workflow. A clinician scanning the findings taps a flagged tongue movement, lands on that moment in the recording, and in a second or two confirms it, or decides a borderline flag isn’t clinically meaningful and moves on. No scrubbing a full session to find the frame in question. Verification stops being a chore and becomes something a clinician actually does, because a score you can check against the source in seconds is a different thing entirely from one you have to take on faith.

The AI watches, the clinician decides

Put the pieces together and the division of labor is clear. The AI does the watching, the tedious work of reviewing footage and flagging every movement. The clinician does the deciding, with the evidence laid out to confirm or overrule in a fraction of the time a manual review would take. None of this makes the tool the decision-maker, and it is not meant to be: TDScreen is a screener, not a diagnosis, and the predicted score is there for the clinician to confirm, with the evidence one click away whenever they want to check it. The timing helps: as of January 2026, CMS recommends tardive dyskinesia screening for every patient on antipsychotics, putting more of these assessments in front of more clinicians who need them to be both fast and defensible.

What behavioral health AI should look like

The broader point outlasts any single feature. Too much health AI is built to be believed. The better bar is built to be checked. It should be useful enough to save a clinician time, and transparent enough that they never have to take its word for anything.

That is the standard the new TDScreen is built to, and the one we think the rest of behavioral health AI should be held to. It is one example of what that looks like in practice; you can see how it works.

Useful enough to save time, transparent enough to trust. That is the bar.

See how TDScreen shows its work