Upcoming Webinar: “Nick’s Lunch & Learn: Spark, The Virtual Patient” — calculating…

Register
Products
PCS Platform
Back to results

Can AI Assess Clinical Communication as Reliably as Human Raters?

Author photo
Michelle Castleberry Oct 6, 2026, 11:04:57 AM

Yes, within what AI is equipped to assess, as long as humans set the standard and the AI applies it. Trained raters often disagree on how to score the same encounter. AI can't fix that by being smarter. It can fix it by applying one expert-defined standard the same way every time, and showing its work.

That was the answer from PCS.ai's Fall 2026 SSH Virtual Learning Lab, If Humans Can't Always Agree, Can AI? Inside AI-Powered Communication Assessment. Here's how we got there, and what to ask before you trust any AI assessment tool, including ours.

Why do human raters disagree?

Most disagreement comes from applying criteria, not defining them. Two trained raters can share the same rubric and still score the same encounter differently, and the same rater can drift from one session to the next.

That points to a clear division of labor:

  • Humans set the standard. Experts define what good looks like, in writing, anchored to published frameworks.
  • AI applies it. The AI scores every encounter against that same definition, the same way, every time.

Humans aren't removed; they're moved upstream. Their judgment goes into setting the standard rather than into hundreds of individual scoring calls.


How does AI score a history-taking checklist?

AI/Assessment 2.0 moves checklist scoring from "Did you ask?" to "Do you differentiate and understand?" Asking a question is no longer enough to earn credit; the learner has to uncover what matters.

Take a simple item: did the learner elicit the patient's medical conditions? The learner asks, "Do you have any health conditions?" The patient says, "Yes." Under keyword-style scoring, that earns the point. Under AI/Assessment 2.0, credit comes only once the learner follows up and identifies the actual conditions.

 

Assessment 1.0

AI/Assessment 2.0

Core question

Did you ask?

Do you differentiate and understand?

Credit

Asking the question earns the point

The learner has to uncover the answer

Weighting

Broad, general items

Weighted by relevance to that patient

When it scores

Question by question, shown in real time

Full transcript, after the session

Reported accuracy

About 86%

About 97%

Three design choices stand out:

  • Scoring grows with the learner. For beginners, credit can be a simple yes or no; for advanced learners, it expects deeper reasoning. Think multiplication facts before algebra.
  • No live score. Holding the result until the end means learners reason through the case instead of chasing points.
  • Shorter can score higher. A routine, script-style interview that scored 100% under Assessment 1.0 can score lower under 2.0, while a shorter, more targeted one scores higher. The score rewards clinical reasoning, not box-checking.

Assessment 1.0 is still available for programs that want simpler "did you ask" credit.


How is an AI scoring item validated?

At PCS.ai, a new checklist credit goes live only after it returns the identical result 1,000 times in a row. It gets there in three steps:

  1. Humans score it first. A group of people score real sessions independently and agree on the right call, across multiple rounds.
  2. An independent AI checks each call. A second AI review scores the same sessions and explains its reasoning. Human and AI have to agree.
  3. It runs 1,000 times. The credit is tested against those sessions 1,000 times. If the answer isn't identical every time, it gets rewritten.

Consistency comes first because it can be tested directly. A human rater drifts from session to session; a credit that passes 1,000 identical runs doesn't. Learners and faculty can also request a review of any credit, so every score can be checked.


Can AI give credible communication feedback?

It can when the feedback is anchored to frameworks the field already trusts, and every judgment points to evidence. Communication is as much art as science, so the bar here is transparency, not just consistency.

PCS.ai's communication feedback is built on the 15-item Observed Learner Evaluation (OLE). Gayle Gliva-McConvey, a founding member and past president of the Association of Standardized Patient Educators (ASPE), developed it by synthesizing nine established communication frameworks: Kalamazoo, Calgary-Cambridge, MIRS, SEGUE, Four Habits, GRS, MAAS-Global, Arizona CIRS and SPIKES.

  • Evidence behind every rating. Each item is rated demonstrated, partially demonstrated or not demonstrated, citing the learner's own words, with a rationale and coaching on how to improve.
  • Educators choose the focus. No one gives 15 items of feedback in one encounter. PCS.ai recommends six or seven, picked to match your objectives, from domains such as information gathering, relationship building, challenging conversations and motivational interviewing.
  • Critical errors and sequencing. Beyond whether a topic was covered, the AI checks whether clinical reasoning happened in a safe order. An unsafe sequence is flagged as a critical error, not just a lost point.

Does AI replace Standardized Patients?

No. SPs and AI-powered virtual patients are two modalities with real overlap, and the goal is to make the most of both. For communication assessment specifically, that overlap is the standard itself: both can be anchored to the same expert-defined framework for what good communication looks like. Where they differ is in what each one can observe.

Standardized patients bring

AI virtual patients bring

The human experience: "This is how that interaction made me feel"

The learner's exact words, cited as evidence

Tone of voice, eye contact and body language

The same criteria applied across every encounter

Nuanced reactions to patient cues

Sequencing and critical-error checks

The SP tells a learner how the communication affected the patient; the AI shows what the learner did and where the evidence is. Together, they give learners a fuller picture than either one alone.

And SP educators are well placed to lead this work. As Gayle put it, defining good communication is what SP educators already do.


How should you evaluate an AI assessment tool?

Ask any vendor these five questions, PCS.ai included. They come from Balazs Moldovanyi, PCS.ai's co-founder and CEO, whose team spent nearly a decade and four full rebuilds getting AI checklist scoring right.

  1. Is it anchored to a validated framework, or did the vendor define "good" itself?
  2. Is the accuracy data a trend over time, or a single number?
  3. Can a learner or faculty member check or challenge a score?
  4. How many times has the vendor rebuilt it, and will they say what didn't work?
  5. Who defined "good": subject-matter experts, or engineering alone?

A polished demo can make almost any tool look like AI assessment. These questions get at what's underneath.


FAQ

What is PCS.AI AI/Assessment 2.0?

PCS.ai's current approach to scoring history-taking and physical-exam checklists. It scores the full transcript after the session and gives credit only when the learner uncovers the relevant information, not just asks about it.

What is the Observed Learner Evaluation (OLE)?

A 15-item communication feedback framework synthesized from nine established frameworks, including Kalamazoo, Calgary-Cambridge and SPIKES. Educators select the items that match their curricular objectives.

What is a critical error in AI assessment?

A mistake in the order of clinical reasoning that makes an encounter unsafe, such as recommending a medication before asking about allergies. It's flagged and explained regardless of how well the rest of the encounter went.

Why don't learners see their score during the session?

A live score turns the encounter into point-chasing. Holding the result until the end keeps learners focused on reasoning through the case.

How accurate is AI checklist scoring?

PCS.ai reports that moving to full-transcript scoring raised accuracy from about 86% to about 97%. Each credit must also return identical results across 1,000 test runs before it goes live.


Watch the full session

Watch If Humans Can't Always Agree, Can AI? on YouTube, including a live demo where checklist scoring and motivational-interviewing feedback come together on one result.

For the earlier generations of this work, see the July 2025 Learning Lab, From Conversation Partner to Communication Coach.

 

Speakers

  • Judit Vigh-Kondor, Director of the AI/Clinical Scenarios & Assessment, PCS.ai | LinkedIn
  • Gayle Gliva-McConvey, Director of Clinical Communication Education & AI/Scenario Design, PCS.ai; founding member and past president of ASPE | LinkedIn
  • Balazs Moldovanyi, Co-founder and CEO, PCS.ai | LinkedIn
  • Michelle Castleberry, Co-founder, PCS.ai (moderator) | LinkedIn

Want to see it with your own scenarios? Email sales@pcs.ai to arrange a custom virtual demonstration.

Share this post
Subscribe to our newsletter
Stay in the know with our exclusive newsletter—delivering the latest updates, tips, and insights straight to your inbox.
By clicking Sign Up you're confirming that you agree with our Terms and Conditions.