Yes, within what AI is equipped to assess, as long as humans set the standard and the AI applies it. Trained raters often disagree on how to score the same encounter. AI can't fix that by being smarter. It can fix it by applying one expert-defined standard the same way every time, and showing its work.
That was the answer from PCS.ai's Fall 2026 SSH Virtual Learning Lab, If Humans Can't Always Agree, Can AI? Inside AI-Powered Communication Assessment. Here's how we got there, and what to ask before you trust any AI assessment tool, including ours.
Most disagreement comes from applying criteria, not defining them. Two trained raters can share the same rubric and still score the same encounter differently, and the same rater can drift from one session to the next.
That points to a clear division of labor:
Humans aren't removed; they're moved upstream. Their judgment goes into setting the standard rather than into hundreds of individual scoring calls.
AI/Assessment 2.0 moves checklist scoring from "Did you ask?" to "Do you differentiate and understand?" Asking a question is no longer enough to earn credit; the learner has to uncover what matters.
Take a simple item: did the learner elicit the patient's medical conditions? The learner asks, "Do you have any health conditions?" The patient says, "Yes." Under keyword-style scoring, that earns the point. Under AI/Assessment 2.0, credit comes only once the learner follows up and identifies the actual conditions.
|
Assessment 1.0 |
AI/Assessment 2.0 |
|
|---|---|---|
|
Core question |
Did you ask? |
Do you differentiate and understand? |
|
Credit |
Asking the question earns the point |
The learner has to uncover the answer |
|
Weighting |
Broad, general items |
Weighted by relevance to that patient |
|
When it scores |
Question by question, shown in real time |
Full transcript, after the session |
|
Reported accuracy |
About 86% |
About 97% |
Three design choices stand out:
Assessment 1.0 is still available for programs that want simpler "did you ask" credit.
At PCS.ai, a new checklist credit goes live only after it returns the identical result 1,000 times in a row. It gets there in three steps:
Consistency comes first because it can be tested directly. A human rater drifts from session to session; a credit that passes 1,000 identical runs doesn't. Learners and faculty can also request a review of any credit, so every score can be checked.
It can when the feedback is anchored to frameworks the field already trusts, and every judgment points to evidence. Communication is as much art as science, so the bar here is transparency, not just consistency.
PCS.ai's communication feedback is built on the 15-item Observed Learner Evaluation (OLE). Gayle Gliva-McConvey, a founding member and past president of the Association of Standardized Patient Educators (ASPE), developed it by synthesizing nine established communication frameworks: Kalamazoo, Calgary-Cambridge, MIRS, SEGUE, Four Habits, GRS, MAAS-Global, Arizona CIRS and SPIKES.
No. SPs and AI-powered virtual patients are two modalities with real overlap, and the goal is to make the most of both. For communication assessment specifically, that overlap is the standard itself: both can be anchored to the same expert-defined framework for what good communication looks like. Where they differ is in what each one can observe.
|
Standardized patients bring |
AI virtual patients bring |
|---|---|
|
The human experience: "This is how that interaction made me feel" |
The learner's exact words, cited as evidence |
|
Tone of voice, eye contact and body language |
The same criteria applied across every encounter |
|
Nuanced reactions to patient cues |
Sequencing and critical-error checks |
The SP tells a learner how the communication affected the patient; the AI shows what the learner did and where the evidence is. Together, they give learners a fuller picture than either one alone.
And SP educators are well placed to lead this work. As Gayle put it, defining good communication is what SP educators already do.
Ask any vendor these five questions, PCS.ai included. They come from Balazs Moldovanyi, PCS.ai's co-founder and CEO, whose team spent nearly a decade and four full rebuilds getting AI checklist scoring right.
A polished demo can make almost any tool look like AI assessment. These questions get at what's underneath.
PCS.ai's current approach to scoring history-taking and physical-exam checklists. It scores the full transcript after the session and gives credit only when the learner uncovers the relevant information, not just asks about it.
A 15-item communication feedback framework synthesized from nine established frameworks, including Kalamazoo, Calgary-Cambridge and SPIKES. Educators select the items that match their curricular objectives.
A mistake in the order of clinical reasoning that makes an encounter unsafe, such as recommending a medication before asking about allergies. It's flagged and explained regardless of how well the rest of the encounter went.
A live score turns the encounter into point-chasing. Holding the result until the end keeps learners focused on reasoning through the case.
PCS.ai reports that moving to full-transcript scoring raised accuracy from about 86% to about 97%. Each credit must also return identical results across 1,000 test runs before it goes live.
Watch If Humans Can't Always Agree, Can AI? on YouTube, including a live demo where checklist scoring and motivational-interviewing feedback come together on one result.
For the earlier generations of this work, see the July 2025 Learning Lab, From Conversation Partner to Communication Coach.
Speakers
Want to see it with your own scenarios? Email sales@pcs.ai to arrange a custom virtual demonstration.