This is the third and final installment of a three-part series on AI, virtual patients, and the evolving role of Standardized Patient Educators — a series that, admittedly, I never intended to write.
I set out to write one piece, ONE article reviewing a peer-reviewed study of ChatGPT as an SP — where the model held up, where it didn't, and how PCS.ai had already run into (and solved) most of the same problems the study documented.
But the more I sat with it, the more the premise itself started to bother me.
Not because researchers were evaluating AI as a standardized patient. That research absolutely needs to happen. What bothered me was that a general-purpose language model was being asked to stand in for an entire category of technology that has been purpose-built for this work.
Because the distinction matters.
If you ask ChatGPT to portray a standardized patient and it struggles with consistency, role portrayal, appropriate disclosure, clinical boundaries, or assessment, the easiest conclusion is that AI isn't very good at being a standardized patient. For a profession (SP Educators) already understandably wary of what AI-powered virtual patients might mean for human SP programs, findings like that can harden the divide: human SPs on one side, AI on the other.
But that's the wrong comparison.
ChatGPT isn't a virtual patient or simulation platform. It wasn't designed around the requirements of healthcare simulation, and its limitations shouldn't define what purpose-built AI patients are capable of any more than the limitations of a generic video call should define what a simulation management platform can do.
And that's what kept gnawing at me. We risk taking the shortcomings of the wrong tool and turning them into conclusions about the technology itself — further polarizing a profession (SP Educators) that should have an enormous role in shaping where this technology goes next.
So one piece became three. Part I made the case that virtual patients should have been embraced as continuity, not competition — a natural extension of what standardized patient programs already do, not a replacement for it. Part II circled back to the study itself, using it to show why "purpose-built" beats "general-purpose" on the things that actually matter in a standardized clinical encounter — prompt engineering, verisimilitude, alignment, trust, feedback.
But making those arguments left me with what, I felt, was an obligation of my own.
If I'm going to say that SP Educators shouldn't be defending their work from AI but helping shape what AI-powered patient simulation becomes, then I need to be able to show what that actually looks like. "AI needs SP Educators" sounds good, but it's vague enough to agree with and do absolutely nothing about.
So — this part — Part III gets specific.
There are three places inside the PCS.ai platform today where I believe SP Educator expertise belongs — not hypothetically, not "in spirit," and not someday when AI gets better. In the actual work of building, testing, and assessing an AI-powered patient.
Consider these a zero-entry point: three places where an SP Educator could walk into the software today, bring the expertise they already have, and immediately make the simulation better.
One of the many takeaways from Cross's paper was how much work it took to get ChatGPT to consistently portray a single patient. Before students ever interacted with the AI, a six-member faculty team iterated on prompts, refined ‘illness scripts’, and adjusted behavioral guardrails until the model behaved appropriately.
That process looks very different on PCS.ai. Instead of writing one enormous prompt, authors build what we call a Patient Concept — a structured clinical profile across fifteen fields: chief complaint, HPI, PMH, and more. And those fifteen fields map cleanly onto Domain 2 of ASPE's Standards of Best Practice — case components like history, affect and demeanor, signs and symptoms to simulate, cues.
What is even further into SP territory is Domain 3: role portrayal. Deciding what a patient volunteers versus withholds, calibrating affect and demeanor so it's consistent and believable, ensuring one learner doesn't get a fundamentally different patient than another — that's the judgment SP Educators have trained SPs on for decades. It's not a categorization skill. Its role portrayal, and authoring a PCS.AI Patient Concept requires exactly that same expertise, just pointed at a different kind of patient.
Because the structure is only as good as what goes into it. Getting the clinically relevant information into those fields — the right symptom clusters, the right level of ambiguity, the details a patient should volunteer versus the ones that should require a well-asked question to surface — is the same skill SP Educators have spent years applying to case writing (or co-authoring alongside faculty) and training for human standardized patients. Authoring an AI portrayal isn't a different discipline from authoring an SP case. It's the same discipline, applied to a new medium.
Who is better suited to build those patients than the people who have spent years training human standardized patients? SP Educators already know what information a patient should volunteer, what should remain hidden until asked, how emotion influences disclosure, and where a scenario should challenge a learner's clinical reasoning. Those are precisely the decisions that shape a believable AI patient.
They aren't prompt-writing skills. They're SP Educator skills.
Creating a patient is only half the work. Once a scenario is built, someone has to interview it — probing for the places where the portrayal is inconsistent, too eager to volunteer information, or not standardized enough across repeated runs.
The question is one SP Educators have asked for decades: "Would a real patient respond this way?"
That requires testing. Running interviews. Trying unexpected questions (every med and nursing student's favorite move right after learning a new zebra diagnosis). Looking for inconsistencies. Ensuring one learner doesn't receive fundamentally different information than another. In other words, standardizing the patient.
That's precisely the job SP Educators already do when they train and calibrate a new human SP: run the case, find where it wobbles, adjust, run it again. Doing that with a digital patient instead of a person doesn't require a new skill set. It requires the same expert clinical interviewing eye, pointed at a different kind of patient.
No one is better prepared to perform that quality assurance than experienced SP Educators. They know when a patient's affect feels wrong. They recognize when an ‘illness script’ drifts. They understand how much information a patient should — and shouldn't — offer without prompting.
The technology generates the conversation. SP Educators ensure the conversation is educationally sound.
Perhaps the opportunity that excites me most comes after the simulation ends.
On PCS.ai, a learner can request review of a specific checklist item on their assessment — did I actually ask about that, did that response count? Someone has to adjudicate that dispute, and those reviews require human judgment.
Who better to provide it than SP Educators? They're trained evaluators who've spent years assessing communication, observing learner performance, and applying standardized scoring criteria. Helping faculty work through a contested credit is a natural extension of work they already do when a student pushes back on a human-SP-based assessment.
As AI becomes increasingly capable of generating first-pass assessments, the need for expert human reviewers doesn't disappear — it becomes more important. Rather than replacing evaluators, AI lets them spend their time where expertise matters most: interpreting nuance, resolving ambiguity, and continuously improving the assessment process.
None of these three are "AI oversight" in the babysitting sense. They're places where the software is explicitly built to depend on the same expertise SP programs were built around: knowing what a good clinical case looks like, knowing how to calibrate a portrayal, knowing how to judge whether a piece of communication actually happened the way a learner claims it did.
If AI-powered patients become commonplace — and I believe they will — the question isn't whether SP Educators still have a role. The question is whether they'll recognize that their role has expanded.
Back in Part I, I mentioned hoping to return to ASPE — not just to attend, but to host the kind of standing-room-only user meetings Balázs Moldoványi and I used to run in the WebSP / LearningSpace days, listening to our clients, gathering workflow feedback and feature requests. This is what that conversation would actually be about: not whether SP Educators belong in this future, but which of these three seats they want to see software developments or improvements in first.
The future of communication training doesn't need fewer SP Educators. It needs their expertise embedded throughout the entire lifecycle of AI-powered simulation — from designing patients, to validating encounters, to ensuring assessments remain fair, meaningful, and educational. The rise of AI in standardized patient work was never a threat to the programs that pioneered simulated clinical encounters. It's an expansion of the same mission, built on the same expertise — just with a few new rooms in the building that need someone who already knows how to do this work.
That's not replacement.
That's leadership.
*All em-dashes in this article were human-generated.