A former president of the Association of Standardized Patient Educators (ASPE) recently emailed me a paper titled Assessing ChatGPT’s Capability as a New Age Standardized Patient: Qualitative Study. In the body of the email, the past-president wrote, “We knew this – but it’s good to see it in print.” She shared it with me because we align on this issue: we are firmly PRO-AI as an adjunct to traditional standardized patients – and equally firm in our stance against SP elimination.
This is this the second article in a 3 part series – Part II: "We Knew This". [Read Part I]
We also share something of an insider's perspective on the paper's findings. Over the past decade at PCS.ai — where I was employee number five and she served on our Clinical Advisory Board — we encountered these same themes in real time: technology constraints, measurable learning gains, realism gaps, and the persistent question of trust. We lived it for years — are still living it — so, for us, the study doesn't introduce a new tension; it validates one we've been navigating in practice. So, while none of it surprised us, seeing it articulated through formal research mattered.
It also underscores a trend we've observed over the past 18+ months: the increasing use of general-purpose large language models (like ChatGPT) as improvised standardized patients. The motivation is understandable. As the paper confirms, SP programs are resource-intensive and difficult to scale. In regions with constrained clinical access — such as parts of the Caribbean — those constraints are even more pronounced.
The drivers are structural, not ideological. But I worry about this trend of using general-purpose large language models. The author of the paper, Dr. Cross, highlights where general-purpose LLMs (ChatGPT) struggle when using it as an SP:
Prompt engineering, for one. Before a single student ever spoke to ChatGPT, a six-member faculty team spent an iterative trial-and-error process just getting the prompt to behave — refining phrasing, tightening the ‘illness script’, adding behavioral guardrails so the model wouldn't wander out of character. And even that carefully engineered prompt needed revising mid-study, because ChatGPT's first instinct was to hand out uniformly glowing feedback regardless of how the encounter actually went. If it takes a faculty committee and multiple rounds of trial and error to get a general-purpose chatbot to convincingly hold still inside a clinical role, that's not a footnote. That's the finding.
Then there's alignment — not in the abstract, philosophical sense, but in the very concrete sense of a model declining to answer when a student asked about sexual history, a completely standard line of questioning in any real clinical intake. Cross calls this the "alignment tax": the guardrails that make ChatGPT broadly safe for a billion different consumer use cases are the same guardrails that make it a worse standardized patient, because a clinical training encounter has different needs than a general-purpose chatbot fielding questions from the open internet.
Verisimilitude took the hit you'd expect. No face. No tone. No shift in the voice when the "patient" is describing something painful or frightening. One student put it plainly: "ChatGPT had the same tone, even if it was saying something sad." Others noted the model tended to over-share relative to what a real patient — or a well-trained SP — would offer unprompted, creating an information overload that doesn't match how actual clinical encounters unfold. A real patient doesn't monologue their entire differential. They wait to be asked.
And trust, while it faded as a concern by the end of the study, is worth sitting with rather than waving off. The faculty team didn't catch a hallucination during their sessions — that's good news. But Cross makes a sharper point that reaches past this one classroom: a single bad answer from a model like this doesn't stay contained to one training exercise. If the same flawed output were to reach multiple healthcare providers, as it plausibly could once tools like this are deployed at scale, that error doesn't just affect one student having a bad day. It affects patients. Scale cuts both directions.
Again, none of this is a surprise to anyone (e.g., me, or us at PCS.AI) who has actually built one of these platforms instead of just prompting one. And that, I think, is the real value of this paper — not that it proves ChatGPT is bad, but that it accidentally draws a clean line between what a general-purpose LLM does and what a purpose-built virtual patient or virtual simulation platform is for.
A general-purpose chatbot has no ‘illness script’ unless someone writes one into the prompt, no memory of what "in character" means beyond what that prompt manages to hold onto, no alignment tuned for clinical education instead of consumer safety, and no built-in mechanism for objective, structured feedback beyond whatever the prompt happens to ask for. Every "technology limitation" Cross's team documented is, functionally, a description of the infrastructure a dedicated platform exists to provide: curated cases instead of ad hoc scripts, guardrails calibrated for the clinical context instead of borrowed from a general-purpose safety policy, and a feedback framework built by educators rather than reverse-engineered by whoever happens to be prompting that day.
Table 4 of the paper says this almost by accident, in a single row: ChatGPT and Claude get credit for flexibility and unlimited practice, but are dinged for "uncurated outputs" and reliance on prompt engineering — while the purpose-built platforms in the comparison (Body Interact, Oscer AI, Soma Lab*) are credited with curated clinical cases instead, with Soma Lab* specifically singled out for tailored feedback. That's not a knock against ChatGPT. It's an accurate description of what it is: a general-purpose tool being asked to do a specialized job. It can do it. It just does it the way a wrench can do a hammer's job — passably, and with real limitations you wouldn't accept in the tool actually built for the task.
I've spent most of this piece describing what general-purpose LLMs get wrong as standardized patients. It's worth being concrete about what a purpose-built platform gets right — so I'll walk through Cross's own categories, applied to the PCS.ai platform — our / PCS.ai’s virtual simulation cloud. What learners actually see and talk to is the digital, virtual patient, delivered through two modalities: Spark and SimVox Ultra. Cross's team didn't include either in their Table 4 comparison, so this is my analysis, not theirs — and frankly, I'm hurt we weren’t included! — but the categories are borrowed directly from the paper.
On prompt engineering and uncurated outputs: nobody on the PCS.ai platform is starting from a blank prompt, or even a paragraph. AI/Scenarios builds a case from a structured, fifteen-field authoring template — reason for visit, main concern and associated symptoms, allergies and medications, chronic illness and medical and family history, relationship status and sexual health, social and substance history, lifestyle, emotional state, even conversational attitude and speech style — rather than asking a faculty committee to hand-write (likely type) and iterate on a single illness script buried inside a prompt. We call this authoring template the “Patient Concept”. The PCS.ai platform also ships with more than 50 pre-authored, curated scenarios spanning primary care, acute presentations, and difficult-news conversations. These can be copied, modified, or created anew. The six-person, trial-and-error process Cross's team went through to get ChatGPT to hold its role is, on the PCS.ai platform, closer to a starting configuration than a prerequisite.
On verisimilitude — the theme every single student in Cross's study raised, in both rounds of interviews: the digital patient, whether a learner is meeting it through Spark or through a SimVox Ultra display in the room, has an avatar, not just text. The voice synthesis supports emotional tone — sad, frightened, cheerful, angry — so a patient describing something painful doesn't sound identical to one describing something routine, which was the exact gap students named (remember the "ChatGPT had the same tone, even if it was saying something sad" from earlier?). Spark and SimVox Ultra display also supports a physical exam: learners select and place instruments on the patient's body — hearing and seeing normal or abnormal results — not just ask questions into a chat window. None of that closes the verisimilitude gap entirely — no digital patient fully replaces a human one, which is the whole point of this series — but it's a direct answer to the specific limitation Cross documented, not a workaround for it.
On trust and feedback: Cross's team had to revise their ChatGPT prompt mid-study because it defaulted to uniformly positive feedback with no diagnostic value. PCS.ai's assessment splits into two structured outputs, not one improvised one. Quantitative feedback is checklist-style scoring built from scenario-specific credit criteria, tested and refined against manually scored transcripts to bring machine scoring in line with human scoring. Qualitative feedback is a narrative writeup — general communication skills, motivational interviewing feedback, strengths and room for improvement, with specific examples cited from the encounter — grounded in recognized communication frameworks rather than generated on the fly. That's a structural difference, not just a better-tuned prompt: the evaluation criteria and the framework live in the platform, not in whatever the last person to edit the prompt happened to ask for.
On alignment — Cross's finding that ChatGPT declined to engage with a standard question about sexual history because of guardrails built for a billion different consumer use cases: the PCS.ai platform doesn't have that problem, on Spark or on SimVox. Both run on the same trained model. We trained it on tens of thousands of examples accumulated over the past decade of our pre-LLM AI work, fine-tuned specifically to perform within the defined context of a clinical encounter, the context of HEALTHCARE SIMULATION — sensitive questions included. It was built for this conversation, not adapted to allow it.
This is the trend I mentioned earlier, and the one that worries me: institutions with real constraints — funding, faculty time, clinical access — reaching for the general-purpose AI over purpose-built virtual patients and virtual simulation platforms because the former is already sitting open in a browser tab, free and available, rather than investing in a platform built to solve the actual problem. I get it. I understand the impulse completely. But "good enough and immediately available" is a different thing than "built for this," and Cross's data quietly proves it.
That's the tension underneath all of this. Cross's argument was that general-purpose AI can function as a standardized patient, limitations included — and the data backs that up. The argument the paper wasn't trying to make, but ends up making anyway, is a different one: those same caveats — prompt engineering, alignment, verisimilitude, trust — read less like limitations and more like a requirements list, one that purpose-built platforms like pcs.ai already exist to fulfill. Cross wasn't building a case for purpose-built AI. He built it anyway, one limitation at a time.
Part III gets specific: three places in the PCS.ai platform where SP educator expertise can serve as an entry point — a real on-ramp for educators who were once hesitant, skeptical, or even resistant to AI, and now want to step in but aren't sure how, or where.
Read Part I: The Rise of AI Isn’t a Threat to SP Programs — Resistance Might Be.