October 6, 2026

How Realistic Is an AI HCP? What We Learned Testing reTrain's Personas

By Ahmet Fatih Şahin, Product Owner, reprai

How Realistic Is an AI HCP? What We Learned Testing reTrain's Personas

Early in reTrain's development, one of our personas went wrong, and it didn't look like a failure.

It never broke character. It never announced it was an AI. It drifted. It became more helpful than the situation called for. It started explaining the product back to the rep. It gave answers far longer than an HCP with a full waiting room would ever give.

Each reply sounded reasonable on its own. Together they made a rehearsal that felt realistic and prepared no one for anything.

My PhD was in computer-aided drug design. The first lesson there is that a model that looks right is where the checking starts, not where it ends. Building AI roleplay for pharma teams, that lesson turned out to be most of the job.

What does "realistic" mean for an AI HCP?

Two different properties get called "realistic," and they are easy to confuse.

The first is fidelity: does the persona behave the way a real HCP, KOL, or access committee would? The second is transfer: does a rep who does well in rehearsal do better in the real visit?

FidelityTransfer
The questionDoes the persona behave like a real HCP?Does rehearsal performance carry into the field?
Whose propertyThe system'sThe world's
How you test itInside the product: character, boundaries, scoringOutside the product: field outcomes, comparison groups, time

What the research says about AI virtual patients

Medical education has been asking this question longer than pharma has.

A 2026 systematic review in JMIR Medical Informatics (Li and Lutfi) covered 39 studies of large language model-based virtual patients. On capability, the systems performed well. Methodologically, the field is still early: cohorts typically ran between ten and fifty students, only a quarter to a third of studies included a controlled comparison, and results varied too much to pool into a meta-analysis.

The strongest work shows what careful evaluation looks like. AIPatient, published in Communications Medicine in December 2025 (Yu and colleagues), tested an agent system across five dimensions and ran a blinded comparison against human simulated patients. The authors were explicit about scope: the simulation is realistic and stable. That is a claim about fidelity, not transfer.

If peer-reviewed clinical education can't yet show transfer, a percentage on a sales slide isn't proof of it either. That includes ours.

How we test reTrain's personas

reTrain is reprai's training product for sales, medical, and market access teams. Team members rehearse real conversations against AI personas built for the conversations they are about to have, then get scored and coached. Here is how we handle fidelity inside it.

Character consistency is a precondition, not a feature

Quote card: Character consistency is a precondition, not a feature. Ahmet Fatih Şahin, Product Owner, reprai

The drift I described above was an early-stage problem. It isn't possible in the product today.

We didn't treat it as a tuning problem. We treated it as an architecture problem and solved it there. The role's boundaries are no longer left to the flow of the conversation. Before every roleplay starts, the system defines who the persona is, what it can talk about, and where it stops. A guardrail layer enforces that definition for the whole conversation. The closed-loop structure leaves the persona no path out of its role.

In reTrain, staying in character is not something the model happens to do in the moment. It is a constraint set when the roleplay is built. In a simulation product, anything else is unacceptable.

Difficulty is an input, not a flaw

"Is the persona too easy or too hard?" is the wrong first question. Difficulty is set by the conversation the company wants its team to rehearse.

A team presenting a new indication for the first time shouldn't face the same persona as a team meeting an HCP who has already committed to another option. We build personas from two sources: the pharma company's own know-how and our accumulated experience. The system also suggests persona adjustments from how it is used, so difficulty isn't a one-time guess.

What we check is where the difficulty comes from. A persona that is hard because it is rude isn't hard. It's unpleasant, and it teaches nothing that transfers. A well-calibrated persona creates friction where real visits get hard: very little time, a decision already made for another option, and an evidence question that needs a precise answer, not an enthusiastic one.

The score is written before the conversation starts

This was our most important design decision. The standard doesn't form after the conversation ends. It is written before it begins.

For every roleplay, the system defines what should be discussed, how the conversation would move in real life, and what tone the rep and the HCP should take. The score is produced against that definition, not against a general idea of "a good conversation."

A real visit is more than transferring the right information, so the measurement isn't limited to words and metrics. A rep's performance forms in three places at once:

  1. Conversation management. Can they set up the opening, hold the sequence, and carry the close?
  2. Word choice. Do their words fit this conversation, inside both the tone and the boundaries?
  3. Delivery. How it is said, pronunciation included.

A sentence with the right content, delivered in the wrong order, in the wrong tone, or in a way that won't be understood, doesn't work in the field. It shouldn't get full marks in reTrain either.

The closed loop exists for exactly this reason. The part that runs the conversation and the part that scores it know the same expected flow and the same expected tone. So the score doesn't answer "how did it sound?" It answers "where, and how far, did this deviate from the conversation that should have happened?"

Diagram of the reTrain closed loop: one expected conversation, defined before the roleplay, feeds both the persona that runs the roleplay and the score that evaluates the rep

Compliance is a boundary, not a test case

We don't "test" whether a persona can be pushed off-label. Leaving the role, off-label discussion, and off-indication discussion are closed at the system level. General compliance rules are defined, and on top of them, each roleplay gets boundaries specific to that conversation.

The deciding choice sits upstream of the model. The single source of truth is the approved label for the relevant market. Not an internal summary. Not another market's version. That matters more than it sounds, because if the source document is wrong, everything downstream is wrong with complete confidence.

In a simulation, compliance can't be something you audit in the output. It has to be defined where the conversation is built.

A reTrain practice session with the AI HCP persona Dr. Erin Caldwell

What we can say about the field today, and what we can't

Everything above is fidelity. Transfer is the harder question.

Today our signal comes from use. Active tenants keep using reTrain, teams are satisfied, and customers tell us they see measurable effects in the field. For a product, that is a serious signal, and I don't dismiss it.

But it is the customer's own observation. Turning it into evidence takes three things:

  1. A field outcome measured independently of the training system.
  2. A comparison between teams that rehearse and teams that don't.
  3. Enough time between rehearsal and measurement to show lasting capability, not fresh memory.

We are working on all three, and we have initial outputs. That is exactly where we want to publish.

Four questions to ask any AI roleplay platform

If you are evaluating AI roleplay for a pharma team, these questions separate a realistic-sounding demo from a rehearsal you can trust:

  1. Where do the persona's boundaries live? In the system's setup, or in whatever the model does in the moment?
  2. What is the score measured against? A conversation defined before the roleplay, or a general impression formed after it?
  3. What is the source of truth? The approved label for your market, or a summary of it?
  4. What evidence connects rehearsal to the field? And if there isn't any yet, will the vendor say so plainly?

What's next

Today, rehearsal in most pharma teams is an event: a cycle meeting, a launch workshop, a new indication. Where reTrain is heading is rehearsal that becomes continuous, part of the job rather than a date on the training calendar.

That makes the standards above more important, not less. When rehearsal happens every week instead of a few times a year, every boundary, every score definition, and every source document has to hold without anyone checking it by hand.

Pharma would never approve a molecule on a convincing story. A rehearsal should be held to the same standard.

FAQ

What is AI roleplay in pharma?

AI roleplay lets sales reps, MSLs, and market access teams rehearse conversations with AI personas of HCPs, KOLs, or payers before the real meeting, then get scored and coached on how it went.

How realistic are AI-simulated HCPs?

It depends on how the persona is built. A persona whose boundaries are defined before the conversation, grounded in the approved label for that market, stays in character and pushes back where real visits get hard. Realism inside a simulation is still different from proven impact in the field.

How is an AI roleplay session scored?

In reTrain, the score is measured against an expected conversation defined before the roleplay starts: what should be discussed, in what sequence, and in what tone. It covers conversation management, word choice and delivery, including pronunciation.

How does AI roleplay stay compliant with off-label rules?

In reTrain, leaving the role, off-label, and off-indication discussions are closed at the system level, with extra boundaries set for each roleplay. The approved label for the relevant market is the single source of truth, and every session is logged for compliance review.

Can AI roleplay replace field coaching?

It changes when coaching happens: before the visit instead of after it. Whether rehearsal improves field outcomes needs independent measurement, comparison groups, and time. reprai is running that work now.

Explore reTrain →

Ahmet Fatih Şahin is Product Owner at reprai. He is a pharmacist with a PhD in computer-aided drug design (Bezmialem Vakıf University) and was a TÜBİTAK 2214-A visiting researcher at the University of Tübingen. LinkedIn

Sources: Li D, Lutfi SL. JMIR Medical Informatics 2026;14:e79039 · Yu H, Zhou J, Li L, et al. AIPatient. Communications Medicine, December 2025


← All posts