AI in Live Clinical Studies: Exploring the Reality of Implementation 

The question most clinical operations teams are sitting with right now is not whether to use AI. It is how to introduce it into workflows that are already running, under regulatory scrutiny, with data that cannot be unwound.

That is a harder problem than the vendor demos suggest. And it is the one worth spending time on.

The Gap Between Interest and Deployment Is Not About Skepticism

Most clinical organizations are actively evaluating AI but have not yet moved it into live studies in any meaningful way. Of those that have, most are still running limited pilots in non-critical workflows, though a smaller group already reports operational use of AI in select workflows.

That caution is well-placed, and the reason for it is consistent: the barrier is not technical. What is holding adoption back is validation and regulatory uncertainty, with, at a distance, the absence of a governance framework. These are process and infrastructure problems, not capability problems. They are solvable. But solving them requires being honest about what AI implementation in a GxP environment actually demands, which is different from what it demands in most other industries.

Not All AI Carries the Same Risk Profile

One of the most practical things a clinical operations team can do before evaluating any AI capability is get precise about what kind of AI it is.

“AI” in clinical development currently covers traditional machine learning algorithms, convolutional neural nets, NLP tools, retrieval-augmented generation systems, and modern large language models. More than 1,500 cleared medical devices already use AI, but the majority rely on classical ML approaches that have been qualified and used for years. These are not interchangeable from a validation or governance standpoint, and treating them as equivalent creates real compliance exposure when you go to document or defend the process.

A more useful framing than “is it AI?” is: does the tool add visibility into the data, or does it stand between the reviewer and the data? A tool that flags more anomalies for a data manager to review is additive. A tool that filters what the reviewer sees, or that generates outputs that go directly into a regulated record, is doing something fundamentally different. The distinction determines how much validation rigor the workflow actually requires.

Retrofitting AI Into a Running Study Is a Different Problem Than Building It In From the Start

This is where a lot of early implementations run into trouble, and it is worth being direct about why.

When AI is introduced into a study that has been running for months or years, the system has to contend with data that was not collected with AI-assisted review in mind. An anomaly detection engine onboarded mid-study is not starting from a clean baseline. It is processing historical data, generating flags, and asking reviewers to adjudicate discrepancies that may or may not reflect genuine issues. If false positive rates are higher than expected, that burden lands on the same team that is already managing the study.

The practical lesson from organizations that have gone through this: pilot on new study starts, not active studies. The implementation overhead is lower, the data is clean, and the team can absorb a learning curve without it affecting an ongoing regulatory record. Selecting the right study type and the right team for early deployments is as important as selecting the right tool.

Timing compounds the problem in a second way. Anomaly detection or AI-assisted review tools need a critical mass of data to produce reliable signal. Introducing them too early in enrollment produces noisy output. Introducing them too late means reprocessing months of accumulated records before the capability is useful. Neither failure is about the technology. Both are about deployment planning.

Validation for AI Requires a Wider Frame

Traditional CSV methodology is built on deterministic logic: given input A, the system reliably produces output B, and you can document and test that relationship. Large language models and probabilistic systems do not work that way. The same input does not always produce the same output, and model behavior can shift if a vendor changes routing or updates parameters on the back end without surfacing that change to the customer.

This does not make AI unvalidatable. It means the validation scope needs to expand beyond the model itself.

Three layers need to be addressed. The orchestration layer is the most important: the system architecture that defines what the AI is permitted to do, in what sequence, with what constraints, and how low-confidence outputs are handled. This is what validation documentation needs to capture, and it is what an inspector will ask to see. A block diagram of the AI’s authorized steps, with clear boundaries around where human judgment is required, is the minimum starting point.

The human use layer is the second component, and it is frequently underweighted. AI-generated outputs tend to be well-formatted even when the underlying data contains errors, which makes errors harder to catch visually. Reviewers can move through formatted AI output quickly without recognizing a problem. Tracking review time as a proxy for genuine engagement, and building in escalation for outputs flagged as low-confidence, are both practical controls that most initial implementations skip.

The model performance layer is third. Benchmark data on model accuracy and sensitivity matters, but it is a baseline, not a guarantee. What matters operationally is whether the full system, orchestration, human process, and model together, is producing reliable and auditable outcomes within the defined context of use.

Context of Use Is the Foundation Everything Else Builds On

If there is one concept that clinical operations teams should internalize before deploying any AI capability in a live study, it is context of use.

Before governance, before validation planning, before model selection, an organization needs to be able to state precisely what the AI is being asked to do and where its authority ends. Not in general terms. Specifically: what inputs does it receive, what outputs does it produce, what decisions does it inform, and what happens if it is wrong.

This matters because AI use cases have a way of expanding. A capability scoped for anomaly detection in data cleaning can quietly absorb query generation, site communication, and escalation routing as the team gets comfortable with it. Each expansion carries different risk and requires its own validation scope. If the original context of use is not documented with precision, those expansions happen without deliberate review.

The E6R3 principle of proportionate oversight applies directly here. Outputs that feed into regulatory submissions, SDTM datasets, safety narratives, submission-ready documentation, require rigorous human review at every step. Internal operational records produced for audit readiness can tolerate more automation with a lighter review footprint. The right level of oversight is not uniform across a study. It is calibrated to what is at stake if a specific output contains an error.

Human in the Loop Is a Starting Point, Not a Complete Control

Routing every AI output through a human reviewer before action is taken is the right instinct for regulated environments. It is not, by itself, a governance strategy.

Human review adds value only when the human has a meaningful opportunity to apply judgment. When AI output volume is high and every document requires sign-off, sign-off becomes a gate rather than a check. The risk is not that reviewers are careless. The risk is that the system design does not give them a reason to engage carefully.

The more durable design pattern builds intelligence into the review process itself. AI outputs evaluated at multiple tiers, with confidence scores assigned and escalation triggered for low-confidence results, give reviewers a focused task rather than a volume problem. A reviewer directed to five flagged items in a fifty-document batch is doing real work. A reviewer asked to approve fifty documents in sequence is a control that looks good on paper and fails in practice.

This also means that as AI systems mature, the measure of quality shifts. Early implementations often track whether the AI performed correctly. Mature implementations also track whether the humans reviewing AI output are performing correctly, and build monitoring into both layers.

Learn more about Sitero’s AI clinical agent – Ash – below:

Pilots That Should Scale Look Different From Pilots That Should Not

Not every AI initiative should move beyond the pilot phase, and treating continuation as the default is a common and expensive mistake.

A pilot that should scale shows a clear signal against a defined metric, has the process and data infrastructure underneath it to reproduce that performance consistently, and has a team that understands the system well enough to operate it without the implementation team present. A pilot that should stop shows weak signal, exposes gaps in data quality or process readiness, or produces results that depend on conditions that cannot be replicated at scale.

Applying AI to an underperforming process does not improve the process. It makes the underperformance faster and more consistent. The prerequisite for AI readiness is process readiness, and the most useful thing a pilot can do is surface where that readiness is missing.

As mentioned in our hosted webinar – Reduce Live Study Risk With AI in eClinical Workflows – by Aman Thukral – eCOA localization is one area where the signal has been strong enough in practice to justify full-scale deployment. It sits on the critical path for study startup, the localization requirements are well-defined, and AI-assisted localization produces measurable cycle time reductions that are reproducible across study types. That combination, clear scope, measurable outcome, reproducible signal, is what a scalable pilot looks like.

Where the Investment Should Go

For clinical ops teams with limited capacity to take on AI readiness work alongside active study obligations, the question of where to focus first matters.

People before technology is the consistent answer from teams that have deployed AI successfully in regulated environments. Not because the technology is secondary, but because the foundational governance frameworks for qualifying AI in GxP settings already exist, borrowed largely from the medical device and radiology precedent. What most clinical organizations lack is not access to tools. It is the operational literacy to deploy them carefully, the training to use them appropriately, and the institutional knowledge to tell a ready use case from one that needs more groundwork.

Data foundations come second. For any AI application operating on structured clinical data, the quality of entity relationships and the integrity of the underlying data model will determine what is actually achievable. An LLM built on poorly structured data does not produce better insights. It produces worse ones, with more confidence.

Governance infrastructure comes third, but it is not optional. Clear accountability for each AI-assisted workflow, defined triggers for reassessment when performance changes, and documented boundaries around what the AI is authorized to do are not overhead. They are what makes the system defensible when it matters.

The Real Risk Is Quiet Drift

The implementation risks most teams plan for are the obvious ones: false positives, validation gaps, team resistance. The risk that tends to go unaddressed is quieter.

AI systems in clinical workflows learn from the data and examples they are given. If that data is incomplete, if the context of use expands beyond what was validated, if human review becomes nominal over time, the system drifts. Not dramatically, and not in ways that are easy to detect without deliberate monitoring. But in ways that compound.

The organizations that deploy AI in live studies successfully are not the ones that move fastest. They are the ones that define precisely what the AI is doing, monitor whether it is still doing that, and build the human processes to catch the gap when it is not.

That is not a reason to wait. It is a reason to be specific from the start.

Need to Reduce Live Study Risk with AI?

Watch our full webinar, “Reduce Live Study Risk With AI in eClinical Workflows,” hosted by Sitero’s Dr. Joby John, Vice President, Service Delivery, on demand:

Sitero’s Mentor eClinical platform is built to support AI-ready clinical operations with the governance infrastructure to deploy safely. Learn more at the link above or connect with our team here.