Industrializing Clinical Biometrics: What the Room Found
Executive Insight Report - Clinical AI Mastermind 2026, Session 2
Event
July 29, 2026
Johnson & Johnson, 320 Bent Street, Cambridge, Massachusetts
Statement


When we set the question for Session 2, we made a claim the room had not yet tested. We said that clinical biometrics is no longer primarily a productivity problem, and that it has become an enterprise AI problem. The distinction matters. A productivity problem is solved by producing more, faster, with fewer people on it. An enterprise problem is solved by deciding who is allowed to do what, on what evidence, with whose signature at the end. On July 29, that claim met a room of people who run this work at scale, and they did not argue with it. They told us what it costs.
In welcoming the room, Dr. Charmaine Demanuele of Johnson & Johnson set the frame. The purpose of scaling biometrics is not scale for its own sake. It is to use AI across the enterprise in ways that reach patients who need a treatment sooner. She also told the room what had changed since Session 1, that there was deliberately more time for discussion, and asked people to stay to the end and take part. They did, and the last hour of the afternoon is where this report spends most of its length.
Dr. Louise Liu, Chief Executive Officer of Hill Research, put the afternoon’s real question in one line as she opened for the co-hosts.

Do we really trust AI during the clinical trials?
— Dr. Louise Liu, Chief Executive Officer, Hill Research
By six o’clock the room had an answer, and it was not a yes or a no. It was a condition.
AI can scale production. Only an accountable operating system can scale evidence.
What the Room Said at the Door
Before the forum, every guest was asked one question.
In your experience, where does clinical biometrics most struggle to move from individual AI pilots to enterprise scale, and what has your organization not yet solved?
Twenty-six people answered, twenty-three of them substantively, and all twenty-three went up on the wall, one to a card, in the words they were written in. The wall was the room’s own evidence, and editing it would have made it ours. During the break, several guests found their own sentence there.

Sorted into five themes, the distribution was lopsided in a way that shaped the whole afternoon.

The Foundation took ten of twenty-three. Close to half the room, answering a question about enterprise AI, wrote about data. Not about models. A director of data science at a top-ten pharma wrote “Lack of standardized input.” A VP of biostatistics and data management in diagnostics wrote “Data lineage and ingestion.” A VP of patient experience at a rare disease company wrote that “The data used for modeling and scaling isn’t yet reliable.” A lead data scientist at a top-twenty pharma gave the longest technical account of why, describing interoperability and site-specific distribution shift as the real obstacles, and adding that a translation layer between sources is fragile and time consuming to maintain. An IT business partner at a rare disease biotech answered with unusual directness: “Our organization is young. We are establishing foundational data and computational capabilities before expanding into AI.”
The People took five. A senior director of trial delivery data strategy at a top-ten pharma wrote “Change management (process and training),” which turned out to be the seed of the afternoon’s most quoted moment. A head of global portfolio statistics named reproducibility and regulatory acceptability, then added something no other card said: “imagination of what can be done, and getting people who can do it.” A chief data innovation officer wrote about the difficulty of gaining the trust of doubters, particularly in smaller companies. An SVP of clinical development at a mid-size biotech began by saying he was not doing anything at enterprise scale, then described the problem anyway, as one of rolling out and training for people who arrive with very different levels of familiarity and from very different roles.
Trust in the Output took three, and they were the sharpest sentences on the wall. A head of biometrics at a biotech wrote that AIs “give wrong answers and stubbornly insists it’s correct.” A senior director of data science and AI at a top-ten pharma put the same thing structurally: the hardest part of moving from pilot to enterprise deployment is “building the guardrails that allow the organization to trust that AI at scale.”
The Pace took three. One principal data scientist answered with a single word, “Regulatory.” A VP of strategic consulting at a CRO wrote that speed and quality both matter, which is the whole tension in seven words.
The Pilot Trap took two. Only two of twenty-three named the pilot-to-enterprise gap directly, and that is worth noticing, because it is the gap we called the session to examine. A data scientist from academia described it precisely: pilots succeed because they are built around one study, one dataset, or one enthusiastic team, while enterprise deployment needs standardized data, validated processes, system integration, governance, and adoption across functions. An SVP of data sciences at an oncology biotech described the version that keeps small companies stuck, which is finding the time and the people to build tools while still delivering the work the business needs today.
The room did not experience its problem as a problem of pilots. It experienced it as a problem of foundations and of people, and it said so before anyone spoke.
The longest answer on the wall came from the CEO of a clinical AI startup, and it did something none of the others did. It sorted itself into our three lenses without being asked to. At scale, data standards and infrastructure are rarely designed for reuse across trials or therapeutic areas, so what works in one program does not transfer cleanly. On agility, adaptive and real-time workflows need a level of cross-functional alignment most organizations have not built. On integrity, the governance frameworks for AI-derived endpoints in submissions are still moving, and that uncertainty slows adoption more than the technology does. Then the sentence that could have opened the session:
The common thread is that the organizational and regulatory layers haven’t kept pace with the technical ones.
— From the wall, card 23
That was written days before anyone arrived. If the constraint were productivity, these answers would have named volume, backlog, and headcount. They named lineage, standards, ownership, consent, trust, and skill. Every one of those is a question about who is responsible for what. The wall made the session’s argument before the session began.
It also set the panel’s first question. When the discussion opened, the room’s most common answer was read back to it, that the data is not standard, not interoperable, and not yet reliable, and the first cross-pillar question was built directly on top of it.
What the Three Lenses Revealed
The same three lenses run through all four sessions of 2026. Scale is the view from industrial volume. Agility is the view from small populations and adaptive design. Integrity is the view that sits across both and asks whether the evidence holds. Three speakers took one lens each, for fifteen minutes, with questions held for the panel.
Scale · Dr. Savina Jaeger, Executive Director, Head of Biometrics, NovaBridge Biosciences

Her opening slide carried a disclaimer that the material is her personal view and not representative of NovaBridge Biosciences, and it holds for everything reported here.
Savina opened by separating a word the industry uses as if it meant one thing. Scale can mean throughput, more trials and more outputs per person. It can mean speed, the calendar time from start to readout. It can mean consistency, the same analysis done the same way everywhere. She asked the room which of the three their leadership means when it says scale the AI, and the question landed, because the answer changes what you build.
Her argument was that AI delivers the first and the third far better than the second. A trial’s duration is set by biology. Even if every patient enrolled at once, an overall survival readout still waits for events to accrue. AI can shorten components of a clinical study. It cannot shorten the disease.
What it can shorten is standardized production, and this is where she made the distinction the session turned on.
What we can scale with AI is standardized production. What doesn’t is the judgment that decides whether the output is right.
— Dr. Savina Jaeger
She then showed where that judgment goes missing, and her examples all sat at the boundary between functions. A clinician asks for a survival analysis and gets a clean curve back in seconds, with the censoring wrong and no statistician anywhere near it. A dataset conforms perfectly to the standard and answers the wrong scientific question. Queries close themselves. Safety narratives get written and lightly reviewed. Her phrase for this was green check, wrong science, and it produced the line the room repeated most.
A green check is necessary but not sufficient.
— Dr. Savina Jaeger
The historical point underneath was the sharpest thing said all afternoon, and it explains why this is new. Before these tools, you could not write analysis specifications without knowing statistics, and you could not touch the database without data management training. Competence was doing two jobs at once. It produced the work, and it quietly controlled who was allowed to produce the work. Competence was an access control that nobody had to write down, because it enforced itself. Automate the production and that check disappears with it. The rules were never written down, because nobody needed to write them.
Her proposed answer came from the teaching hospital, where a resident can see patients and an attending physician signs. The mechanism is not trust, it is privileging: a named person, credentialed for a defined scope, accountable for the decision. Translated into biometrics, she set out five gates that are mostly already familiar. Provenance. The privilege to promote an output into evidence. Conformance treated as separate from correctness. Methods pre-specified rather than chosen after the result. Independent reproduction, meaning the result can be rebuilt from source. Alongside them she fixed accountability to the function, so that anyone may generate and only the accountable function signs. What is new is not the controls. It is the need to state them, now that competence no longer states them for us. Her operating instruction was to separate work that is AI-produced from work that is evidence-validated, and to name who has the authority to move something from the first category into the second.
Four times during fifteen minutes she stopped and put a question to the room rather than answering it herself. That is unusual in a keynote and it set the temperature for the rest of the afternoon.
Agility · Dr. Roberto Araujo, Senior Medical Director, Pompe and Fabry Disease, Sanofi

Roberto opened by saying the first keynote set his up, and then quoted his own thesis.
The limit on adaptive design is governance rather than technology, and the answer is timing.
— Dr. Roberto Araujo
Agility in rare disease means adaptive design, which he defined as the capacity to modify the trial architecture, the statistical methods, the data collection framework, and the population feasibility constraints, while holding evidentiary integrity and regulatory compliance in place. Read that definition slowly and AI has a role in every clause. That was his point. The technology is not the binding constraint. The constraint is the governance around it, and the fact that most of it must be settled before anything starts.
He grounded this in the programs he works on. Pompe disease and Fabry disease are both rapidly progressing, both heterogeneous, both multisystemic, with trial populations of twenty patients, perhaps two hundred at the outside. Those are the numbers enrolled in a Phase Two or Phase Three study, not the size of the disease population. In infantile onset Pompe, the trajectory is measured in weeks and sometimes days. An episodic trial design cannot see a disease moving that fast, and a control arm would not be ethical, so the comparator has to come from historical registry data. That is workable, but only if the comparator is defined prospectively. Decide afterward and it is no longer evidence, because the comparison was chosen once the result was already known.
Patient safety and evidence integrity remain primary objectives.
— Dr. Roberto Araujo
From there he named four places where adaptive infrastructure fails operationally. Decision authority, because AI produces more interim signals and someone must already hold the right to act on them. Monitoring, because detecting an inflection point is not the same as being accountable for what follows. Evidence translation, because a signal a clinician cannot act on through a defined pathway is not an improvement. And what he called evidentiary attrition, where AI synthesizes across registries and historical trials and the chain of accountability thins as it goes. His rule was that accountability structures must be established before protocol formalization, not after a model produces something interesting.
The Fabry work showed the other side. Because effective treatments exist, a trial without an active comparator would not be ethical, so the design carries comparator harmonization, imaging biomarkers, and small-population statistics all at once, with AI holding the analytical consistency across them. He ran out of time with material still to give, which the room noticed.
Integrity · Dr. Bhaskar Dutta, Head of Digital Health and Medical Affairs Technologies, Alexion

Bhaskar began where the other two ended.
What does it serve if we cannot trust the evidence?
— Dr. Bhaskar Dutta
Then he told a story rather than showing a framework. In 2006, work at Duke University claimed that an algorithm built on clinical and genomic data could predict which chemotherapy regimen would work for a given tumor. It was not theoretical. Real patients in active trials were assigned on that basis. When independent statisticians tried to reproduce the results in 2010, the underlying data did not hold up. Eleven papers were eventually retracted, and misconduct was not formally confirmed until 2015, nine years after the original claim, and the part he deliberately left off the slide was the patients who were assigned to the wrong treatment while it lasted. He mentioned, almost in passing, that he had nearly joined that group as a postdoctoral researcher. Integrity was not an abstraction in the room after that.
His question was what AI changes, since integrity has always mattered. Three things. Compromise at scale, because a flaw in a system that produces at volume is reproduced at volume. Compressed review windows, where the weeks between database lock and submission were already the place where problems went uncaught. And methods that are not interpretable by construction, so an error can exist without being visible even to someone looking for it.
He then mapped AI across the evidence lifecycle, from protocol design through data capture, statistical analysis, medical writing and submission, and on into pharmacovigilance, and made the point that these stages are not one problem. The models differ, the risks differ, the time scales differ. So he abstracted the failure modes instead of the applications. Provenance gaps, where nobody can say in minutes which version of which data produced a result. Model drift and versioning, made harder when the model is hosted by a third party and changes without you. Automation bias, which he tied directly to Savina’s green check, where a human is in the loop on paper and a checkbox in practice. And the explainability gap, where with a Cox model or a regression you can state the parameters, and here you often cannot.
Against those he set four pillars, in deliberate order. Data provenance first, because without it the rest do not matter. Model reproducibility, meaning same inputs and same model return the same result, which is difficult when continuous deployment is the norm and the honest answer is predetermined change control specified in advance. Accountability, with a person who signs, since the company that built the model will not be taking it. And auditability, which produced the line that best summarizes his lens.
Show me how you got the results.
— Dr. Bhaskar Dutta
An auditor does not want the result. The result is not in question. What is in question is the path, and if provenance, reproducibility, and sign-off are real, then the audit is a retrieval exercise rather than a forensic one.
Two further points deserve to survive. Consent collected years ago for clinical use should not be assumed to authorize training a model today, and that assumption is quietly everywhere. And the regulators are moving, from good machine learning practice and predetermined change control at the FDA through to the EU AI Act, which will treat most healthcare uses as high risk and carries real penalties. His closing observation was that every organization is building its own governance model, and some are now building governance models to govern their governance models.
He closed on a single slide carrying Warren Buffett’s test for the people he hires, from the Berkshire Hathaway annual meeting in 2005. Everything in the preceding forty minutes, he said in effect, is intelligence and energy. Process, infrastructure, governance. None of it substitutes for the quality the quote names first, which is the only one of the three a person has to choose.
In looking for people to hire, you look for three qualities: integrity, intelligence, and energy. And if you don’t have the first, the other two will kill you.
— Warren Buffett
The Cross-Pillar Panel: Each Speaker Answered a Question from Another Lens
The series is built on a cross-pillar design, so the three lenses do work on each other rather than in parallel. Each keynote generates one question, and that question is routed to a different lens, so the speaker who has to answer it is the one who did not raise it. In Session 2 the cross-pillar design ran in full for the first time, and all three routings landed.
The questions were written by the curation team from the final decks, and given to the panel in advance, so nobody was caught by surprise. What could not be planned was what the room did to them.

First cross-pillar question · Scale to Integrity
Savina had argued that what AI scales is standardized production. The wall said, ten times over, that the data underneath is not standardized. Bhaskar’s four pillars begin with provenance and assume a lineage you can trace. So the question went to him: what do you tell an organization where that lineage does not exist yet? Do they start there, or do they start with the pillars?
He did not defend the pillars. He answered as someone who has had to work without them. You assemble what you can. Control arms pooled across studies to reach statistical power. Published figures read with image processing to build a cohort out of the literature. Registries that were never designed as trial infrastructure, used anyway. His position was that a missing foundation is not a reason to stop, it is a description of the work.
Roberto, from Agility, said the same thing from his own programs, where disparate data has been used to derive quantitative systems pharmacology models and then to design control arms. Not easy, he said, but doable.
Savina was asked to react to a position that ran against her own, and she kept her position without dismissing theirs. You do not get away from standards. What changes is where the standardizing happens. The unstandardized data still has to be converted, and that conversion is itself one of the better uses of these tools.
That is the finding. The three answers are not a disagreement, they are a sequence. Start with what you have, convert as you go, and be honest that you are converting rather than pretending the standard was there. Nobody in the room had to choose between the pillars and the mess.
Second cross-pillar question · Integrity to Agility
Bhaskar’s first pillar says having the data is not the same as being allowed to use it. Roberto’s Fabry registry holds twenty-five years of follow-up, so the question went to him. When AI ingests it, whose consent are you relying on, and who checked?
His answer was the most concrete of the afternoon. The Fabry registry runs to roughly eight thousand patients worldwide over twenty-five years, and other rare disease registries hold different numbers. It did not name AI in its consent form, because nothing did then. Patients are re-consented periodically and for specific analyses, which he described as a genuine hurdle rather than a formality. Some have been lost to follow-up. Some have passed away. Only the data of patients who have re-consented for that specific purpose can enter a model.
Pressed on the guardrails, on what stops unconsented data from being pulled in without anyone noticing, he described a chain rather than a control. Data is classified into buckets at ingestion. It carries tags when it is pooled for analysis. An automated check runs, then a human check, with audit trails behind both. Beyond the system there are advisory boards of physicians who raise privacy directly, and more recently the patient associations as well.
Bhaskar, whose keynote had produced the question, then answered the operational half of it. His proposal was to treat data the way the industry already treats a drug product, with the equivalent of packaging and a label. Attach to the data what it may be used for and what it was consented for, so authorization is a property of the data rather than a check someone has to remember to run. He also opened a door Roberto had left closed: consent can be reasoned about patient by patient, not only study by study, so failing to re-contact everyone does not disqualify the group you did reach.
Roberto came back once more, and this is the exchange to keep. The technology is real, he said, and it still cannot be fully automated. There has to be a human check. Then he made the argument that being open about the work achieves more than compliance alone. His teams invite clinicians and patient associations to the table precisely because those groups have been critical about how data is used, and the answer to that criticism is not a better control, it is a seat.
It’s never too much to ensure patient rights and privacy is ensured.
— Dr. Roberto Araujo
The routing did something a parallel panel could not. It made the person who raised the principle answer for its cost, and it made the person paying the cost describe the machinery. Consent stopped being a compliance topic and became an operating design.
Third cross-pillar question · Agility to Scale
Roberto’s answer to agility runs on infrastructure most of the room does not have. Savina had said the twelve-person biotech and the twelve-thousand-person pharma face the same standard. So the question went back to her. What does the twelve-person version look like, and where does the claim break?
She separated the parts that are cheap from the part that is not. Provenance, traceability, and named accountability hold at any size, and none of them are expensive. She said she had been discussing exactly this on the Monday before the session, and that these controls can be implemented in a one-page document. What costs is validation, and small companies outsource it. That is where her claim gives way a little rather than failing: you can rent the expertise, but then you are managing a relationship instead of a capability, and you have to be able to judge the quality of what you are handed.
Alexandre pushed into the human version of the same problem. A ten-person biotech with one statistician who also covers quality assurance, because nobody else understands the regulatory requirement. That person does the work and checks the work. How does a company like that meet the standard?
Bhaskar answered from the opposite extreme, and was candid that it was not his experience, since his company has more than a hundred thousand people. What he offered instead was the asymmetry underneath. A twelve-person company with no product on the market carries a different kind of risk from a company with thirty products in a hundred countries. The large organization checks twice and three times because it has more to lose, and that defensiveness is not caution for its own sake. Read against Savina’s answer, it says something uncomfortable and true. The standard is the same, the consequence of missing it is not, and the smaller organization is the one with less to spend on getting it right.
What the cross-pillar design produced
Then, in the middle of that third answer, Savina set down the question that outlasted the afternoon. If AI absorbs the junior work, where does the next generation of senior judgment come from? She said she had no way through it, and Roberto said two minutes later that he had the same question in his own talk.
The cross-pillar design did not produce that question. Both of them had prepared it before they arrived, separately, on different lenses, without having spoken to each other. What the design did was make that visible. Nothing in three parallel keynotes would have shown it, and it came out live rather than in our notes afterward. The question has its own section below, because it is not finished.

Roberto put the relationship between the first two lenses better afterward than anyone managed on the day.
Competence as the binding constraint and governance as the binding constraint are two sides of the same argument.
— Dr. Roberto Araujo, writing to the host team after the session
What the Room Asked
At twenty past five the room joined the panel, and stayed with it for the next half hour. Guests are identified here by role and setting rather than by name, as they are on the wall.

The first question came back to the registry. If evidence is derived from a registry and assembled with AI, how does the FDA treat it, and what has been agreed with them? Roberto’s answer was about sequence rather than submission. In rare disease his teams meet the agency before anything is written down, and the relationship runs on frequent contact rather than formal milestones. He described regulators who are more open to experiment where the population is small, who will sometimes say plainly what they want to see, and who have accepted a different approach when his team came back with reasons. Asked how such data gets validated, he was direct that there is no standard answer. Guidelines hold and cannot be departed from, and everything above them is customized case by case, because of small numbers, fast progression, and endpoints whose clinical meaningfulness is itself under discussion.
Then Alexandre brought in the guest whose form answer had named change management, process and training, and asked him to put it to the panel. What followed is the image most people took home.
His teams do clinical data management, biostatistics, and statistical programming. They know the road from Boston to New York. They have been driving it for years, and they are the domain experts, in the sense that they know where the road goes. Now they are handed a jetpack. There are two problems, not one. The first is learning to fly it. The second is standing in front of an auditor afterward and explaining how it works.
The panel answered in three directions and none of them contradicted the others. Bhaskar said there is no single fix and the problem has to be worked in several dimensions at once. Most of the workforce needs a basic understanding of these tools, not expertise, because the risk of not having it is real. Alongside that, organizations need specialists, sometimes hired and sometimes found inside, and he said his own organization was doing both. He also noted that some of the people who had spent their careers driving the car were becoming expert with the jetpack on their own, which is happening organically rather than by program. Roberto answered from his own experience, as a physician who had no AI in his training and went after it deliberately, not to become an expert but to become a competent user. Everyone, he said, is working out where they sit: doing the work, interpreting what comes out of it, or designing how the tool gets used at all.
The guest came back, and his second question was better than his first. Those are individual paths, he said. What about organizational design? The frameworks, the data, the governance all need to be there, but how does an organization move from individuals who have adapted to a workforce that has?
Savina answered by moving the question from skill to accountability.
You cannot hold a model accountable. You hold a human accountable.
— Dr. Savina Jaeger
The production layer can be shared with AI. Responsibility cannot. What is needed is a framework that holds regardless of who or what produced the work, with critical checks that ask whether this actually meets the standard. She returned to what she had called confident hallucination in her keynote, the human version, where something looks right and is wrong, and said she sees it most in smaller organizations where one person is wearing several hats and a clinician may be running the statistics. The answer is not to prohibit these tools. It is to be explicit about what can be promoted into evidence, and to layer responsibility behind it.
Later the discussion turned to what human in the loop means when it is written into a protocol, and both men treated the phrase as too loose to be useful. Roberto put it through an example everyone in the room had lived. You ask a model something you already know, and you can tell whether the answer is right, or could be better, because you know the subject. Give the same tool to someone who does not, and the discernment is not there to be applied. Bhaskar said the phrase has to be resolved into questions rather than used as a standard: what decision is being made, how mature is the use, how critical is the outcome, and who is the right human for that particular judgment. He also mentioned that his teams have used two independent models to check each other, which is verification of a kind, though it does not move the accountability anywhere.
Three questions from the room then went at the enterprise problem from three different angles.

One guest raised cost, and did it by reading the wall’s question back before answering it. The impression so far has been that AI is cheap. The models keep getting bigger, the context windows keep growing, everyone is now dependent, and the companies selling it will eventually need to make money. The phrase he used was that the honeymoon is over. Savina answered honestly that she did not know where it goes, and that the biotech world is already struggling for investment, which makes the dependence expensive in a way nobody has priced.
Another returned to integrity and made a point about where it lives. It cannot sit on the model alone, it has to run across the pipeline. But as models get more complex, explainability becomes part of integrity, and explainability is still largely a research subject with little applied use in the industry, let alone a settled regulatory view. Bhaskar agreed that it is the step that sets the pace, said that research is actively focused on building it in, and then declined the obvious conclusion. The absence of explainability is not a reason to stop using these tools today. It is a reason to match the use to the stakes.
The last of the three asked whether the familiar chain, from data to insight to decision to action to value, applies to clinical submissions, and whether every link needs a human. Bhaskar’s answer closed the loop on his keynote. The chain applies, but each link is a different problem, with different models and different risks, so the human belongs in the loop on a risk basis rather than a uniform one. An early protocol draft written to save time is not the same object as a final submission table, and treating them alike is how review becomes a formality.
A final question about whether AI can be trusted to predict what has not happened yet brought Roberto back to something already running. His teams train models to find undiagnosed patients in electronic health records and claims data, with a probability threshold that can be set high or low depending on what is being looked for. The model learns from each patient it identifies. He described it as working well and improving, and as one of the places where the technology is doing something people cannot do at all rather than doing faster what people already do.
The Question We Could Not Answer
Every session in this series is meant to leave the room with more than it arrived with. Session 2 also left it with something missing, and that is the part worth reporting most carefully.
The question came up in the third cross-pillar routing, while Savina was answering something else. AI can do a great deal of the junior work. That is the point of it, and the efficiency is real. But the junior work is not only output. It is the mechanism by which judgment is made. She used the hospital again, as she had in her keynote, where the resident sees patients under supervision and eventually becomes the attending who signs. Take away the residency and the attending still has to exist. Where does that person come from?
Roberto took it up immediately and said he had the same question in his own material. That was not politeness. Both had written it down before the session. Savina closed her slide on who watches the watchers with it, adding that she genuinely did not have this one, and listed the talent pipeline among the four positions she wanted the room to attack. Roberto ended his keynote with four unresolved questions for the scientific and regulatory community, the last of which asks what mechanism develops clinical judgment in the next generation if AI reduces junior-level analytical work. Two speakers, two lenses, two companies, no contact beforehand, and the same unsolved problem in both decks.
An hour later she came back to it at the end of a long answer, and finished the thought. There is a moment, she said, when we become too dependent on the easy answer. So how do we keep learning, and how do we stay able to defend what we produce? Then, without softening it: she does not have an answer, and she can see the moment happening now.
The room had already said this before the session started, without recognizing that it was the same problem. Five of the twenty-three cards on the wall were about people. One asked for training that fits people arriving with very different levels of familiarity. One named the difficulty of gaining the trust of doubters. One, from a head of global portfolio statistics, asked for something harder than either: the imagination to see what could be done, and the people who can do it. Those cards read as a training problem in July. Read against Savina’s question, they are the same problem at a different point in time. Training is how you bring people to the tool. Her question is about what happens to expertise once the tool is doing the work that used to build it.
We are not going to resolve it in a report. What we can say is that it was the one question two senior people in different parts of this industry independently prepared, both marked as unanswered, and then declined to answer in front of a room that wanted them to. Convergence of that kind is worth more than an agreement reached in the room, because neither of them was influenced by the other. That is a finding, not a gap.
Savina has since written that the question stayed with her, that she does not think any of us have the answer yet, and that she would welcome taking it further. It will be on the table in October.
I’m glad the question about where the next generation will learn judgment and competence, if junior jobs are delegated to AI, and only senior jobs are involving human oversight and competence, stayed with people. It stayed with me too, and I don’t think any of us have the answer yet.
— Dr. Savina Jaeger, writing to the host team after the session
What Session 2 Sets Up
Something happened three times on July 29 without anyone pointing it out.
Savina’s gates require methods to be pre-specified, which means chosen before the result exists. Roberto said accountability structures must be established before protocol formalization, and that if the structure is not settled before anything starts, the rest is a recipe for failure. Bhaskar’s answer to models that change under you was predetermined change control, which means agreeing in advance what may evolve and how far. Three speakers, three lenses, three different sets of vocabulary, and the same instruction underneath all of it. Decide first. Everything the afternoon described as an operating system turned out to be a set of decisions that only work if they are made before the work begins.
That is why this report ends by pointing upstream. The controls we spent the afternoon describing are applied to a trial that has already been designed. If the decisive moment is earlier, then the design itself is where the next question lives.
Session 3, October 7. The Architectural Shifts in Phase Three.
Phase Three is where the cost of a late decision is highest and the room to change anything is smallest. If AI is now part of the engine that produces evidence, the architecture of a Phase Three trial is not the same object it was, and the field is still treating it as though it were. That is the reframing Session 3 will test, and it inherits two things directly from Session 2. Savina’s unanswered question about where judgment comes from, which she has asked to take further. And Roberto’s account of small populations and adaptive design, which is what Phase Three looks like when the usual assumptions are removed.
Session 4, November 18. Data Provenance and Speed.
Ten of twenty-three cards on the wall in July were about the foundation, and provenance was the first pillar of the Integrity keynote. Session 4 puts the two words that were never really reconciled in the same room and asks whether they can be.
The year’s question stands where it did in January. How is clinical AI redrawing the architecture of drug development? Session 1 argued the submission last mile is no longer a data problem but an architecture problem. Session 2 argued that clinical biometrics is no longer a productivity problem but an enterprise AI problem. Both are the same move, made at a different altitude. A familiar difficulty, still described in the language of the thing it used to be, turns out to have become something else while everyone was busy with it. October takes that move one step further upstream.
Closing Note from the Curation Team
About forty people came to Cambridge on a Wednesday afternoon in heavy rain, from roughly twenty-eight organizations, and stayed past six.
What we heard from them afterward said more about the value of the room than about the program. One guest wrote that it helped to hear other people facing the same problems, and that they had thought they were the only one saying these things. That sentence is worth more than a strong review. Everyone in that room is senior enough to be the person others look to for the answer, which is a lonely position when the answer does not exist yet. An invitation-only room is not for the exclusivity. It is so that people who are expected to be certain can say, in front of peers, that they are not.
What we take from Session 2 is that the controls holding evidence together are mostly not new. They used to be enforced by competence, and now they have to be stated, assigned, and signed. That is why the room’s own line is the one we have carried everywhere since.
AI can scale production. Only an accountable operating system can scale evidence.
Our thanks to Dr. Charmaine Demanuele of Johnson & Johnson, who anchored the session, welcomed the room, and asked for more discussion time than Session 1 allowed, which turned out to be the right instinct. To Angellica White-mann, who made the venue work, and to the J&J team. To Dr. Louise Liu of Hill Research, who opened with the question the afternoon was really about. To David Hall, who held the rundown and kept a full program to time. To Dr. Alexandre Duprey, who moderated the cross-pillar panel in its first working outing and made three routings feel like a conversation rather than a mechanism. To AKT Health and the Tech Impact Foundation as co-hosts, and to Boston International Media Consulting as operations partner. To Kaige Zheng, whose notes from the panel are the record this report is built on, and to Sonny Zhao for the photographs in these pages. And to the supporting team who set the room, ran the door, and carried the microphones.
Above all, to Dr. Savina Jaeger, Dr. Roberto Araujo, and Dr. Bhaskar Dutta, who each prepared fifteen minutes on a lens they did not choose, then answered a question drawn from someone else’s material, and who between them were willing to say out loud that one of the most important questions in front of this field is one they cannot yet answer.

Session 3 is on October 7 in Cambridge. The Architectural Shifts in Phase Three. We are not going there with an answer in hand, and we would rather not. Bring the question you could not answer in July, and we will work on it together.
Fion Liao · Curation Team Member, Clinical AI Mastermind 2026
Appendix A · The Controls, in One Page
What each lens asked this room to check. Everything here comes from the three keynotes.
Scale · five gates, and what sits behind them
Provenance. Where the output came from. The privilege to promote, so that only the accountable function moves work into evidence. Conformance treated as separate from correctness, because a clean format is not right science. Methods pre-specified rather than chosen once the result is known. Independent reproduction, meaning the result can be rebuilt from source.
Behind the gates sit two rules. Tier everything, so exploratory work is labelled and evidence is validated and signed. And fix accountability to the function: anyone may generate, only the accountable function signs.
Agility · four places governance fails
Decision authority, because AI produces more interim signals and the right to act on them has to be explicit. Monitoring, because detecting an inflection point is not the same as being accountable for what follows. Evidence translation, because a signal a clinician cannot act on through a defined pathway is not an improvement. And evidentiary attribution, where AI synthesizes across registries and historical trials and the chain of accountability thins as it goes.
The rule underneath: accountability structures are established before protocol formalization, not after a model produces something interesting.
Integrity · four risks, four pillars
The risks. Provenance gaps, where nobody can say in minutes which version of which data produced a result. Model drift and versioning, made harder when the model is hosted elsewhere and changes without you. Automation bias, where a human is in the loop on paper and a checkbox in practice. And explainability gaps, where the parameters cannot be stated.
The pillars, in this order. Provenance, because without it the rest do not matter. Reproducibility, meaning same inputs and same model version return the same result. Accountability, with a named person who signs. And auditability, so the path can be shown years later.
Appendix B · What the Room Said at the Door
All twenty-three substantive responses, in the words submitted, as printed on the cards in the room. Attribution is by role and organization type. No names appear.
The Foundation · 10 of 23
| # | Role and setting | Answer, as submitted |
|---|---|---|
| 01 | Director, R&D Digital and Data Strategy · top-10 pharma | Data management |
| 02 | VP, Biostatistics and Data Management · diagnostics | Data lineage and ingestion |
| 03 | Director, Data Science · top-10 pharma | Lack of standardized input |
| 04 | Global Head of Biometrics and Data Management · specialty pharma | The data privacy part |
| 05 | Global Head of Data Management · biotech | Data Quality (including Medical Data Review) |
| 06 | Director, Clinical Programming and Data Science · rare disease pharma | Make Data Available for AI and governance |
| 07 | VP, Patient Experience and Insights · rare disease pharma | The data used for modeling and scaling isn’t yet reliable. |
| 13 | IT Business Partner for R&D · rare disease biotech | Our organization is young. We are establishing foundational data and computational capabilities before expanding into AI. |
| 14 | Director, Business Development · clinical AI | Infrastructure should be built up to adapt to the AI data capture and all downstream processing. |
| 20 | Lead Data Scientist · top-20 pharma | In my experience, data interoperability, site specific data distribution shift are the biggest technical challenges. The data interoperability is a bigger issue it is not easy to force standardization across sources. Building a translation layer is fragile and time consuming. |
The People · 5 of 23
| # | Role and setting | Answer, as submitted |
|---|---|---|
| 08 | Sr. Director, Trial Delivery Data Strategy · top-10 pharma | Change management (process and training) |
| 09 | Associate Director, AI and Digital Sciences · top-10 pharma | Clear literacy and harmonization across therapeutic areas |
| 15 | Chief Data Innovation Officer · biotech | Investments to scale. Difficult to gain the trust of doubters especially for small to midsize companies. |
| 16 | Head of Global Portfolio Statistics · top-10 pharma | Reproducibility. Regulatory acceptability. But also imagination of what can be done, and getting people who can do it. |
| 21 | SVP, Clinical Development · mid-size biotech | Admittedly, I am not doing anything at “enterprise scale” but I think this is likely a transferable issue between small and scaled: proper roll out and training protocols fit for users from varying backgrounds of familiarity and prowess with AI tools and from different roles within development teams (scientist vs exec vs clinical etc etc) |
Trust in the Output · 3 of 23
| # | Role and setting | Answer, as submitted |
|---|---|---|
| 10 | Medical Affairs Director · top-10 pharma | Compliance, trust, safety (from the platform) |
| 17 | Head of Biometrics · biotech | Trusting AI regarding data breach flawless deliverable. AIs give wrong answers and stubbornly insists it’s correct. |
| 18 | Sr. Director, Data Science and AI · top-10 pharma | The biggest challenge in moving from an AI pilot to enterprise-deployed AI is building the guardrails that allow the organization to trust that AI at scale. |
The Pace · 3 of 23
| # | Role and setting | Answer, as submitted |
|---|---|---|
| 11 | Principal Data Scientist · top-20 pharma | Regulatory |
| 12 | VP, Strategic Consulting · CRO | Speed and quality are really important |
| 23 | CEO · clinical AI startup | The shift from pilot to enterprise exposes gaps across all three dimensions. At scale, the challenge is that data standards and infrastructure are rarely designed with reuse across trials or therapeutic areas in mind — what works in one program doesn’t transfer cleanly. On agility, adaptive and real-time workflows require a level of cross-functional alignment that most organizations haven’t built yet. On integrity, the governance frameworks for AI-derived endpoints in regulatory submissions are still evolving, and that uncertainty slows adoption more than the technology does. The common thread is that the organizational and regulatory layers haven’t kept pace with the technical ones. |
The Pilot Trap · 2 of 23
| # | Role and setting | Answer, as submitted |
|---|---|---|
| 19 | SVP, Data Sciences · oncology biotech | We need to find time/resources to build tools while continuing to support the business needs of a small nimble biotech company |
| 22 | Data Scientist · academic | Clinical biometrics often struggles to scale AI because successful pilots are built around one study, dataset, or enthusiastic team, while enterprise deployment requires standardized data, validated processes, system integration, governance, and adoption across functions. |
Clinical AI Mastermind 2026 · Session 2 · Executive Insight Report.