Language Fluency in LLMs: A Critical Obstacle to Multilingual Evidence Interpretation for EU Joint Clinical Assessments


August 6, 2026

Large language models (LLMs) are rapidly becoming part of the Health Technology Assessment (HTA) workflow. From literature screening to evidence summarization, AI is increasingly being evaluated as a tool to accelerate evidence synthesis and support the preparation of Joint Clinical Assessments (JCAs) under the European HTA Regulation. Among these applications, one of the most promising is the use of autonomous AI agents to extract structured information from HTA reports and consolidate evidence across multiple European jurisdictions.

Our recent proof-of-concept presented at ISPOR 2026 demonstrated that autonomous LLM-based agents can successfully extract not only traditional Population, Intervention, Comparator, and Outcome (PICO) elements, but also context-specific HTA evidence — including methodological requirements, reasons for rejecting comparators or outcomes, and country-specific critiques — from reports written in Spanish, Dutch, and French. Expert-guided prompt refinement substantially improved extraction quality while reducing hallucinations, supporting the feasibility of multilingual AI-assisted HTA landscape analyses.

While these findings are encouraging, they also raise a much more important question. As AI agents move beyond information extraction and begin supporting evidence interpretation across the European Union, how much does language fluency influence what these systems actually understand?

This question extends beyond translation accuracy. HTA reports do not merely describe clinical evidence; they interpret it. They communicate uncertainty, methodological concerns, regulatory expectations and value judgments that are deeply embedded in language. An agent that correctly extracts every PICO element but misunderstands the reasoning behind an HTA recommendation may produce outputs that are technically complete while scientifically misleading.

Recent work from Stanford University, the University of Pretoria, and The Asia Foundation highlights that this challenge is not unique to healthcare. Their report, Mind the (Language) Gap, argues that modern LLMs consistently underperform in many non-English and lower-resource languages because of limitations in both the quantity and quality of language data used during training. More importantly, they emphasize that machine translation frequently fails to preserve contextual and cultural meaning, limiting downstream reasoning rather than simply reducing translation quality.

For HTA, this observation has profound implications.

 

From multilingual extraction to multilingual interpretation

Extracting structured information from multilingual reports is fundamentally different from interpreting the regulatory reasoning contained within those reports.

PICO extraction is, in many ways, a well-defined information retrieval task. Population characteristics, interventions, comparators, and endpoints are generally presented explicitly within HTA documents. Modern LLMs have become increasingly proficient at identifying these structured elements, particularly when guided through carefully engineered prompts and expert-designed extraction frameworks.

The challenge begins once the agent moves beyond factual extraction.

National HTA agencies rarely communicate their conclusions using simple statements such as “the evidence is insufficient.” Instead, reports contain nuanced discussions about methodological uncertainty, external validity, endpoint relevance, comparator selection, indirect treatment comparisons and clinical meaningfulness. These discussions frequently rely on language that reflects national regulatory traditions as much as scientific evidence.

An AI agent must therefore perform a fundamentally different task. Rather than recognizing predefined entities, it must infer regulatory intent.

This distinction is likely to become increasingly important as AI systems evolve from supporting individual HTA reviews toward consolidating evidence across multiple European countries for EU JCA preparation.

 

Regulatory fluency: A missing dimension of AI evaluation

Current discussions surrounding multilingual AI typically focus on language coverage. Vendors often report that their models support dozens or even hundreds of languages, creating the impression that multilingual capability is largely a solved problem.

However, language support should not be confused with regulatory fluency.

Regulatory fluency can be defined as the ability of an AI system to correctly interpret the methodological reasoning, evidence judgments, and decision-making context expressed within a regulatory document written in a particular language.

This distinction is particularly relevant in HTA.

Two reports evaluating the same oncology therapy may arrive at similar reimbursement recommendations while expressing very different methodological concerns. One agency may emphasize uncertainty regarding overall survival maturity. Another may question the relevance of surrogate endpoints. A third may accept the evidence but reject the proposed comparator because it does not reflect national clinical practice.

Although these reports may ultimately lead to similar conclusions, they represent different regulatory reasoning pathways.

An AI system that reduces all three to a generic statement such as “insufficient evidence” has failed to preserve the information that matters most.

 

Language carries regulatory meaning

Language in HTA is rarely neutral.

Expressions describing uncertainty, clinical benefit, methodological limitations or evidence maturity often carry highly specific meanings that extend beyond their literal translation.

For example, one agency may state that additional clinical benefit has not been demonstrated, while another concludes that the available evidence remains associated with considerable uncertainty. Although both statements may ultimately influence reimbursement decisions, they describe fundamentally different evidence gaps.

Similarly, terminology surrounding comparators varies considerably across Europe. Some agencies discuss accepted standards of care, while others refer to appropriate comparators, relevant comparators or clinically meaningful comparators. These terms are not always interchangeable.

The same applies to discussions surrounding external validity, indirect comparisons, subgroup analyses and patient-reported outcomes. The wording itself often reflects national methodological frameworks developed over decades of HTA practice.

Consequently, multilingual AI agents must understand not only vocabulary but also the regulatory context in which that vocabulary is used.

 

Why machine translation may not be enough

One intuitive solution is to translate every HTA report into English before analysis.

Indeed, machine translation has become remarkably accurate over the past decade and represents one of the major approaches proposed to overcome language resource limitations. However, the Stanford report highlights important limitations of this strategy. Machine translation frequently loses contextual knowledge, introduces unnatural linguistic patterns known as “translationese,” and may flatten language-specific connotations that are essential for downstream reasoning.

For everyday communication these limitations may be acceptable.

For regulatory evidence synthesis they may not.

An EU JCA workflow may involve dozens of HTA reports produced by agencies across Europe. If each report is translated before evidence extraction, small semantic shifts may accumulate across the pipeline. A subtle distinction between methodological uncertainty and evidence inadequacy may disappear during translation. Criticism directed toward endpoint selection may become indistinguishable from criticism of study design. Regulatory nuances embedded within the original language may be replaced by more generic English expressions.

The final evidence synthesis may therefore appear coherent while masking important differences between national assessments.

This phenomenon deserves far greater attention than it has received to date.

 

Lessons from our multilingual HTA agent

Our ISPOR proof-of-concept provides an interesting perspective on this challenge.

The objective of our work was not to compare language performance, but to evaluate whether autonomous AI agents could extract both traditional PICO elements and context-specific HTA evidence across multilingual oncology reports. The agents successfully completed extraction across reports written in Spanish, Dutch, and French, with prompt refinement substantially improving overall performance and reducing hallucinations.

Interestingly, performance was not identical across reports. The French HTA document produced the highest overall extraction accuracy, while hallucinations were observed only within the Spanish report when using the initial prompt.

These observations should not be interpreted as evidence that current LLMs perform better in French than Spanish. The study evaluated only three reports, and numerous factors — including document structure, writing style, and complexity — may have influenced performance.

However, the findings suggest an important hypothesis for future research.

Differences in multilingual extraction performance may not solely reflect prompt design or document characteristics. They may also reflect differences in the underlying linguistic fluency of the model within specific regulatory languages.

Testing this hypothesis systematically across additional European languages represents an important next step for AI research in HTA.

 

The next challenge for EU JCA

The European HTA Regulation fundamentally changes the role of multinational evidence synthesis.

Rather than preparing country-specific submissions independently, manufacturers increasingly need to understand how multiple HTA bodies evaluate similar evidence, where methodological expectations converge, where they diverge, and how these perspectives can inform future JCA strategies.

AI agents are exceptionally well positioned to support this work.

However, future systems will need capabilities extending well beyond multilingual extraction.

They must identify recurring methodological concerns across countries, distinguish between evidence uncertainty and evidence rejection, compare national perspectives on comparators and endpoints, recognize implicit critiques, and consolidate these insights without losing the regulatory meaning embedded within each language.

In other words, they must become regulatory interpreters rather than multilingual translators.

 

Looking forward

The conversation surrounding multilingual AI has largely focused on language coverage, translation quality, and benchmark performance. For health technology assessment, these metrics are no longer sufficient.

As AI becomes integrated into evidence synthesis and EU Joint Clinical Assessment workflows, evaluation should increasingly measure regulatory fluency — the ability to preserve scientific intent, methodological reasoning, and regulatory context across languages.

Our proof-of-concept demonstrated that autonomous multilingual agents can extract structured HTA information across European reports. The next challenge is ensuring that these systems understand not only what an HTA agency wrote, but why it wrote it.

Ultimately, the success of AI in HTA will not be determined by how many languages a model supports. It will be determined by whether those languages preserve the evidence, reasoning, and regulatory judgment upon which healthcare decisions depend.

 

Interested in learning more?

Read our ebook, “Navigating EU JCA: Submissions and Local HTA Decision-Making”:

Download your copy today!
Subscribe to our newsletter

Manuel Cossio

Head of AI Solutions, Real-World Evidence, Value, and Access

Manuel Cossio is Head of AI Solutions, Real-World Evidence, Value, and Access at Cytel. Manuel is an AI engineer with over a decade of experience in healthcare AI research and development. He currently leads the creation of generative AI solutions aimed at optimizing clinical trials, focusing on hierarchical multi-agent systems with multistage data governance and human-in-the-loop dynamic behavior control.

Manuel has an extensive research background with publications in computer vision, natural language processing, and genetic data analysis. He is a registered Key Opinion Leader at the Digital Medicine Society, a member of the ISPOR Community of Interest in AI, a Generative AI evaluator for the EU Commission, and an AI researcher at UB-UPC- Barcelona Supercomputing Center.

He holds an M.Sc. in Translational Medicine from Universitat de Barcelona, a Master of Engineering in AI from Universitat Politècnica de Catalunya, and a M.Sc. in Neuroscience from Universitat Autònoma de Barcelona.

Read full employee bio

Claim your free 30-minute strategy session

Book a free, no-obligation strategy session with a Cytel expert to get advice on how to improve your drug’s probability of success and plot a clearer route to market.

glow-ring
glow-ring-second