Clinical data standards are designed to bring consistency. SDTM, in particular, gives us a common structure for organizing clinical trial data so that studies can be reviewed, submitted, integrated, and analyzed with a shared language.
But anyone who has worked across multiple studies knows the uncomfortable truth: using the same standard does not always mean implementing it the same way!
That was the central idea behind my CDISC Europe Interchange 2026 presentation, “Originality Matters: Preserving Standards Integrity from Clinical Data Models to Tiramisù.” The title may sound playful, but the message is serious. Standards matter. Their interpretation matters even more.
And yes, tiramisù has something to teach us about SDTM.
There is a standard, and then there are variations
Tiramisù has a traditional recipe. It is not strawberry tiramisù. It is not pumpkin tiramisù. It is not a deconstructed dessert in a glass with a familiar name attached to it. But most importantly, it is alcohol-free. Any additions belong in a “supplemental” dish, not in the original recipe. If you would like some sweetness, enjoy a glass of dessert wine on the side.
The same is true for clinical data standards. A standard defines structure, expectations, and intent. But in practice, people adapt. They interpret. They extend. Sometimes those extensions are justified. Sometimes they are simply “local habits” repeated across studies.
In SDTM terms, those “variations” often show up as supplemental qualifiers, non-standard domains, inconsistent controlled terminology, or different choices in mapping the same concept. One study may look perfectly reasonable in isolation. But when you look closely at hundreds of studies, patterns begin to emerge.
That is where metadata repository becomes powerful!
From qualitative review to data-driven assessment
Previous SDTM quality review efforts, including my own presentation from around seven years ago, were largely based on manual assessment of submission packages. Similar approaches have also been used by other presenters and, in a regulatory context, by FDA reviewers. These reviews were valuable because they highlighted recurring implementation issues.
The next step was obvious: could we move from “anecdotal” knowledge to measurable evidence?
To explore that question, we built an SDTM metadata warehouse using historical SDTM metadata from more than 350 anonymized SDTM packages Cytel has either developed or received from other vendors. The goal was not to analyze patient-level data. Instead, we looked at metadata and aggregated data: domains, variables, QNAMs, labels, test codes, units, breakdown by SDTM IG versions, and study characteristics such as study phase (Figure 1).
This allowed us to move from “we often see this issue” to “we can quantify how often this issue appears.”
Figure 1: Metadata-Warehouse Retrospective Build Plan
What metadata at scale can tell us
A single study can tell you whether one submission package is internally consistent. A portfolio-level metadata warehouse can tell you something much broader: how SDTM is actually implemented in practice.
The analysis was designed to answer practical questions such as:
- How many SDTM domains are typically present in a study?
- Which domains form the core SDTM backbone?
- How often are newer domains adopted after they are introduced?
- Which standard domains most often require supplemental qualifiers?
- Which non-standard domains appear repeatedly?
- Where do laboratory units, test codes, or mappings vary across studies?
These questions are difficult to answer from theory alone. They require a view across many real implementations.
One clear finding was that SDTM is structurally consistent, but content is much less consistent. Core domains such as demographics (DM), adverse events (AE), concomitant medications (CM), exposure (EX), disposition (DS), and laboratory (LB) data are widely used. However, the way content is represented within and around those domains varies substantially.
Supplemental qualifiers usage: A signal of implementation drift
Supplemental qualifiers (SUPP) are sometimes necessary. They provide flexibility when a standard domain does not include every variable required for a specific study or therapeutic area.
But SUPP usage can also be a signal. When the same parent domains repeatedly require supplemental qualifiers across many studies, it may suggest that implementation practice is drifting beyond the standard structure, or that the standard itself may need to evolve (Figures 2 and 3).
In the metadata analyzed, AE, CM, and DM showed especially strong dependency on supplemental qualifiers, with SUPPs present in more than 85% of studies where those parent domains existed. AE was particularly consistent in this pattern, with SUPPAE appearing in the vast majority of AE studies.
This does not mean SUPP is wrong. It means SUPP usage deserves attention. At scale, it becomes a quality signal, a standardization signal, and potentially an input into future standards development.
Figure 2: Twenty Years of SDTM Implementation Guidance
Figure 3: Usage of New SDTM Domains Since They Were Introduced
Laboratory data: Standard structure, variable content
Laboratory data provided another useful example.
The LB domain is familiar to everyone working with SDTM. It is one of the most common and important domains. Yet laboratory metadata revealed substantial variability in units, mappings, and representation of similar concepts.
Some parameters showed multiple Standard International (SI) unit choices. In some cases, different unit expressions represented the same meaning. In other cases, parameters such as creatinine clearance or aPTT showed considerable variation across studies (Figure 4).
Laboratory data also show how implementation can drift: the same or similar results may appear in LB, MB, or IS across different studies, even though FDA expectations and CDISC Controlled Terminology provide guidance on where those data should be represented (Figure 5).
This is not just a cosmetic issue. Laboratory data inconsistency can affect review, integration, automation, and downstream analytics.
Figure 4: SI Consistency
Figure 5: LB vs IS vs MB
Why this matters for submissions, integration, and analytics
Conformance checks are essential, but they do not replace expert review, metadata review, or sponsor-level consistency assessment.
For integration, even when studies follow SDTM, harmonization effort is still required. Domain structures may align, but content-level metadata such as TESTCD, QNAM, QLABEL, units, labels, and supplemental qualifiers can differ significantly.
For analytics, metadata is critical. A metadata warehouse can support benchmarking, pattern detection, outlier identification, and quality monitoring. It can help identify missing expected domains, unexpected unit variability, rare test codes, inconsistent mappings, or repeated non-standard practices.
In other words, metadata is not administrative residue. It is reusable intelligence!
From metadata review to reusable CDISC intelligence
The most important takeaway from this work is simple: SDTM is standardized. Implementations are not!
That is not criticism of SDTM. It is a realistic observation about how standards operate in the real world. Structure is often consistent, content varies. Domain usage, supplemental qualifiers, terminology, units, and mapping decisions reflect study design, sponsor practice, therapeutic area needs, SDTM IG version, and sometimes habit.
By analyzing metadata at scale, we can see patterns that are invisible in a single study. We can identify where standards are working well, where implementation is drifting, and where future guidance may be useful.
This approach also opens the door to the next opportunity: applying the same metadata-driven analysis to ADaM. If SDTM metadata can reveal implementation variability in domains, QNAMs, units, and test codes, ADaM metadata may reveal similar insights in Analysis Parameters (or concepts), derivations, analysis flags, traceability, and endpoint standardization.
The future of standards is not only about defining models. It is about learning from how those models are used.
Standards do not fail. Implementations drift. And metadata may be one of the best tools we have to bring them back into alignment.
Stay tuned, more insights to come from SDTM Metadata Warehouse.
Interested in learning more?
Download your copy of Angelo’s ebook, The Good Data Submission Doctor on Data Submission and Data Integration to the FDA:
Download your copy today!Subscribe to our newsletter
Angelo Tinazzi
Senior Director, Statistical Programming, Clinical Data Standard & Submission
Angelo Tinazzi is Senior Director, Statistical Programming, Clinical Data Standard & Submission, at Cytel. Angelo is a well-published and recognized expert in statistical programming, with over 25 years’ experience in clinical research. In particular, his core expertise lies in the application of CDISC standards across different therapeutic areas, such as data submission to health authorities like the FDA and PMDA.
As well as being an authorized CDISC instructor, Angelo is former member of the CDISC European Committee, and co-lead of the Italian-speaking CDISC User Network. Angelo is also conference co-chair for PHUSE EU Connect 2026 and conference chair for PHUSE EU Connect 2027.
Prior to joining Cytel, Angelo worked at Merck Serono, SENDO Foundation, Phamarcia & Upjohn, Simbologica SAS Quality Partner, the UK Medical Research Council, and the Institute for Pharmacological Research “Mario Negri.”
Read full employee bioClaim your free 30-minute strategy session
Book a free, no-obligation strategy session with a Cytel expert to get advice on how to improve your drug’s probability of success and plot a clearer route to market.




