Real-World Data and RCTs: The Methodological Debate Over Routine Data

Data from routine medical practice is becoming increasingly important in research. But while enthusiasm for “big data” is growing, experts and institutions such as the IQWiG are calling for methodological caution. A debate.

Photo: Generated by AI
Hanna Sachse
June 17, 2026
IQWiG; BMJ; Roche

What It's All About

Not every data source is a substitute for controlled clinical research. Yet the promise is compelling: Data is generated every day throughout the healthcare system—during doctor’s visits, through digital health applications (DiGA), during hospital stays, or via billing records. This “routinely collected data” (RCD) is available in large quantities and is less expensive to obtain than traditional studies specifically designed to address a research question. However, there is growing concern in the scientific community that the desire for quick results is obscuring the focus on methodological quality. A recent article in the British Medical Journal (June 2026), to which Tim Mathes of IQWiG, among others, contributed, provides context on the topic.

But first: What exactly are RWD and RCD?

  • RWD (Real-World Data): This is not a scientifically standardized term, but rather a descriptive umbrella term for all data generated outside the controlled setting of clinical trials (RCTs). It is a very broad concept and has been criticized by expert bodies such as the IQWiG due to its vagueness and its often promotional (“marketing-heavy”) use.
  • RCD (Routinely Collected Data): This is the more scientifically precise term. It refers to a specific subset within RWD. The defining characteristic is its origin: these data are “byproducts” of administrative processes and clinical documentation requirements.

‍

The IQWiG's Position: Methodological Excellence Over Data Volume

The IQWiG takes a very clear stance in this debate: While observational studies based on routine data are a necessary complement (for example, for rare events or long-term effects following approval), they cannot replace randomized controlled trials (RCTs).

The reason is methodological: In an RCT, randomization ensures that patients are comparable. With routine data, however, the choice of treatment is usually targeted: a doctor prescribes a medication based on specific symptoms or preexisting conditions. When comparing the outcomes of these patients, one often measures not the drug’s efficacy but rather the differences in the baseline characteristics of the patient groups. Without complex adjustment procedures, this leads to massive distortions, known as “bias.”

The Role of AI in a Broader Context

Artificial Intelligence (AI) acts as an “amplifier.” On the one hand, machine learning provides tools for cleaning large, heterogeneous datasets and identifying more complex, nonlinear relationships within the data. On the other hand, researchers warn of risks:

  • Statistical “shortcuts”: AI models tend to identify correlations in the data that are based on administrative processes (e.g., “Testing is more common in better-equipped hospitals”) rather than on medical effects.
  • The "black box" problem: Because many algorithms do not disclose their decision-making processes, methodological errors or discrimination against certain population groups often go unnoticed.
  • Overinterpretation: AI is no substitute for a study design. Without clinical expertise that understands the context of the data, algorithms may produce results that are statistically significant but clinically irrelevant or even misleading.
The debate is not an “either/or” situation

Experts in the field, including those at IQWiG, see the solution in a hybrid strategy:

  1. Strengthening RCTs: Greater efforts should be made to conduct traditional studies in everyday clinical practice in a way that is simpler, faster, and more cost-effective.
  2. Targetedly closing gaps in the evidence: Routine data should be used primarily in areas where prospective studies are hardly feasible (e.g., rare diseases, benefit assessment after market entry).
  3. Methodological rigor: Analyses of routine data require interdisciplinary collaboration.
  4. Methodological standards must remain high to prevent erroneous decisions in health care.
"Whether they are statisticians, clinicians, epidemiologists, HTA experts, or AI experts—they all agree: The use of routine data presents unique challenges..."
Tim Mathes, Head of the IQWiG Health Economics Division

‍

He adds that “targeted strategies are needed to make effective use of them. We show that in many cases, conducting an RCT that uses routine data is the better approach.”

If you'd like to know more

An Overview of the Key Points

#Data source: Routine data (RCD) are administrative byproducts and are not data collected primarily for research purposes.

#Methodological Risk: The main problem is “bias” (systematic distortion). Since there is no randomization, the results often cannot be interpreted causally.

#IQWiG Position: RCTs remain the gold standard. Routine data provide valuable supplementary information but are no substitute for methodologically sound evidence generation.

#AI as a Tool: AI can improve data quality and identify complex patterns, but it carries the risk of taking clinically nonsensical “shortcuts” through black-box decisions.

‍#Interdisciplinarity: A thorough analysis of routine data absolutely requires the integration of epidemiology, clinical experience, and statistics. AI alone is not sufficient as a methodological foundation.

RWD encompasses an enormous range of data sources:

  • Administrative data: Health insurance billing data, discharge summaries, DRG data.
  • Electronic Health Records (EHR): Documentation from outpatient and inpatient care.
  • Registry Data: Specific disease registries that collect data according to a specific protocol.
  • Patient-Generated Health Data (PGHD): Information from wearables, fitness trackers, health apps (DiGA), or Patient-Reported Outcome Measures (PROMs).

The key characteristics of the RCD are:

  • Not a research design: The data were not collected primarily to test a specific clinical hypothesis, but rather to document care or ensure billing.
  • Lack of standardization: Because they are optimized for administrative purposes, they often lack the controlled measurement conditions necessary for research.
  • Limited control: Researchers have no influence over which variables were collected, how often they were measured, or who coded the data and according to what criteria.

The representativeness problem (bias):

Routine data provide a distorted picture of the population:

  • Systematic underreporting: Minorities, low-income groups, and rural populations are underrepresented and are also less likely to receive diagnostic testing.
  • Fragmentation: These groups often use different institutions, which results in incomplete data sets that are difficult to consolidate.
  • Risks Associated with Data Integration
  • False comparability: Datasets are often combined with inappropriate control groups (e.g., young children as controls for the general patient population), which skews predictive models due to statistical artifacts.

BMJ Publication: Sabine Hoffmann et al. (2026). Using routinely collected data for research purposes: challenges and mitigation strategies. BMJ 2026;393:e087812. Published online June 2, 2026.

More Articles
Deepfake Medical Professionals: “You want real truth in advertising”
July 30, 2026
Agent-Based AI in Medicine: Potential, Practice, and Guidelines
July 23, 2026
Pharmaceutical Communication: The Celebrity Factor vs. Patient Trust
July 9, 2026

Kicking Off the Industry Dialogue: The New Forward Pharma Podcast

Kicking Off the Industry Dialogue: The New Forward Pharma Podcast
August 10, 2026
Listen now
Listen to the podcast
"FORWARD PHARMA"
analyzes the key trends in the healthcare and pharmaceutical industries. We provide context for current topics and speak with experts.
Click here for the
FORWARD PHARMA
Podcast