Harmonizing Data Collection for Precision Oncology: AACR Project GENIE® Data Model

Precision oncology relies on large amounts of data, often combining multiple datasets from different institutions, to produce reliable, representative, and clinically meaningful results to guide cancer prevention, diagnosis, and treatment. Despite standardization efforts, inconsistencies in the practices of data capturing, definition, and storage persist and represent a big challenge for researchers.

In an attempt to simplify data harmonization, the AACR Project GENIE® (Genomics Evidence Neoplasia Information Exchange) consortium created a new data model, called GENIE Data Model (GDM), an open-access framework that helps standardize longitudinal clinical and genomic data collection across solid tumors at different stages of a patient’s cancer journey from diagnosis through long-term outcomes. 

GDM, which was the result of a collaboration among subject matter experts from academia and the government and included the patient advocate perspective, was released publicly on April 28 and presented during a Methods Workshop at the AACR Annual Meeting 2026. The model was also described in a paper published in Cancer Research Communications.

What Is a Data Model?

Jennifer Hoppe, MPH

A data model is a blueprint that defines what data will be collected, how the data will be organized, and the rules for storing and using the information in a consistent way.

“There is a lot of clinical information that can be repurposed from study to study,” said Jennifer Hoppe, MPH, a senior clinical project manager at Project GENIE and the first author of the paper. “Suppose you are an academic team that is conducting a research project for which you need to collect patient data across different clinical studies, such as demographics, cancer diagnosis, what drugs the patients received, and when their treatment started and ended. Rather than the researchers having to build a model for how they collect all this information for every single study, we have done that work in advance and built a model every research team can directly use to start entering data right away. I think most researchers would find this work really helpful.”

Why a New Model?

“Many data models already exist, each with its own complexity and built with a specific purpose in mind,” said Hoppe. “The GDM was built to be applicable to many different solid tumors and different use cases while aligned to currently used standards.”

During the Annual Meeting session, chair Jeremy L. Warner, MD, a professor of medicine at Brown University, who was one of the co-lead principal investigators on the GDM project, discussed the importance of data harmonization. He pointed out that oncology data, particularly real-world data, are still fragmented, with incompatible electronic health records, mismatches and different nomenclatures in genomic annotations, inconsistently recorded outcomes, and a lack of granular longitudinal data.

The GDM was created with these challenges in mind. It was designed as a modular, tumor-agnostic core, into which disease-specific extensions can be plugged in as needed, and to be extensible and scalable in the future, Warner said. He emphasized that the GDM model is about breaking down data silos and enabling a connected ecosystem for oncology research.

The work to develop the GDM builds on more than a decade of experience of the Project GENIE consortium in assembling and growing the Project GENIE registry of real-world genomic data through data sharing among 20 leading international cancer centers.

“The GDM was specifically designed to facilitate the collection of more key clinical variables programmatically, allowing GENIE to get closer to our goal of documenting complete patient cancer experiences for every patient in the registry,” said Shawn M. Sweeney, PhD, senior director of the AACR Project GENIE Coordinating Center. “Ultimately, the GDM lays the foundation for fully federating the registry in the near future.”

Hoppe likened the registry to an iceberg. “What we have available is just the tip, but there is all that ice submerged still, all that information that exists on which we don’t have visibility yet because we can’t collect it,” she said. “But as we deploy this model and we are able to collect more data, eventually we will see more and more of the full picture of that iceberg for all the patients in the GENIE registry.”

The Development of the GENIE Data Model

The researchers reached out to members of the oncology community to assemble a team of 13 subject matter experts who were divided into four working groups. Hoppe explained that their first task was to establish which essential variables need to be collected to characterize a tumor exhaustively. These variables were organized into four content areas or clinical domains, spanning initial diagnosis through treatment and long-term outcomes, and each working group focused on its area of expertise: diagnosis, treatment, outcomes, or person-centered metrics, the latter including demographic variables, social determinants of health, and comorbidities.

The GDM organizes the data related to a patient’s cancer journey into four clinical domains: diagnosis, treatment, outcomes, and person-centered metrics.

“Using their expert opinion, the working groups really helped shape what we would need to include so that the model is not composed of thousands of variables but encompasses the core group of variables that can give a reliable picture for almost any solid tumor,” Hoppe explained. “We wanted the model to be widely applicable, but also specific enough to provide the information a researcher would need.”

Once the experts defined the variables and the corresponding parameters and value sets, the next step was to build that information into a database collection tool. Specifically, the team used REDCap (Research Electronic Data Capture), a specialized web-based software platform, but the GDM can be deployed within any data collection tool, added Hoppe.

One important goal for the GDM was interoperability, a term that describes the ability of different information systems to communicate and exchange data seamlessly. In this case, each data variable in the GDM has a shared definition and parameters that are aligned to existing standards and terminologies and captured with the same value set.

In addition, the GDM can work both with manually entered data and with data that have already been collected and stored in a structured way (for example, in Excel files or enterprise data warehouses), to minimize the curation effort.

“Eventually, in the future, we would like to have 100% structured data supporting the Project GENIE registry, but since we are not there yet, we needed to develop something that was flexible enough to be able to incorporate both manually collected and structured data,” said Hoppe.

Deploying the Model

The GDM is hosted in a repository on the cloud-based platform GitHub and can be accessed through a Creative Commons license. Researchers have access to three components: the database file, which can be loaded into REDCap or other data collection tools; an Excel file that contains the data dictionary, which provides an-easy-to-read overview of the variables and corresponding definitions and value sets; and a curation directives manual, or data completion guideline, which provides additional details to ensure homogenous use of the model across institutions.

Limitations and What Comes Next

Since it was designed as a large pancancer model, GDM may not be entirely specific for each research project, said Hoppe. However, the team built it to be modifiable, so that the users can add additional fields as needed to accommodate their research questions.

Shawn M. Sweeney, PhD

Another limitation is that the model is currently only suitable for solid tumors and cannot be utilized for studies on hematologic malignancies, she added. “We are working on an expansion that we will hopefully be able to deploy within the next year or two.”

The team is also developing packages of prebuilt algorithms to be applied by end users to perform calculations on the data collected by the model and obtain new, derived variables to create an analytic dataset. “For example, one could use a provided algorithm to calculate the duration of a treatment from the start and end dates of that treatment,” said Hoppe.

“The community can expect annual updates, and we hope end users will share any derivatives they create back with the broader research community so that we can all get to cures as quickly as possible,” concluded Sweeney.

AACR Annual Meeting sessions are available for virtual viewing for all registered attendees through October 2026.