Skip to main content
An official website of the United States government

Questions Answered: Data Commons in Cancer Research

, by Tanja Davidsen, Ph.D.

In this series, NCI CBIIT experts answer commonly asked questions (via search engines and generative AI platforms) about technology and data in cancer research. So, whether you're a researcher wanting to better understand computational approaches, or a data scientist wanting to learn how your expertise can accelerate discovery, this blog series is for you!

In this blog, Dr. Tanja Davidsen answers questions about data commons and their use in cancer research. Dr. Davidsen oversees the NCI team developing NCI’s Cancer Research Data Commons (CRDC), a secure, cloud-based, data science infrastructure that helps researchers submit, share, access, and analyze petabytes of data.

Question: How do data commons differ from data repositories? 

Answer:

A data commons offers more than your standard data repository. In both data repositories and data commons, you can search across data sets and possibly access that data through a web interface. You may also be able to interact with the data programmatically through an Application Programming Interface (API). But, in a data commons, you have an extra layer on top of that.
For example, in NCI’s CRDC, one of our “extra layers” is that we link to our cloud resource, which is hosted on the Seven Bridges Cancer Genomics Cloud (SB-CGC), powered by Velsera. That way, you can analyze all the data that CRDC makes available in the cloud. So, our cloud resource provides you with the availability to analyze the data without having to download it to a local compute environment. Additionally, some of the resources in the CRDC provide an additional level of functionality for users. For example, our Genomic Data Commons offer specific genomic tools to analyze their data. They recently launched a new Correlation Plot Tool that allows users to correlate GDC molecular data with patient clinical and survival data. 

Question: What data commons are available at NCI?

Answer:

Currently, we have seven data commons available to the public. I’ll list them by date of release:

  • Genomic Data Commons: The Genomic Data Commons, or GDC, is the oldest and most established of our data commons. It holds genomic data, including important data sets like The Cancer Genome Atlas (TCGA) and the Therapeutically Applicable Research to Generate Effective Treatments (TARGET). The powerful thing about the GDC is that all the data has been harmonized with the same human genome standard and variant calling pipeline. So, everything within GDC is apples-to-apples.
  • Proteomic Data Commons: The Proteomic Data Commons, or PDC, focuses on proteomic data like mass spectra, peptide, and protein data. It holds the proteomic data from large-scale studies like NCI’s Clinical Proteomic Tumor Analysis Consortium (CPTAC) and the Applied Proteogenomics OrganizationaL Learning and Outcomes (APOLLO) program, among others.
  • Imaging Data Commons: We also have the Imaging Data Commons, or IDC, which focuses on DICOM medical images including CTs, MRIs, and PET scans. You can find the imaging data for TCGA and CPTAC here, as well as data from the Childhood Cancer Data Initiative (CCDI) and Human Tumor Atlas Network (HTAN).
  • General Commons: In addition, we have the General Commons or GC. This is for data that does not otherwise fit another CRDC data commons. Much of the data in our GC is actually genomic data that the GDC was unable to accept. The GDC’s harmonization step is relatively expensive because they realign all the BAMs and recall all the variants (a preprocessing step so the GDC results are consistent and comparable). So, the GC offers that place for “as-submitted" genomic data as well as other data types.
  • Integrated Canine Data Commons: Our Integrated Canine Data Commons, or ICDC, hosts data from pet dogs undergoing clinical trials for their cancer. For example, you can find transcriptomic data for canine patients from NCI’s Comparative Oncology Program (COP).
  • Clinical and Translational Data Commons: Our Clinical and Translational Data Commons, or CTDC, hosts data from NCI’s human clinical trials and immune-oncology studies. This includes the clinical and molecular data from NCI’s Cancer MoonshotSM Biobank.
  • Population Science Data Commons: The Population Science Data Commons, or PSDC, our seventh data commons, was launched recently in March and will focus on epidemiology studies. There, you can find data from NCI’s Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial.

Question: How has artificial intelligence (AI) impacted data commons? 

Answer:

AI is definitely making an impact on our data commons as it is doing across all technology. As an example, I’ll talk a bit about an effort that we’ve been involved in for the last couple of years with Advanced Research Projects Agency for Health (ARPA-H). ARPA-H is similar to Defense Advanced Research Project Agency, but for health, supporting bold, innovative programs, that aim to transform the field. In 2023 NCI and ARPA-H worked together to come up with a new program called the Biomedical Data Fabric Toolbox (or “BDF Toolbox”). ARPA-H’s model is to have a short couple of years of funding for programs that are revolutionary, not evolutionary. So, the idea is to take a risk and push the field forward. A number of ARPA-H projects funded for BDF Toolbox integrate AI with CRDC APIs and data. Because ARPA-H funds things for a relatively short amount of time, successful completed projects transition to partners so that the community can benefit. As an ARPA-H transition partner, NCI plans to integrate several AI tools coming out of the BDF Toolbox into the CRDC. 

For example, in the future, you’ll be able to search across all CRDC data from a convenient chat box. You can currently search for data through each of the data commons’ websites or through their APIs. So, these new AI tools will allow us the ability to offer a way to search across all data sets using natural language. Additionally, the AI tools will give you more ways to visualize the data. So, if you can think of what you want, you can ask the AI to do it and get that result. We can also use AI tools to help clean up and harmonize the data that researchers submit to CRDC.

There are many different applications of AI that will make CRDC more powerful, but I want to add a note of caution. AI is a useful tool, but you should always double-check what it is telling you. 

Question: What are the benefits of data commons to a cancer researcher?

Answer:

If you’re a researcher interested in finding out what data exists and what you have access to, then you’ll find benefit in using a data commons. Our data commons help researchers access a wide variety of data types and data sets from large and important programs. As I shared, you can find data from TCGA, TARGET, HTAN, APOLLO, and others in our CRDC. We also strive to be interoperable with similar data hosted in data commons across NIH. Our CRDC really focuses on cancer data, but you may also be interested in the GTEx data from NHGRI’s AnVIL or the TOPMed data from NHLBI’s BioData CATALYST platform. We’ve worked with these platforms to allow connections between the data sets, so you can analyze these data without having to move them. You can even analyze these, and other data sets beyond the CRDC, in the Seven Bridges Cancer Genomic Cloud. 

Question: What are best practices for managing large-scale biomedical data sets? 

Answer:

First, it’s important we have well-structured data. I think everyone has heard the phrase “garbage in, garbage out.” If you don’t have harmonized data, it will be impossible to search across all the data sets, or that search won’t be meaningful. We are working on a standard data submission portal where you can request and eventually submit the data to CRDC. With this data portal, we will set the same standards to make sure that all data coming can be successfully searched.
Secondly, we have standards for the data we store. We have about 17 petabytes of data right now, and that number is just going to get larger. So, we evaluate which data is utilized the most or which data is not being accessed. That allows us to prioritize storage so that the data sets most used by the research community are available.  

Have another question?

If we missed your question about data commons or NCI’s CRDC, email NCI CBIIT. We’ll connect you with a contact. If it is commonly asked question, we’ll update the blog with an answer.

Interested in more Q&A Blogs?

If you enjoyed this blog, check out the remainder of the series. You'll find interviews with CBIIT staff who share their expertise in cancer research-related technology, informatics, and more.

Author

Tanja Davidsen, Ph.D.
As the chief of NCI CBIIT’s Data Ecosystems Branch, Dr. Tanja Davidsen leads the team developing the CRDC. She is also a biomedical informatics specialist with research interests focusing on cancer, infectious diseases, biological databases, cloud applications, and large-scale genomic analysis.
 

< Older Post

Questions Answered: Generative AI in Bioinformatics

Email