Decoding the Data Ecosystem Institute (DeCoDE Institute) 

Decoding the Data Ecosystem Institute (DeCoDE Institute) 

Congratulations to the DeCoDE Institute Capstone Event winners!

 

The Decoding the Data Ecosystem (DeCoDE) Institute is a comprehensive virtual institute and capstone project event on open science, reproducible research, knowledge exchange, and skill development through Common Fund Data Ecosystem (CFDE) web portals, tools, and resources. Grounded in FAIR (Findable, Accessible, Interoperable, Reusable) principles, the CFDE is a unified network for discovering, accessing, and analyzing diverse biomedical datasets to accelerate scientific discovery.

From July 28 through July 30, the Training Center hosted the DeCoDE Capstone Event, marking the conclusion of the 2026 DeCoDE Institute. Participants showcased their creativity, collaboration, and hard work marking the culmination of a successful program. Below are the winning projects from this year’s event.

Additional Information

1st Place: Identification of Common Extracellular RNA Biomarkers Across Neurodegenerative and Cognitive Disorders

  • This used CFDE’s exRNA to determine if they could identify extracellular RNA (exRNA) signatures in plasma across Alzheimer's disease (AD), Parkinson's disease (PD), Lewy body dementia (LBD), and mild cognitive disorder (MCD) that could serve as shared biomarkers of neurodegeneration.
  • The team identified differentially expressed exRNAs in neurodegenerative diseases and assessed if the identified exRNAs are associated with key neurological function.
  • Team members included Marangelie Criado-Marrero, Munira Haque, Deepti Jain, Marina Rice, Nur Shahir, Zaynab Shakkour

2nd Place: Metabolomic Profiling of Human Breast Cancer Plasma Using LC-MS to Identify Disease-Associated Metabolic Signatures

  • Breast cancer is a leading cause of cancer-related mortality in women. Imaging and histopathology remain the primary diagnostic tools, but metabolic alterations occur early in tumor development and may offer complementary biomarkers for diagnosis, staging, and prognosis. This analyzed CFDE’s Metabolomics Workbench ST000355 plasma dataset (LC-MS) to identify metabolites and multivariate patterns that distinguish breast cancer patients (Stages I–IV) from healthy controls.
  • Team members included Ujjalkumar Subhash Das, Bidisha Sengupta, Akanksha Gupta

3rd Place: Cardiac Sex-Specific & Time-Course Multi-Omic Response to Endurance Training

  • This used CFDE’s MoTrPAC data to assess if sedentary heart tissue shows sex-specific differences in gene expression and enriched biological pathways, similar to those observed in subcutaneous white adipose tissue.
  • By applying a published sexual-dimorphism and exercise-training analysis framework to a different tissue, the project evaluated whether cardiac gene expression, protein abundance, and biological pathways show shared or tissue-specific patterns.
  • Team members included Upama Roy Chowdhury, Ebuka Ezeika, Xiuqi “Jade” Li, Sarah Myer, Viveka Patil

2026 DeCoDE Session Recap

Couldn't join us for this year's DeCoDE Institute or missed a session you were excited about? No problem! We've put together a recap just for you to catch up on all the highlights and insights.

DATE AND TIME TITLE AND OVERVIEW RESOURCES

June 2

 

Part 1: Overview of CFDE and Data Scavenger Hunt

This session provided a foundational overview of the Common Fund Data Ecosystem (CFDE), highlighting its purpose, structure, and role in advancing biomedical research. Participants engaged in a hands-on Data Scavenger Hunt designed to familiarize them with CFDE resources, data and tools, fostering an interactive and practical understanding of the ecosystem.  

View Recording: https://youtu.be/0660k8DZAhk

View Slides/Session Materials: Part 1: Overview of CFDE and Data Scavenger Hunt

 

June 4

 

Part 2: Introduction to Omics and CFDE Web Portal Exploration

This session introduced participants to the fields of omics (e.g., genomics, transcriptomics, proteomics, metabolomics, epigenomics), covering its significance and applications in biomedical research. Participants also learned about and explored CFDE web portals and tools, gaining insights into its functionalities and learning how to navigate and utilize the web portals for data-driven research. 

View Recording: https://youtu.be/0HGXoreFU5s

View Slides/Session Materials: Part 2: Introduction to Omics and CFDE Web Portal Exploration

 

June 15

 

Part 1: Applied AI for Biologists Using Galaxy and CFDE Cloud Resources

This training introduced participants to Galaxy and the CFDE Cloud Workspace.

View Recording: https://youtu.be/oF2K4H1C0bQ

View Slides/Session Materials: Part 1: Applied AI for Biologists Using Galaxy and CFDE Cloud Resources

 

June 17

 

Part 2: Applied AI for Biologists Using Galaxy and CFDE Cloud Resources 

Building on the foundational knowledge of Galaxy and the Cloud Workspace, the session dove into practical applications by integrating CFDE datasets with AI tools. Participants utilized hands-on practice in leveraging AI-driven methodologies to answer meaningful biological questions and enhance their data processing skills. 

View Recording: https://youtu.be/35R7Acc7_E0

View Slides/Session Materials: Part 2: Applied AI for Biologists Using Galaxy and CFDE Cloud Resources 

 

 

June 30

 

Part 1: R Programming Basics 

This session introduced the fundamentals of programming in R using an interactive R Notebook environment. Participants learned core concepts such as objects, data types, packages, and basic data exploration with basic hands-on exercises. 

View Recording: https://youtu.be/i49_td5F9Z8

View Slides/Session Materials: Part 1: R Programming Basics 

 

 

July 1

 

Part 2: Data Wrangling in R with tidyverse 

This workshop introduced the tidyverse ecosystem in R with a focus on data manipulation using the dplyr package. Participants utilized hands-on real-world scenarios with essential functions for selecting, filtering, transforming, summarizing, and grouping data. 

View Recording: https://youtu.be/IEFb3fodfpo

View Slides/Session Materials: Part 2: Data Wrangling in R with tidyverse 

 

 

July 20

 

Reproducible Data Science with Git and GitHub 

This workshop introduced the fundamentals of version control using Git and GitHub and explains their role in reproducible and collaborative data science. Participants learned how to track changes, document their work, and manage project history using Git within RStudio. 

View Recording: https://youtu.be/VwLlhFkNjmw

View Slides/Session Materials: Reproducible Data Science with Git and GitHub 

 

 

July 22

 

Data Visualization Essentials in R

This workshop introduced the principles of data visualization and demonstrates how to create clear, effective visualizations using R. Participants learned how to structure plots, map data to visual elements, and produce reproducible graphics suitable for research and communication. 

View Recording:

View Slides/Session Materials: Data Visualization Essentials in R

 

 

 

The institute is hosted by the CFDE Training Center managed by ORAU in collaboration with BioData Sage LLC.