Search NASA⌕ Search

SEARCH · Search NASA

Results for “Open Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Developing a Vision for Heliophysics Infrastructure: The LIKED Resource and the DIARieS Ecosystem

Heliophysics data and computational infrastracture are not equipped for 21st science, suffering from holes in the know-how to build better systems. Without a clear vision, efforts to improve the infrastructure have been incremental and incoherent. This poster presents both the vision and the technology required: an online LIbrary KnowledgE and Discovery (LIKED) resource for discovering and implementing knowledge, data, and infrastructure resources; and an online analysis ecosystem to simplify Discovery, Implementation, Analysis, Reproducibility, and Sharing (DIARieS) of scientific results and environments. The LIKED and DIARieS solutions adopt FAIR data principles and the best practices from the budding field of open science. The proposed new infrastructure components will close many of the current gaps in heliophysics’ infrastructure, such as the ability to search for data and knowledge by phenomenon across domains, and to find software and examples relevant to the desired data set (including model data). Further, these components will enable community members to more efficiently use the resources already present and improve upon the content via a community-curated and trusted library. Combining these solutions lowers the barriers to heliophysics resources for all, increasing the return on our investments. Finally, the structure behind these ideas are topic-agnostic, so they are fully extensible to other fields, leading to invaluable connections to other disciplines. Just as with the development and construction of a long-term satellite mission, we must work together as a community to build a vision of the infrastructure that will most benefit the community, and then collaborate to construct, assemble, and test all the necessary pieces individually and as a unit. Our purpose in presenting this work is to not only describe the proposed vision, but also to gather feedback from the community on this topic.

infrastructure↗

Spaceflight Biospecimen and Data Sharing in Support of Science Discovery and Exploration

For decades, NASA and international partners have conducted biological experiments in space to understand effects of spaceflight and address potential hazards. To enable spaceflight back to the Moon, and then to Mars and beyond, it is imperative to further understand basic science and health risks associated with spaceflight, along with developing countermeasures. The sending of experiments and organisms into space is a costly endeavor. To maximize scientific return, sharing with the scientific community both space-flown biospecimens and data from completed experiments is essential. New fundamental, applied, and bioinformatic science insights can be gained from specimen and data sharing efforts. Data reuse enables spaceflight health risk modeling, analyzing adverse outcomes across spaceflight hazards, and deep space autonomous support for the flight medical officer. Space-flown biospecimens not required by mission Principal Investigators are regularly archived and made available for scientific request. The largest biorepository of these samples are found within NASA’s Institutional Scientific Collection at Ames Research Center (ISC-ARC), which stores over 32,000 specimens mostly from Shuttle and International Space Station (ISS) missions, but also some ground-based analog samples. The Ames Life Sciences Data Archive manages the ISC-ARC. Tissues are predominantly from mice and rats, though samples are also available from bacteria and quail. Only a handful of other similar collections exist worldwide. Rodent biospecimens exposed to simulated space radiation at Brookhaven National Laboratory are archived under the purview of NASA HRP Space Radiation Element. Microbial collection and analyses from 20 years of routine environmental monitoring of air, surfaces, and water systems of the ISS were performed to ensure a safe environment for astronauts. Samples from the ISC-ARC, space radiation and microbial collections are searchable and requestable through the NASA Life Sciences Data Archive (LSDA). Decades of planetary protection microbial isolates derived from spacecraft bioburden are archived in JPL’s microbial collection. Rodent biospecimens from spaceflight investigations conducted by the Japan Aerospace Exploration Agency (JAXA) are archived and available at the JAXA Biorepository at Tsukuba Space Center. The Russian Institute of Biomedical Problems also has a collection of animal, microbial, cellular, and fungi available for research from ground analog experiments. Several data repositories exist for scientists to utilize. The LSDA is the primary NASA source of life sciences research data and information. It contains decades of spaceflight and ground-analog research involving human, microbial, cellular, plant, and animal subjects. Data is collected from NASA-funded investigations through the Human Research Program and the Space Biology Program. The NASA Lifetime Surveillance of Astronaut Health collects and grants access to clinical and occupational health monitoring data from astronauts, with a list and description of data collected available for request through the LSDA. NASA GeneLab at ARC collects genomic, transcriptomic, proteomic, and metabolomic data from any species. It is a repository and platform for collaborative open-science bioinformatic approaches. JAXA is establishing an ‘omics-based repository in collaboration with the Tohoku Medical Megabank (ToMMo), called the JAXA-ToMMo Integrated Biobank for Space Life Science. Overall, the sharing of these biospecimen and data resources can assist researchers worldwide in understanding spaceflight effects on biology, along with enabling next generation data science applications for space exploration platforms. Websites: https://lsda.jsc.nasa.gov/ ; https://www.nasa.gov/ames/research/space-biosciences/isc-bsp ; https://www.nasa.gov/ames/research/space-biosciences/alsda

Ryan T. Scott↗

Data Integrity Challenges in NASA Giovanni

The Geospatial Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni) is an online tool developed by the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), one of 12 NASA Science Mission Directorate Data Centers (DAACs) to analyze and visualize NASA remote sensing and model data without downloading data and software. As of this writing, over 2000 Earth satellite and model variables are available in Giovanni, including several well-known NASA satellite missions (e.g., TRMM, GPM) and projects (e.g., MERRA-2, GPCP). There are twenty-two plots provided by Giovanni that can be used to analyze, compare, and explore Earth data across disciplines. Results can be shared with colleagues and downloaded for further analysis. Giovanni has helped publish over 3000 referral papers over the years. As open science policies roll in, data integrity has become a major challenge for Giovanni and other tools. For integrity, both data and workflows must be transparent. FAIR-compliant data, including input, intermediate, and result products, as well as their associated statistics, metadata, and information, are needed. The NASA Data Product Development Guide for Data Producers provides a key resource on how to develop FAIR-compliant data products. Data quality information is also needed from data producers and analysis services like Giovanni. The workflow part is quite challenging and requires workflow management improvements, such as recording workflows and making them available to users. In this presentation, we will discuss the data integrity challenges in Giovanni.

data analysis, visualization↗

The NASA Ames Research Center Institutional Scientific Collection: History, Best Practices and Scientific Opportunities

The NASA Ames Life Sciences Institutional Scientific Collection (ISC), which is composed of the Ames Life Sciences Data Archive (ALSDA) and the Biospecimen Storage Facility (BSF), is managed by the Space Biosciences Division and has been operational since 1993. The ALSDA is responsible for archiving information and animal biospecimens collected from life science spaceflight experiments and matching ground control experiments. Both fixed and frozen spaceflight and ground tissues are stored in the BSF within the ISC. The ALSDA also manages a Biospecimen Sharing Program, performs curation and long-term storage operations, and makes biospecimens available to the scientific community for research purposes via the Life Science Data Archive public website (https:lsda.jsc.nasa.gov). As part of our best practices, a viability testing plan has been developed for the ISC, which will assess the quality of archived samples. We expect that results from the viability testing will catalyze sample use, enable broader science community interest, and improve operational efficiency of the ISC. The current viability test plan focuses on generating disposition recommendations and is based on using ribonucleic acid (RNA) integrity number (RIN) scores as a criteria for measurement of biospecimen viablity for downstream functional analysis. The plan includes (1) sorting and identification of candidate samples, (2) conducting a statiscally-based power analysis to generate representaive cohorts from the population of stored biospecimens, (3) completion of RIN analysis on select samples, and (4) development of disposition recommendations based on the RIN scores. Results of this work will also support NASA open science initiatives and guides development of the NASA Scientific Collections Directive (a policy on best practices for curation of biological collections). Our RIN-based methodology for characterizing the quality of tissues stored in the ISC since the 1980s also creates unique scientific opportunities for temporal assessment across historical missions. Support from the NASA Space Biology Program and the NASA Human Research Program is gratefully acknowledged.

ALSDA↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

NASA's GeneLab: An Integrated Omics Data Commons and Workbench

GeneLab (http://genelab.nasa.gov) is a NASA initiative designed to accelerate “open science” biomedical research in support of the human exploration of space and the improvement of life on earth. The GeneLab Data Systems (GLDS) were developed to help investigators corroborate findings from “omics” (genomics, transcriptomics, proteomics, and metabolomics) assays and translate them into systems biology knowledge and, eventually, therapeutics, including countermeasures to support life in space. Phase I of the project (completed) emphasized developing key capabilities for submission, curation, storage, search, and retrieval of omics data from biomedical research in and of space environments. The development focus for Phase II (completed) was federated data search and retrieval of these kinds of data from other open-access repositories. The last phase of the project (in work) entails developing an omics analysis tool set, and a portal to visualize processed omics data, emphasizing integration with the data repository and search functions developed during the prior phases. The final product will be an open-access system where users can individually or collaboratively publish, search, integrate, analyze, and visualize omics data.

genome↗

NASA's GeneLab: An Integrated Omics Data Commons and Workbench

GeneLab (http://genelab.nasa.gov) is a NASA initiative designed to accelerate "open science" biomedical research in support of the human exploration of space and the improvement of life on earth. The GeneLab Data Systems (GLDS) were developed to help investigators corroborate findings from "omics" (genomics, transcriptomics, proteomics, and metabolomics) assays and translate them into systems biology knowledge and, eventually, therapeutics, including countermeasures to support life in space. Phase I of the project (completed) emphasized developing key capabilities for submission, curation, storage, search, and retrieval of omics data from biomedical research in and of space environments. The development focus for Phase II (completed) was federated data search and retrieval of these kinds of data from other open-access repositories. The last phase of the project (in work) entails developing an omics analysis tool set, and a portal to visualize processed omics data, emphasizing integration with the data repository and search functions developed during the prior phases. The final product will be an open-access system where users can individually or collaboratively publish, search, integrate, analyze, and visualize omics data.

genome↗

NASA's GeneLab: An Integrated Omics Data Commons and Workbench

GeneLab (http://genelab.nasa.gov) is a NASA initiative designed to accelerate "open science" biomedical research in support of the human exploration of space and the improvement of life on earth. The GeneLab Data Systems (GLDS) were developed to help investigators corroborate findings from "omics" (genomics, transcriptomics, proteomics, and metabolomics) assays and translate them into systems biology knowledge and, eventually, therapeutics, including countermeasures to support life in space. Phase I of the project (completed) emphasized developing key capabilities for submission, curation, storage, search, and retrieval of omics data from biomedical research in and of space environments. The development focus for Phase II (completed) was federated data search and retrieval of these kinds of data from other open-access repositories. The last phase of the project (in work) entails developing an omics analysis tool set, and a portal to visualize processed omics data, emphasizing integration with the data repository and search functions developed during the prior phases. The final product will be an open-access system where users can individually or collaboratively publish, search, integrate, analyze, and visualize omics data.

Berrios, Daniel C.↗

Open-Source Science-Driven Development of the Science Data System (SDS) for Earth System Observatory (ESO) Atmospheric Missions

The NASA Earth System Observatory (ESO) atmospheric missions will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The Science Data System (SDS) will deploy the adaptive processing system (APS) developed within the Cloud to manage the research and operational processing of ESO atmospheric mission orbital and suborbital sensors and curate these data for near real-time and collection reprocessing and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage and distribution. Further, the SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The SDS follows NASA’s commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the SDS system components will be developed with open-source concepts including components of APS itself as well as ESO atmospheric mission algorithms. This presentation describes the framework of the SDS and its integral part in facilitating OSS within the ESO atmospheric missions.

David M. Giles↗

PubChemLite Plus Collision Cross Section (CCS) Values for Enhanced Interpretation of Nontarget Environmental Data

Finding relevant chemicals in the vast (known) chemical space is a major challenge for environmental and exposomics studies leveraging nontarget high resolution mass spectrometry (NT-HRMS) methods. Chemical databases now contain hundreds of millions of chemicals, yet many are not relevant. This article details an extensive collaborative, open science effort to provide a dynamic collection of chemicals for environmental, metabolomics, and exposomics research, along with supporting information about their relevance to assist researchers in the interpretation of candidate hits. The PubChemLite for Exposomics collection is compiled from ten annotation categories within PubChem, enhanced with patent, literature and annotation counts, predicted partition coefficient (logP) values, as well as predicted collision cross section (CCS) values using CCSbase. Monthly versions are archived on Zenodo under a CC-BY license, supporting reproducible research, and a new interface has been developed, including historical trends of patent and literature data, for researchers to browse the collection. This article details how PubChemLite can support researchers in environmental and exposomics studies, describes efforts to increase the availability of experimental CCS values, and explores known limitations and potential for future developments. The data and code behind these efforts are openly available.

PubChem↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge and support human space missions. Through artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in space biosciences and engineered astronaut health systems, to enable Earth-independence and mission operations autonomy. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated mission biomonitoring, and 8) a Precision Space Health system. AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the space biology field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics to phenotypic data using an ensemble model to infer causality of rodent liver health disruption, 2) usage of explainable ML to interrogate muscular underpinnings of muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interactions, and 5) a suite of benchmarked open science datasets enabling programmers to identify best algorithms to answer space biology questions.

space biology↗

Enhancing accessibility and usability of Algorithm Theoretical Basis Documents through the Algorithm Publication Tool

Effective communication of scientific theories is crucial for transforming raw instrument data into valuable Earth observation products. The NASA Earth science data community disseminates this knowledge through Algorithm Theoretical Basis Documents (ATBDs). Historically, these documents lacked a standardized format, were designed for human readability rather than machine interpretation, and were challenging to locate due to the absence of a centralized repository. The Algorithm Publication Tool (APT) transforms how ATBD content is presented, simplifying the processes of creating, updating, and locating these documents. APT offers authors the option to use its user-friendly cloud-based interface or standardized templates for ATBD development. The primary advantage of the interface is its capability to manage the entire ATBD creation process within a single environment, ensuring comprehensive tracking of all activities and facilitating user tasks. Conversely, the use of standardized ATBD templates allows users to create documents using familiar tools like Google Docs, Microsoft Word, or Overleaf for LaTeX. APT also provides a centralized repository, enabling easy search and discovery of published ATBDs. This presentation showcases APT's functionalities, illustrates its contributions to advancing open science, and highlights potential benefits for broader community adoption.

Bradley Baker↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology↗

Biospecimen Culling: Temporal RNA Integrity Analysis Across Spaceflight Missions Dating from 1985 to 2011

The Ames Life Science Data Archive (ALSDA) at NASA Ames Research Center is managed by the Space Biosciences Division and has been operational since 1993. The ALSDA is responsible for archiving information and biospecimens collected from life science spaceflight experiments and matching ground control experiments. They are stored in the Ames biobank, which is located in the Biospecimen Storage Facility (BSF). The ALSDA also manages a Biospecimen Sharing Program, performs curation and long-term storage operations, and makes biospecimens available to the scientific community for research purposes via the Life Science Data Archive public website (https:lsda.jsc.nasa.gov). The BSF maintains both fixed and frozen spaceflight and ground tissues, collected from recent and past spaceflight missions. Due to the ever increasing demand for space to preserve current and future flight biospecimens, the ALSDA has initiated the development of a culling plan for biospecimens currently stored in the BSF. Culling enables the ALSDA to assess the quality of archived samples, and supports the development of standardized culling procedures that improve the operational efficiency of the BSF. The culling plan focuses on generating disposition recommendations for samples in the BSF, and currently is based on measuring ribonucleic acid (RNA) integrity number (RIN). The culling process includes (1) sorting and identification of candidate samples for RIN analysis, (2) completion of RIN analysis on select samples, and (3) development of disposition recommendations for specimens based on the RIN values. Furthermore, our approach allows for unique scientific opportunities, including development of a RIN-based methodology for culling, and temporal assessment of the quality of the tissues that have been stored in BSF since the 1980s. Results of this work will also support NASA open science initiatives.

biospecimen↗

Open-Source Science-led Development of the Atmosphere Observing System (AOS) Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David Giles↗

Open-Source Science-led Development of the AOS Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David M. Giles↗

Brain‐age prediction: Systematic evaluation of site effects, and sample age range and size

Abstract Structural neuroimaging data have been used to compute an estimate of the biological age of the brain (brain‐age) which has been associated with other biologically and behaviorally meaningful measures of brain development and aging. The ongoing research interest in brain‐age has highlighted the need for robust and publicly available brain‐age models pre‐trained on data from large samples of healthy individuals. To address this need we have previously released a developmental brain‐age model. Here we expand this work to develop, empirically validate, and disseminate a pre‐trained brain‐age model to cover most of the human lifespan. To achieve this, we selected the best‐performing model after systematically examining the impact of seven site harmonization strategies, age range, and sample size on brain‐age prediction in a discovery sample of brain morphometric measures from 35,683 healthy individuals (age range: 5–90 years; 53.59% female). The pre‐trained models were tested for cross‐dataset generalizability in an independent sample comprising 2101 healthy individuals (age range: 8–80 years; 55.35% female) and for longitudinal consistency in a further sample comprising 377 healthy individuals (age range: 9–25 years; 49.87% female). This empirical examination yielded the following findings: (1) the accuracy of age prediction from morphometry data was higher when no site harmonization was applied; (2) dividing the discovery sample into two age‐bins (5–40 and 40–90 years) provided a better balance between model accuracy and explained age variance than other alternatives; (3) model accuracy for brain‐age prediction plateaued at a sample size exceeding 1600 participants. These findings have been incorporated into CentileBrain ( https://centilebrain.org/#/brainAGE2 ), an open‐science, web‐based platform for individualized neuroimaging metrics.

60 APPLIED LIFE SCIENCES↗