Search NASA⌕ Search

SEARCH · Search NASA

Results for “Findable”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Filling in Subsurface Storage Open Data Gaps - Updates to CCS Data Availability on EDX and EDX Spatial (FWP-1022465)

There is a need to preserve and efficiently access data resources to drive the next generation of research and development while ensuring compliance with DOE regulations. Over the last 10+ years, there has been ongoing efforts by the DOE Carbon Storage Program to ensure that there is effective data curation and preservation of DOE funded research leveraging the NETL-FECM data repository, the Energy Data eXchange (EDX). This talk presents updates about ongoing efforts to continue to support the mission of ensuring that carbon storage data is findable, accessible, interoperable, and reusable to the carbon storage stakeholder community through EDX and EDX Spatial. Presented at the NETL Carbon Management Review Meeting, Pittsburgh, 2024.

Morkner, Paige↗

Workflows Community Summit 2024: Future Trends and Challenges in Scientific Workflows

The 2024 Workflows Community Summit report presents the outcomes of a three-day international gathering that brought together 109 experts from 18 countries to discuss future trends and challenges in scientific workflows. The summit focused on six key areas: time-sensitive workflows, convergence of AI and HPC workflows, multi-facility workflows, heterogeneous HPC environments, user experience and interfaces, and FAIR computational workflows. Discussions highlighted emerging challenges such as integrating AI with traditional HPC, managing workflows across diverse facilities, addressing heterogeneity in computing environments, and ensuring workflows are findable, accessible, interoperable, and reusable (FAIR). The report outlines recent advances, ongoing challenges, and provides recommendations for each topic area, emphasizing the need for standardization, improved interoperability, and the development of more sophisticated tools and frameworks to support the evolving landscape of scientific workflows in the era of exascale computing and AI integration.

97 MATHEMATICS AND COMPUTING↗

AIR Framework for Physics-Inspired Artificial Intelligence in High Energy Physics (Final Report)

The FAIR4HEP project was a collaboration between Argonne National Laboratory, University of Illinois Urbana-Champaign, Massachusetts Institute of Technology, University of Minnesota, and the University of California San Diego funded by the US Department of Energy, Office of Science, Office of Advanced Scientific Research (ASCR) from 2020 to 2023. The primary focus of the FAIR4HEP project was to advance our understanding of the relationship between data and artificial intelligence (AI) models by exploring relationships among them through the development of findable, accessible, interoperable, and reusable (FAIR) frameworks. Using high-energy physics (HEP) as the science driver, this project developed a FAIR framework to advance our understanding of AI, provide new insights to apply AI techniques and provide an environment where novel approaches to AI can be explored. This final report summarizes the accomplishments of the MIT group.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HPC-FAIR: A Framework Managing Data and AI Models for Analyzing and Optimizing Scientific Applications

The increasing reliance on machine learning (ML) to analyze and optimize large-scale scientific applications on supercomputers faces a significant bottleneck: the lack of readily available, high-quality training datasets and the difficulty in reusing existing AI models. This project was motivated by the urgent need to address the “FAIR” principles (Findability, Accessibility, Interoperability, Reusability) for both training datasets and AI models in the high-performance computing (HPC) domain. The project developed HPC-FAIR, a high-performance computing data management framework designed to centralize HPC-related datasets and AI models within a unified hub. To ensure interoperability, the framework established a standardized representation and vocabulary (ontology) for both data and models. HPC-FAIR also implemented automated workflows to streamline data processing, model access, and benchmarking. Additionally, the project focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics.

97 MATHEMATICS AND COMPUTING↗

FAIR Framework for Physics-Inspired Artificial Intelligence in High Energy Physics (Final Report)

The FAIR4HEP project was a collaboration between Argonne National Laboratory, University of Illinois Urbana-Champaign, Massachusetts Institute of Technology, University of Minnesota, and University of California San Diego funded by the US Department of Energy, Office of Science, Office of Advanced Scientific Research (ASCR) from 2020 to 2024. The primary focus of the FAIR4HEP project was to advance our understanding of the relationship between data and artificial intelligence (AI) models by exploring relationships among them through the development of findable, accessible, interoperable, and reusable (FAIR) frameworks. Using high-energy physics (HEP) as the science driver, this project developed a FAIR framework to advance our understanding of AI, provide new insights to apply AI techniques, and provide an environment where novel approaches to AI can be explored. This final report summarizes the accomplishments of the University of Illinois group.

43 PARTICLE ACCELERATORS↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

The Heliophysics Data Environment Today

Driven by the nature of the research questions now most critical to further progress in heliophysics science, data-driven research has evolved from a model once centered on individual instrument Principal investigator groups and a circle of immediate collaborators into a more inclusive and open environment where data gathered ay great public cost must then be findable and useable throughout the broad national and international research community. In this paper and as an introduction to this special session, we will draw a picture of existing and evolving resources throughout the heliophyscs community, the capabilities and data now available to end users, and the relationships and complementarity of different elements in the environment today. We will cite the relative roles of mission and instrument data centers and resident archives, multi-mission data centers, and the growing importance of virtual discipline observatories and cross-cutting services including the evolution of a common data dictionary. We will briefly summarize our view of the most important challenges still faced by users and providers, and our vision in ow the efforts today can evolve into a more and more enabling data framework for the global research community to tap the widest range of existing missions and their data to address a full range of critical science questions from the scale of microphysics to the heliospheric system as a whole.

Fung, Shing F.↗

FAIRness and Usability for Open-Access Omics Data Systems

Omics data sharing is especially crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the "FAIRness" of NASA's GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. 14 metrics. The range of Pass ratings was 29-79% of the 14 metrics, Partial Pass 0-21%, and Fail 7-50%. The range of overall FAIRness scores was 5-12 (out of 14). The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. We propose two new principles that Big Data systems, in particular, should consider for increasing data accessibility. We relate our experiences implementing semantic integration of omics data from several systems for the federated querying and retrieval functions of the GLDS, given the shortcomings in data interoperability of these systems.

Berrios, Daniel C.↗

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the "FAIRness" of NASA's GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.↗

Linking Asteroid Detections from the Large Synoptic Survey Telescope

We have conducted a detailed simulation of the Large Synoptic Survey Telescope (LSST) in order to understand the system’s ability to link detections of asteroids within and across nights in order to populate a catalog of asteroid orbits. We show that LSST, using its baseline survey cadence, should be able to successfully link and catalog asteroids. In our simulation of a single monthly observing cycle, LSST produced 66 million candidate detections of main belt asteroids (MBAs) and near Earth objects (NEOs), of which 77% were spurious detections related to detector noise or image processing. Using the Moving Object Processing System, we were able to assemble single-night “tracklets” with negligible losses, but a purity of only 43%. The next stage of linking led to three-night orbits with data sets no more than 12 days in length, and it is at this stage that the false detections are readily removed from the data stream. Main-belt linkages were essentially complete and 99.8% pure. Similarly, only 0.02% of linked detections involving NEOs were spurious. On the other hand, NEO linking was 93.6% complete, indicating that 6.4% of potentially findable NEOs were not successfully linked. We believe that this rate can be improved with careful tuning of the MOPS linking algorithms. The NEO catalog was affected by main-belt confusion so that mis-linked MBAs appeared as NEOs, and many correctly linked MBAs were consistent with NEO orbits. We show that these cases arise primarily from MBAs detected at lower solar elongations and we postulate that this is an artifact of a one month simulation that will be readily resolved by surveying over many months.

Chesley, Steven R.↗

NASA ESDIS Standards Office

This poster describes the function of the NASA ESDIS Standards Office, lists the findable, interoperable, accessible and the reusable standards that have been reviewed and endorsed for broader use in NASA data and information systems, and the impacts of these endorsed standards.

Lynnes, Chris↗

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the “FAIRness” of NASA’s GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.↗

Perspectives on Data Reproducibility and Replicability in Paleoclimate and Climate Science

This paper summarizes the current state of reproducibility and replicability in the fields of climate and paleoclimate science, including brief histories of their development and applications in climate science, new and recent approaches towards improvement of reproducibility and replicability, and challenges. Recommendations for addressing those challenges include: development of searchable, auto-updated, interlinked, multi-archive public paleoclimate repositories for raw and processed digital datasets; cross-center standardized code base cases, improved data storage techniques, and a focus on replicability for climate simulation storage and access; and support of the development and community awareness of findable, accessible, interoperable and reusable (FAIR) principles by funding agencies and publishers. This paper is largely based on the May 2018 presentations of a panel of researchers to the Committee on Reproducibility and Replicability in Science, part of the National Academies of Science, Engineering, and Medicine. The commentary and recommendations made here are in alignment with those of its Consensus Study Report on Reproducibility and Replicability in Science (2019).

data repositories↗

Spaceflight Biospecimen Sharing in Support of Science Discovery and Exploration

For decades, NASA and international partners have flown non-human biological experiments in space to understand the effects of spaceflight and address potential biological hazards. Sending organisms into space is a costly endeavor which makes space-flown biological specimens a valuable resource. To enable maximum scientific return, samples not required by the Principal Investigators are harvested and collected mostly by NASA’s Space Biology Biospecimen Sharing Program. These specimens are collected according to well-established SOPs that maintain quality and integrity. The specimens are then preserved, archived, and made available to the international scientific community through NASA’s Institutional Scientific Collection (ISC) at Ames Research Center (ARC). The ISC-ARC biospecimens and descriptive metadata are findable and accessible for request through the Life Sciences Data Archive (LSDA). The NASA ISC-ARC currently stores over 32,000 specimens from Shuttle, International Space Station, and ground-based investigations (spaceflight analog experiments involving either hindlimb unloading, centrifugation, or partial weight-bearing study designs). Tissues are predominantly from mice and rats, though samples are also available from bacteria and quail. The specimens include tissues from many physiological systems including musculoskeletal, neurosensory, reproductive, respiratory, circulatory, and digestive. Tissues are stored at -80°C, -20°C, +4°C, or ambient and preserved in various fixatives. Descriptive metadata is available for all samples. Historically, these tissues have been used for a wide range of analyses, including histology, genomics, and transcriptomics. Plans are underway to expand the ISC-ARC beyond the mostly-rodent contents, to include a space-relevant microbial culture collection including bacteria, fungi, and yeast. This expansion of the ISC-ARC will now involve identifying and standardizing best practices for microbial curations. To ensure safe long-term storage of microbial isolates, a microbiology laboratory will be dedicated for identification, cell culture, and lyophilization. Awarding of tissue to public science investigators has resulted in 33 publications since 2011, with 48 requests being submitted since 2016. Of note, NASA GeneLab has been awarded ISC-ARC biospecimens in the past few years. GeneLab processes the biospecimens to generate various levels of ‘omics’ data, which are published on GeneLab’s open access online platform for bioinformatics analysis and visualization. This has helped a systems biology community grow around the processed-biospecimens’ datasets, resulting in many new publications and insights. Websites: https://www.nasa.gov/ames/research/space-biosciences/isc-bsp ; https://lsda.jsc.nasa.gov/Biospecimen

Ryan T. Scott↗

Expanding Biological Repository Data Available for Sharing and Knowledge Discovery

Biology has developed next-generation data science and alternative analytical approaches with methodologies which require principal investigator (PI) experimental assay data be re-used. This new approach involves mining multiple datasets at once from various hierarchical organizations of biological complexity, while concurrently evaluating how experimental factors affect endpoints of standard assays. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make findable, accessible, interoperable, and reusable (FAIR) all non-human space-relevant biological data. These data include mission metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery, and subject-experienced telemetry (radiation, temperature, humidity, acoustics, vibrations). ALSDA has transformed to bring current biological repository data and all future collected data into this new scientific data mining reality. It has integrated into the ‘NASA Open Science’ group of projects to facilitate a suite of new tools and workflows to improve data accessibility and reusability by implementing data management plans, automating data submission agreements, and adopting the single-point-of-entry data submission portal, originally developed by NASA GeneLab. These systems required ALSDA to develop science assay configurations for the submission portal, capturing essential assay parameters according to established norms in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. ALSDA datasets are curated to maintain rich metadata, accuracy of datasets, data transparency, provenance, and additionally ensure data are machine-readable (e.g., R and Python languages). ALSDA integration with GeneLab and its analysis portals enable higher-order physiological-level datasets be mined in conjunction with -omics datasets. As ALSDA physiological-level datasets are published (micro-computed tomography, histology, intraocular pressure, hormonal assays, immunostaining, ultrasonography), the merging of hierarchical organizations of biological complexity from spaceflight will enable new knowledge discovery approaches.

Ryan T Scott↗