Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

NASAs GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate open science biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics (omics) data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

genome↗

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES↗

G2Aero Database of Airfoils - Curated Airfoils

This dataset contains a curated set of 19,164 airfoil shapes from various applications and the data-driven design space of separable shape tensors (PGA space), which can be used as a parameter space for machine-learning applications focused on airfoil shapes. We constructed the airfoil dataset in two main stages. First, we identified 13 baseline airfoils from the NREL 5MW and IEA 15MW reference wind turbines. We reparameterized these shapes using least-squares fits of 8-order CST parametrizations, which involve 18 coefficients. By uniformly perturbing all 18 CST coefficients by +/-20% around each baseline airfoil, we generated 1,000 unique airfoils. Each airfoil was sampled with 1,001 shape landmarks whose x-coordinates followed a cosine distribution along the chord. This process resulted in a total of 13,000 airfoil shapes, each with 1,001 landmarks. In the second phase, we gathered additional airfoils from the extensive BigFoil database, which consolidates data from sources such as the University of Illinois Urbana-Champaign (UIUC) airfoil database, the JavaFoil database, the NACA-TR-824 database, and others. We undertook a thorough pre-processing step to filter out shapes with sparse, noisy, or incomplete data. We also removed airfoils with sharp leading edge and those exceeding our threshold for trailing edge thickness. Additionally, we thinned out the collection of NACA airfoils-- parametric sweeps of NACA airfoils with increasing thickness and camber present in BigFoil database-- by selecting every fourth step in the parameter sweeps. Finally, we regularized the airfoils by reparametrizing them with an 8-order CST parametrization (with 1,001 shape landmarks with x coordinated following cosine distribution along the chord) and removing airfoils with high reconstruction errors. This data pre-processing resulted in a set of 6,164 airfoils. In total, our curated airfoil dataset comprises 19,164 airfoils, each with 1,001 landmarks, and is stored in the curated_airfoils.npz file. Using this curated airfoil dataset, we utilized the separable shape tensors framework to develop a data-driven parameterization of airfoils based on principal geodesic analysis (PGA) of separable shape tensors. This PGA space is provided in PGAspace.npz file.

airfoils↗

Enabling Space Biological Knowledge Discovery Through Image and Video Data Sharing

Increased biomedical risks associated with deep space crewed missions (cis-Lunar, Mars transit/surface) require development of health countermeasures, novel ecosystem support, risk modeling, and fundamental space biological knowledge discovery. Molecular-omics, physiological-phenotypic-behavioral, and environmental-radiation telemetry data from space biological and health studies are needed for reuse by scientists to address these tasks. The data as well as space-relevant biospecimens are being made more findable, accessible, interoperable, and reusable through NASA’s Open Science Data Repository (OSDR). This new OSDR umbrella grouping includes NASA GeneLab, the NASA Ames Life Sciences Data Archive (ALSDA), and the NASA Biological Institutional Scientific Collection. The OSDR system design appropriately handles metadata and processed-tabular results from ALSDA studies collected from space experiments. But raw and processed ALSDA bioimage and video datasets require an expansion of OSDR’s data architecture to handle ingestion, curation, and egress. The academic-industry bioimaging field saw a scientific renaissance in the past several years through leveraging open-source software, international collaborations, machine learning, and other open science/programming approaches. As crewed missions and more biological experiments are on the deep space horizon, OSDR is embracing data stewardship through listening to feedback from subject matter experts and designing an expanded architecture which is appropriate for NASA’s goals to enable analysis and reuse of bioimaging and video data for the public science community.Discovery Through Image and Video Data Sharing

space biology↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Cross Kingdom Analysis of Data Within the GeneLab Repository Identifies a Potential Conserved Response of Life to the Stress Associated with Spaceflight

It is important to determine the health risks and potential survival for astronauts associated with long-term space missions. This entails not only understanding the impact the space environment will have on humans, but also how it will affect other organisms needed for humans to survive in space such as plants. In addition, it has been reported in the literature that hundreds of genes seem to be conserved and/or transferred between different organisms from bacteria, archaea, fungi, microorganisms, and plants to animals. Since space travel involves humans in a closed environment over a long period of time, we hypothesize that potential conserved biological factors will occur between the different organisms in that environment possibly due to transfer of genes. Determining the conserved factors that are commonly being regulated in space can shed insight into possible universal master regulators and also determine the symbiotic relationship between the organisms in space. Utilizing NASA's GeneLab Data Repository (a rapidly expanding, curated clustering of spaceflight-related ‘omics-level datasets for all organisms), we were able to uncover a novel pathway and factors that were commonly shared between humans, mice, plants, C. Elegans, and drosophilas. Through ChIP-Seq enrichment analysis techniques utilizing various GeneLab datasets from each species that were flown in space, we found the following factors to be conserved across all species: oxidative stress, DNA damage (through GABPA/NRFs and NFY), SIX5, GTF2B and glutamine synthetase. Such commonalities would likely reflect the effects of factors such as microgravity and the increased radiation exposure inherent in spaceflight on basic physical processes shared by all biological systems at the cellular level. Differences between organismal responses revealed by GeneLab's data should also help understand the unique reactions to life in space that arise from the very different lifestyles of microbes, animals and plants.

Barker, Richard↗

Event Log / Raw Data

The WFIP3 event log is a curated record spanning 578 days of meteorological phenomena and field observations that complements the campaign’s high-frequency measurements. The log combines manually documented daily weather discussions with automatically derived indicators of key atmospheric processes, providing standardized, publicly available context to support model evaluation, forecast verification, and case-study selection for offshore boundary-layer research.

17 WIND ENERGY↗

Exploring our Changing Planet through NASA’s Earth Information Center and Novel Methods for Earth Science Communication

In June 2023, NASA unveiled the Earth Information Center (EIC), an interagency initiative which invites the global community to explore how our planet is changing through interactive installations, captivating visualizations, and novel storytelling. Existing in both physical and virtual space, the EIC serves as a gateway to actionable information collected through an expanding fleet of Earth observing satellites and sensors. The EIC features data driven visualizations, near real-time information, immersive experiences, and curated stories that highlight the applications of publicly available data to address environmental challenges across nine thematic areas: agriculture, biodiversity, disasters, greenhouse gases, air quality, sea level rise, sustainable energy, water resources, and wildfires. Designed by an interdisciplinary team, EIC exhibits are developed to reach a wide base of end-users across multiple learning modalities. During the first year of operation, the inaugural location of the EIC at NASA Headquarters in Washington, DC, welcomed an estimated 24,000 visitors, hosted over 145 tours for domestic and international organizations, supported NASA’s Earth Day 2024 programming, and led six STEM programs for diverse student communities. With lessons learned from the first year of being open to the public and a variety of novel science communication tools in development, the EIC is expanding the mission’s reach by collaborating with museums and visitor centers. Through these collaborations, we aim to inform a broader demographic about the unprecedented changes observed in Earth’s climate and inspire communities to learn more about their one and only home, planet Earth.

Nicole Ramberg-Pihl↗

GeneLab: Current and Future Omics Data Integration Between Space Biology and HRP

For the past five years, the Biological and Physical Sciences Division has pioneered Open Science in Space Biology by funding the NASA GeneLab project. Along with the Ames Life Sciences Data Archive, GeneLab has quickly become the world leader in archiving and scientifically curating spaceflight and spaceflight relevant multi-omics data. Specifically, the GeneLab Data System has become a full enterprise solution providing advanced mining capabilities, several application programming interfaces for data federation and machine learning approaches, and delivering to the world an analytical and visualization platform which has enabled collaboration within the scientific community. Over the past three years, large meta-analysis and modeling studies have been published by the GeneLab Analysis Working Groups (AWGs), which are comprised of ~200 volunteer scientists. One natural extension of GeneLab data reuse has recently turned towards linking animal data with human data, which is the next necessary step to further validate animal models for inferring biological risks to humans conducting LEO, lunar or Martian missions. As such, data from the Human Research Program are an essential component of GeneLab and ALSDA. At the moment, simulated space radiation experiments conducted at Brookhaven National Laboratory make the most of HRP GeneLab data, and the scientific community has been eager to also link their animal spaceflight results to actual Astronaut data and human analog data. We will discuss further the current status of knowledge and future approaches to accelerate our basics understanding of the impact of space stressors on humans using latest omics technology.

omics↗

Pu(IV) quantification via visible–near-infrared absorption spectroscopy: tackling interferences using D-optimal design and partial least squares

Here, this study presents a novel analytical approach for quantifying Pu(IV) in glove box environments using fiber-optic-based visible–near-infrared absorption spectroscopy in combination with partial least squares regression (PLSR) and design of experiments. The method addresses significant challenges posed by overlapping spectral features arising from Nd(III), which is a common fission product impurity, and the speciation variability of Pu(IV) nitrato complexes in HNO 3 concentrations ranging from 2.5 to 11 M. A curated training set consisting of data from 20 samples was developed via D-optimal design to enable robust PLSR model calibration for Pu(IV) using the near-infrared band near 1050 nm. The training set was acquired from samples in cuvettes with a 1-cm path length and was used to build the PLSR model. The robustness of the model was validated with data collected using a dip probe with a 1-cm path length and varying Pu(IV) concentrations. The strong performance of the model indicates good model transfer from cuvette to dip probe and highlights the potential for in situ measurements and online monitoring of reactions in a crystallization reactor vessel. The results demonstrate that this combined spectroscopic and chemometric approach can accurately and simultaneously quantify Pu(IV) and HNO 3 , thereby offering a promising tool for real-time monitoring in process environments.

Actinide↗

Gold-Standard Chemical Database 137 (GSCDB137): A Diverse Set of Accurate Energy Differences for Assessing and Developing Density Functionals

We present GSCDB137, a rigorously curated benchmark library of 137 data sets (8377 entries) covering main-group and transition-metal reaction energies and barrier heights, (intra- and intermolecular) noncovalent interactions, dipole moments, polarizabilities, electric-field response energies, and vibrational frequencies. Legacy data from GMTKN55 and MGCDB84 have been updated to today's best reference values; redundant or low-quality points were removed, and many new, property-focused sets were added. Testing 29 popular density functional approximations (DFAs) confirms the expected Jacob's-ladder hierarchy overall but also reveals notable exceptions: functional performance for frequencies and electric-field properties correlates poorly with that for other ground-state energetics. ωB97M-V and ωB97X-V are the most balanced hybrid meta-GGA and hybrid GGA, respectively; B97M-V and revPBE-D4 lead the meta-GGA and GGA classes. Double hybrids lower mean errors by about 30% versus their hybrid analogues but demand careful frozen-core, basis set, and spin contamination treatment. GSCDB137 offers a comprehensive, openly documented platform for rigorous validation of DFA and universal machine learning potentials, and training of the next generation of exchange-correlation functionals.

Liang, Jiashu [University of California, Berkeley,↗

Comparative Performance Evaluation of Large Language Models for Extracting Molecular Interactions and Pathway Knowledge

Understanding the interactions and regulatory relationships among biomolecules is essential for deciphering complex biological systems and elucidating the mechanisms behind diverse biological functions. Traditionally, the collection of such molecular interaction data has relied on expert curation, a process that is both time-consuming and labor-intensive. To address these limitations, this study explores the use of large language models (LLMs) to automate the genome-scale extraction of molecular interaction knowledge. Here, we evaluate the performance of various LLMs on key biological tasks, including the identification of protein-protein interactions, detection of genes associated with pathways influenced by low-dose radiation, and inference of gene regulatory relationships. Our findings demonstrate that larger LLMs tend to perform better, particularly in extracting intricate gene and protein interactions. Despite their strengths, these models face challenges in recognizing functionally diverse gene groups and highly correlated regulatory relationships. Through a comprehensive analysis using established molecular interaction and pathway databases, we show that LLMs possess the potential to identify relevant biomolecules and predict their interactions, offering valuable insights and marking a significant step toward AI-driven biological knowledge discovery.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

A Review of Resilience and Long-Term Planning in Power and Water Systems in the United States

There is recognition among power and water utilities that the frequency and magnitude of high consequence and low probability events could increase as a result of climate change. The interconnected nature of energy-water systems raises the possibility of cascading failures, increasing complexity and risks. Resilience and long-term planning are important ways of weathering the effects of climate change. First, to understand more about resilience, we reviewed existing literature on resilience definitions, metrics, and modeling, focusing on integrated water-power systems. Second, to understand how resilience and planning are being applied in practice, we interviewed utilities and organized, curated, and synthesized the interview data to arrive at several key findings, which are presented here. We found that there is not a consistent definition for resilience, yet it is something that utilities regularly plan for, often with different names and varying methods/measures. However, there is a tangible shift in the industry towards defining and determining measurable resilience metrics. While the exact metrics are a work in progress, utilities are taking steps forward by (1) putting people and culture at the center of resilience, (2) recognizing their own interdependencies, and (3) pursuing better cross-sector collaboration.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Dark Energy Survey Year 6 Results: Photometric Dataset for Cosmology

We describe the photometric dataset assembled from the full 6 yr of observations by the Dark Energy Survey (DES) in support of static-sky cosmology analyses. DES Y6 Gold is a curated dataset derived from DES Data Release 2 (DR2) that incorporates improved measurement, photometric calibration, object classification and value-added information. Y6 Gold comprises nearly 5000 deg$^{2}$ of grizY imaging in the south Galactic cap and includes 669 million objects with a depth of i$_{AB}$ ∼ 23.4 mag at a signal-to-noise ratio ∼ 10 for extended objects and a top-of-the-atmosphere photometric uniformity <2 mmag. Y6 Gold augments DES DR2 with simultaneous fits to multiepoch photometry for more robust galaxy shapes, colors, and photometric redshift estimates. Y6 Gold features improved morphological star–galaxy classification with an efficiency of 98.6% and a contamination of 0.8% for galaxies with 17.5 < i$_{AB}$ < 22.5. Additionally, it includes per-object quality information, and accompanying maps of the footprint coverage, masked regions, imaging depth, survey conditions, and astrophysical foregrounds that are used for cosmology analyses. After quality selections, benchmark samples contain 448 million galaxies and 120 million stars. This publication is complemented by data access and documentation.

79 ASTRONOMY AND ASTROPHYSICS↗

Reuse of Software Assets for the NASA Earth Science Decadal Survey Missions

Software assets from existing Earth science missions can be reused for the new decadal survey missions that are being planned by NASA in response to the 2007 Earth Science National Research Council (NRC) Study. The new missions will require the development of software to curate, process, and disseminate the data to science users of interest and to the broader NASA mission community. In this paper, we discuss new tools and a blossoming community that are being developed by the Earth Science Data System (ESDS) Software Reuse Working Group (SRWG) to improve capabilities for reusing NASA software assets.

Mattmann, Chris A.↗

X-Ray Micro-Computed Tomography of Apollo Samples as a Curation Technique Enabling Better Research

X-ray micro-computed tomography (micro-CT) is a technique that has been used to research meteorites for some time and many others], and recently it is becoming a more common tool for the curation of meteorites and Apollo samples. Micro-CT is ideally suited to the characterization of astromaterials in the curation process as it can provide textural and compositional information at a small spatial resolution rapidly, nondestructively, and without compromising the cleanliness of the samples (e.g., samples can be scanned sealed in Teflon bags). This data can then inform scientists and curators when making and processing future sample requests for meteorites and Apollo samples. Here we present some preliminary results on micro-CT scans of four Apollo regolith breccias. Methods: Portions of four Apollo samples were used in this study: 14321, 15205, 15405, and 60639. All samples were 8-10 cm in their longest dimension and approximately equant. These samples were micro-CT scanned on the Nikon HMXST 225 System at the Natural History Museum in London. Scans were made at 205-220 kV, 135-160 microamps beam current, with an effective voxel size of 21-44 microns. Results: Initial examination of the data identify a variety of mineral clasts (including sub-voxel FeNi metal grains) and lithic clasts within the regolith breccias. Textural information within some of the lithic clasts was also discernable. Of particular interest was a large basalt clast (approx.1.3 cc) found within sample 60639, which appears to have a sub-ophitic texture. Additionally, internal void space, e.g., fractures and voids, is readily identifiable. Discussion: It is clear from the preliminary data that micro-CT analyses are able to identify important "new" clasts within the Apollo breccias, and better characterize previously described clasts or igneous samples. For example, the 60639 basalt clast was previously believed to be quite small based on its approx.0.5 sq cm exposure on the surface of the main mass. These scans show the clast to be approx.4.5 g, however (assuming a density of approx.3.5 g/cc). This is large enough for detailed studies including multiple geo-chronometers. This basalt clast is of particular interest as it is the largest Apollo 16 basalt, and it is the only mid-TiO2 basalt in the Apollo sample suite. By identifying the location of interesting clasts or grains within a sample, we will be able to make more informed decisions about where to cut a sample in order to best expose clasts of interest for future study. Moreover, knowing the location of internal defects (e.g., fractures) will allow more precise chipping and extraction of clasts or grains. By combining micro-CT scans with compositional techniques like micro x-ray fluorescence (particularly on sawn slabs), we will be able to provide even more comprehensive information to scientists trying to best select samples that fit their scientific needs.

Ziegler, R. A.↗