Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures

The advent of single-particle cryo-electron microscopy (cryo-EM) has brought forth a new era of structural biology, enabling the routine determination of large biological molecules and their complexes at atomic resolution. The high-resolution structures of biological macromolecules and their complexes significantly expedite biomedical research and drug discovery. However, automatically and accurately building atomic models from high-resolution cryo-EM density maps is still time-consuming and challenging when template-based models are unavailable. Artificial intelligence (AI) methods such as deep learning trained on limited amount of labeled cryo-EM density maps generate inaccurate atomic models. To address this issue, we created a dataset called Cryo2StructData consisting of 7,600 preprocessed cryo-EM density maps whose voxels are labelled according to their corresponding known atomic structures for training and testing AI methods to build atomic models from cryo-EM density maps. Cryo2StructData is larger than existing, publicly available datasets for training AI methods to build atomic protein structures from cryo-EM density maps. We trained and tested deep learning models on Cryo2StructData to validate its quality showing that it is ready for being used to train and test AI methods for building atomic models.

59 BASIC BIOLOGICAL SCIENCES↗

Millimeter-wave observations of Euclid Deep Field South using the South Pole Telescope: A data release of temperature maps and catalogs

Context. The South Pole Telescope third-generation camera (SPT-3G) has observed over 10,000 square degrees of sky at 95, 150, and 220 GHz (3.3, 2.0, 1.4 mm, respectively) and will significantly overlap the ongoing 14,000 square-degree Euclid Wide Survey. The Euclid collaboration recently released Euclid Deep Field South (EDF-S) observations of 23 square degrees at wide field depths in the first quick data release (Q1). Aims. With the goal of releasing complementary millimeter-wave data and encouraging legacy science, we performed dedicated observations of a 57-square-degree field overlapping the EDF-S. Methods. The observing time totaled 20 days, and we reached noise depths of 4.3, 3.8, and 13.2 $μ$K-arcmin at 95, 150, and 220 GHz, respectively. Results. In this work we present the temperature maps and two catalogs constructed from these data. The emissive source catalog contains 601 objects (334 inside EDF-S) with 54% synchrotron-dominated sources and 46% thermal dust emission-dominated sources. The 5$σ$ detection thresholds are 1.7, 2.0, and 6.5 mJy in the three bands. The cluster catalog contains 217 cluster candidates (121 inside EDF-S) with median mass $M_{500c}=2.12 \times 10^{14} M_{\odot}/h_{70}$ and median redshift $z$ = 0.70, corresponding to an order-of-magnitude improvement in cluster density over previous tSZ-selected catalogs in this region (3.81 clusters per square degree). Conclusions. The overlap between SPT and Euclid data will enable a range of multiwavelength studies of the aforementioned source populations. This work serves as the first step toward joint projects between SPT and Euclid and provides a rich dataset containing information on galaxies, clusters, and their environments.

Archipley, M. [Chicago U., Astron. Astrophys. Ctr.↗

Predicting Critical Transitions in Multiscale Data

Predicting the dynamics of complex nonlinear systems remains a challenging problem both in dynamical systems theory as well as real world science and engineering applications. Data-driven methods utilizing the latest advances in machine learning (ML) provide a promising new paradigm for this task. Our work centered on Reservoir Computing (RC), which has shown itself to be capable of skillfully predicting chaotic dynamics in multiscale systems. In the first part of the work, the focus is on how to improve predictions of critical transitions in a class of slow-fast metastable systems in which the equations are known. An additional goal was to determine whether a relationship exists between RC and Koopman operator theory, to improve the efficiency and broaden the applicability of the approach. In the second part of this work, a variation on the RC model known as Reconstructive Reservoir Computing (RRC) is applied to real-world data to identify anomalies.

97 MATHEMATICS AND COMPUTING↗

CROCUS Tipping Bucket Rain Gauge Data from Argonne Deployable Mast Deployed at Argonne National Laboratory During Urban Flooding Campaign

The Tipping Bucket Rain Gauge (TBRG) dataset contains data from a non-heated Met One 12-inch tipping bucket rain gauge that was mounted on the Argonne Deployable Mast (ADM). The ADM is a rapid deployable meteorological trailer that can be outfitted with instrumentation to measure urban heat island effects, urban flooding or urban flux measurements. During the urban flooding field campaign, the ADM was outfitted with multiple precipitation measurement systems, including the TBRG. This dataset contains one minute measurements for precipitation accumulation during the ADM's deployment at the Argonne Testbed for Multiscale Observational Science (ATMOS) site. These data are helpful for identifying periods of precipitation, leading to potential flooding. TBRGs can be used to validate optical rain gauge data and disdrometer data collected during the CROCUS urban flooding campaign. Data were collected at ATMOS, a 20-acre prairie site at Argonne National Laboratory in Lemont, Illinois. The data is presented as daily NetCDF (.nc) files, each containing approximately 24 hours of observations. Files follow the naming convention of: the project (CROCUS), location (ADM-atmos), instrument name (tbrg), data level (raw, a1), and date (year, month, day). The NetCDF format can be accessed using common scientific software such as Python using xarray, netCDF4 or act-doe.

1-min Precipitation Accumulation↗

PSTN-019: The LSST Science Pipelines Software: Optical Survey Pipeline Reduction and Analysis Environment

The NSF-DOE Vera C. Rubin Observatory is executing the Legacy Survey of Space and Time (LSST) as its prime mission, producing a series of data releases over the ten-year survey. The LSST Science Pipelines Software will be used to create these data releases and to perform the nightly prompt processing and alert production. This paper provides an overview of the LSST Science Pipelines Software, describing the components and their integration into pipelines that generate science-ready data products.

79 ASTRONOMY AND ASTROPHYSICS↗

Data-Driven Approach for Controlled Icosahedral Boron- Rich Compound Growth

This final technical report summarizes the research accomplishments and research highlights at the end of the funding period. This project aimed to leverage existing and new computational data produced from first-principles and molecular dynamics simulations to understand the thermodynamic, mechanical, and electronic properties of icosahedral boron compounds. The goal is to achieve targeted material properties by controlling the synthesis routes of these boron-rich compounds.

36 MATERIALS SCIENCE↗

CMIP7 Data Request: atmosphere priorities and opportunities

This paper presents a comprehensive overview of the Coupled Model Intercomparison Project Phase 7 (CMIP7) request for data unlocking key research avenues in atmospheric science and provides justification for the resources needed to produce this data. Topics within the CMIP7 Atmosphere Theme centre around processes and feedbacks in atmospheric science such as clouds, aerosols and atmospheric chemistry, atmospheric circulation, temperature variability and extremes, radiative forcings, and Earth system model evaluation. These topics are summarised in this paper as scientific “opportunities” which will be realised through CMIP7 experiments and Earth system model outputs. These opportunities were submitted by a thematic group of atmospheric science community representatives combined with an extended consultation process. The production of these variables will close key gaps and uncertainties identified during previous rounds of CMIP, and will be broadly used by scientific, policy, governmental, industry, and other communities that rely on climate model projections for research and decision making, including supporting the 7th Intergovernmental Panel on Climate Change Assessment Report (AR7). As an author group, we also reflect on the process used to collate this data request and make recommendations to future CMIP governance on implementing a consultation on this scale in the future.

58 GEOSCIENCES↗

Evaluation of daily gridded climate products using in situ FLUXNET data and tree growth modeling

Gridded climate data products have facilitated research in climate and ecology by providing meteorological data continuously across large spatial scales. However, the sensitivity of scientific outcomes to dataset choice remains poorly understood, and evaluation using station-based records can favor datasets built heavily on weather stations. Here, we evaluate seven high-resolution daily gridded datasets covering the contiguous United States using independent meteorology from the FLUXNET2015 dataset, with a focus on the implications of dataset choice for process-based tree growth modeling. We find that gridded products tend to capture temperature accurately while consistently overestimating the magnitude and frequency of precipitation and its extremes. Moreover, datasets vary in how they define a ‘day,’ which significantly affects temporal alignment with FLUXNET2015 observations. Despite differences among the datasets, the interannual variability in tree ring simulations is insensitive to dataset choice, likely because daily-scale biases are averaged out through accumulated growth across several months. However, inaccuracies in temperature and precipitation can significantly bias modeled xylem cell production, with systematically higher annual precipitation in the gridded datasets leading to greater xylem production compared to simulations using in situ data. Our results suggest that model applications, especially those that integrate to time scales longer than one day, are likely insensitive to climate dataset choice, but applications that are sensitive to daily climate variations or to absolute climate values need to carefully consider biases in gridded climate products.

54 ENVIRONMENTAL SCIENCES↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

Building morphologies of the USA structures database; a gauntlet feature set

In recent years there has been a proliferation of methods and data to extract building footprints from satellite imagery. However there has been very little effort to provide additional insight about these buildings beyond their spatial location and shape. Features derived from their geometries can be used to better characterize these buildings which are critical for further research and development. In this work a set of 65 unique features for every building for more than 131 million buildings covering the US has been developed. This rich feature dataset will enable researchers, policymakers and various agencies to derive additional building characteristics like height, occupancy type, and help to gain valuable and new insights of the built environment.

Environmental sciences↗

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

Data from a multi-year targeted proteomics study of a longitudinal birth cohort of type 1 diabetes

The deployment of liquid chromatography-mass spectrometry-based plasma proteomics experiments in a large cohort is sparse, leading to a lack of data available for benchmarking, method development or validation. Comprised of 6,426 plasma analyses, The Environmental Determinants of Diabetes in the Young (TEDDY) proteomics validation study constitutes one of the largest targeted proteomics experiments in the literature to date. The proteomics data from this study were generated over the course of 2.5 years from over 900 study subjects, each providing up to 29 longitudinal samples. The data also includes 916 quality control samples. The targeted mass spectrometry assay was comprised of 694 peptides mapping to 167 proteins and the panel was measured in each subject and QC sample. The targeted proteomic dataset presented here can be used as a resource for new computational method development, such as for batch correction, as well as for benchmarking and comparing the performance of different methods/tools.

60 APPLIED LIFE SCIENCES↗

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A data-driven framework for predicting machining stability: employing simulated data, operational modal analysis, and enhanced transfer learning

Chatter, a self-excited vibration phenomenon, presents a significant challenge in machining operations, particularly in high-speed milling, where it can degrade tool life, reduce material removal efficiency, and compromise workpiece quality. Addressing this challenge requires a reliable predictive model that can accommodate the complex dynamics of various machining scenarios. This study introduces a novel, data-driven approach to predicting machining stability, leveraging over 140,000 simulated datasets and employing advanced techniques such as operational modal analysis (OMA), enhanced transfer learning (TL), and receptance coupling substructure analysis (RCSA). By integrating these methodologies, the framework effectively classifies and predicts chatter across diverse operational modes, achieving robust and accurate outcomes. Our model utilizes a Random Forest (RF) classifier trained with the comprehensive dataset, which demonstrates substantial improvements in both predictive accuracy and robustness. Specifically, the RF model achieved an accuracy rate of 85%, an area under the curve (AUC) of 0.90, and an F1 score of 0.88, underscoring its capability to adapt to varying machining configurations. These results highlight the framework’s potential to enhance operational efficiency and machining quality by providing reliable chatter predictions across a broad range of machining parameters. In conclusion, this research thus offers a significant advancement in predictive maintenance for machining processes, enabling more stable and efficient manufacturing operations.

42 ENGINEERING↗

Characterizing Wet Season Precipitation in the Central Amazon Using a Mesoscale Convective System Tracking Algorithm

To comprehensively characterize convective precipitation in the central Amazon region, we utilize the Python FLEXible object TRacKeR (PyFLEXTRKR) to track mesoscale convective systems (MCSs) observed through satellite measurements and simulated by the Weather Research and Forecasting model at a convection-permitting resolution. This study spans a 2-month period during the wet seasons of 2014 and 2015. We observe a strong correlation between the MCS track density and accumulated precipitation in the Amazon basin. Key factors contributing to precipitation, such as MCS properties (number, size, rainfall intensity, and movement), are thoroughly examined. Our analysis reveals that while the overall model produces fewer MCSs with smaller mean sizes compared to observations, it tends to overpredict total precipitation due to excessive rainfall intensity for heavy rainfall events (≥10 mm hr –1 ). These biases in simulated MCS properties could vary with the constraints on the convective background environment. Moreover, while the wet bias from heavy (convective) rainfall outweighs the dry bias in light (stratiform) rainfall, the latter can be crucial, particularly when MCS cloud cover is significantly underestimated. A case study for 1 April 2014 highlights the influence of environmental conditions on the MCS lifecycle and identifies an unrealistic model representation in both stratiform and convective precipitation features.

54 ENVIRONMENTAL SCIENCES↗

Dataset of tensile properties for sub-sized specimens of nuclear structural materials

Mechanical testing with sub-sized specimens plays an important role in the nuclear industry, facilitating tests in confined experimental spaces with lower irradiation levels and accelerating the qualification of new materials. The reduced size of specimens results in different material behavior at the microscale, mesoscale, and macroscale, in comparison to standard-sized specimens, which is referred to as the “specimen size effect.” Although analytical models have been proposed to correlate the properties of sub-sized specimens to standard-sized specimens, these models lack broad applicability across different materials and testing conditions. The objective of this study is to create the first large public dataset of tensile properties for sub-sized specimens used in nuclear structural materials. We performed an extensive literature review of relevant publications and extracted over 1,000 tensile testing records comprising 55 columns including material type and composition, manufacturing information, irradiation conditions, specimen dimensions, and tensile properties. The dataset can serve as a valuable resource to investigate the specimen size effect and develop computational methods to correlate the tensile properties of sub-sized specimens.

36 MATERIALS SCIENCE↗