Search NASASearch

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

DDCP framework

DDCP protocol software 1.0 This repository contains the C++ implementation of version 1.x of the Distributed Data Communications Protocol (DDCP). DDCP provides request/reply, feature discovery, data transfer, control, interrupt, and transaction support for communicating with accelerator instrumentation over UDP. The standard server port is 65000. The framework is a source dependency for services that communicate directly with DDCP hardware. It is not a deployable service by itself.

Joshi, Shreya [Fermi National Accelerator Laborato

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige

Long-term measurements of ice nucleating particles at Atmospheric Radiation Measurement (ARM) sites worldwide

Ice nucleating particles (INPs) play a critical role in cloud microphysics and precipitation formation, yet long-term, spatially extensive observational datasets remain limited. Here, we present one of the most comprehensive publicly available datasets of immersion-mode INP concentrations using a single analytical method, generated through the U.S. Department of Energy's (DOE) Atmospheric Radiation Measurement (ARM) user facility. INP filter samples have been collected across a broad range of environments – including agricultural plains, Arctic coastlines, high-elevation mountain sites, marine regions, and urban areas – via fixed observatories, mobile facility deployments, and vertically-resolved tethered balloon system operations. We describe the standardized processing and quality assurance pipeline, from filter collection and processing using the Ice Nucleation Spectrometer to final data products archived on the ARM Data Discovery portal. The dataset includes both total INP concentrations and selectively treated samples, allowing for classification of biological, organic, and inorganic INP types. It features a continuous 5-year record of INP measurements from a central U.S. site, with data collection still ongoing. Seasonal and site-specific differences in INP concentrations are illustrated through intercomparisons at −10 and −20 °C, revealing distinct regional sources and atmospheric drivers. We also outline mechanisms for researchers to access existing data, request additional sample analyses, and propose future field campaigns involving ARM INP measurements. This dataset supports a wide range of scientific applications, from observational and mechanistic studies to model development, and provides critical constraints on aerosol-cloud interactions across diverse atmospheric regimes (Creamean et al., 2024, 2020b; https://doi.org/10.5439/1770816).

Creamean, Jessie M. [Colorado State Univ., Fort Co

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew

CoRE MOF DB: A curated experimental metal-organic framework database with machine-learned properties for integrated material-process screening

Here, we present an updated version of the Computation-Ready, Experimental (CoRE) Metal-Organic Framework (MOF) database, which includes a curated set of computation-ready MOF crystal structures designed for high-throughput computational materials discovery. Data collection and curation procedures were improved from the previous version to enable more frequent updates in the future. Machine-learning-predicted properties, such as stability metrics and heat capacities, are included in the dataset to streamline screening activities. An updated version of MOFid was developed to provide detailed information on metal nodes, organic linkers, and topologies of an MOF structure. DDEC6 partial atomic charges of MOFs were assigned based on a machine-learning model. Gibbs ensemble Monte Carlo simulations were used to classify the hydrophobicity of MOFs. The finalized dataset was subsequently used to perform integrated material-process screening for various carbon-capture conditions using high-fidelity temperature-swing adsorption (TSA) simulations. Our workflow identified multiple MOF candidates that are predicted to outperform CALF-20 for these applications.

CoRE MOF database

2024 NMDC Ambassador Training Materials [Slides]

The NMDC is a sustainable data discovery platform that promotes open science and shared-ownership across a broad and diverse community of researchers, funders, publishers, societies, and other collaborators. The NMDC aims to enable multi-omic microbiome research to accelerate scientific discovery. The NMDC is a Department of Energy funded program that is a collaboration between 3 National Laboratories: Lawrence Berkeley National Laboratory (LBNL), Los Alamos National Laboratory (LANL), and Pacific Northwest National Laboratory (PNNL).

54 ENVIRONMENTAL SCIENCES

HTESP (High-throughput electronic structure package): A package for high-throughput ab initio calculations

High-throughput ab initio calculations are the indispensable parts of data-driven discovery of new materials with desirable properties, as reflected in the establishment of several online material databases. The accumulation of extensive theoretical data through computations enables data-driven discovery by constructing machine learning and artificial intelligence models to predict novel compounds and forecast their properties. Efficient usage and extraction of data from these existing online material databases can accelerate the next stage materials discovery that targets different and more advanced properties, such as electron–phonon coupling for phonon-mediated superconductivity. However, extracting data from these databases, generating tailored input files for different ab initio calculations, performing such calculations, and analyzing new results can be demanding tasks. Here, in this work, we introduce a software package named “HTESP” (High-Throughput Electronic Structure Package) written in Python and Bash languages, which automates the entire workflow including data extraction, input file generation, calculation submission, result collection and plotting. Our HTESP will help speed up future computational materials discovery processes.

36 MATERIALS SCIENCE

Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots

Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given input/output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of the variables of the hypergraph to discover its structure and approximate its unknown functions. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most scientific problems can be framed as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). We introduce an interpretable Gaussian Process (GP) framework for such (Type 3) problems that does not require randomization of the data, access to or control over its sampling, or sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super-exponential complexity of causal inference methods, is enabled by the nonlinear ANOVA capabilities of GPs used as a sensing mechanism.

Science & Technology - Other Topics

Data Science-Driven Discovery of Multimetallic Oxygen-cycle Electrocatalysts for Enhanced Energy Conversion

The overarching objective of this effort has been to combine state-of-the-art data science techniques, first principles analyses, and molecular-level characterization of electrocatalyst structure and reactivity to identify both in-situ mechanisms for degradation and transformation of electrocatalysts with highly complex catalytic structures and the impact of these transformations on catalytic activity. The primary catalysts of interest have been multielemental alloys, including high entropy alloys (HEA’s), which are characterized by a high degree of disorder and up to 20 different elements within a single nanoparticle. We have applied these strategies primarily to energy-critical oxygen cycle electrocatalytic reactions, including oxygen reduction (ORR), but we have also considered extensions to non-electrochemical chemistries such as ammonia synthesis and decomposition. We have made strong progress in the development of computational methods on both the level of machine learning methods development as well as first principles-based treatments of HEA’s, and we have leveraged these insights to propose promising HEA catalysts for the ORR. On the experimental side, we developed new HEA synthesis and characterization protocols relevant to these reactions and developed a database combining our experimental results with corresponding computational tools.

36 MATERIALS SCIENCE

Introducing Molecular Hypernetworks for Discovery in Multidimensional Metabolomics Data

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and address challenges of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used in the Global Natural Products Social Molecular Networking platform, are a prominent strategy for exploring and visualizing molecular relationships and improving annotation. MNs are mathematical graphs showing the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated with a single molecular identity. Here, this paper introduces “molecular hypernetworks” (MHNs) as more complex MN models able to natively represent multiway relationships among observations. Compared to MNs, MHNs can more parsimoniously represent the inherent complexity present among groups of observations, initially supporting improved exploratory data analysis and visualization. MHNs also promise to increase confidence in annotation propagation, for both human and analytical processing. We first illustrate MHNs with simple examples, and build them from liquid chromatography- and ion mobility spectrometry-separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Derivation of physical equations for high-speed laser welding using large language models

It is challenging to formulate complex physical phenomena that occur in a manufacturing process, particularly when the available data are limited, rendering conventional data-driven approaches ineffective. This study aims to predict humping onset in high-speed laser welding by introducing a novel framework, namely text-to-equations generative pre-trained transformer (T2EGPT). This method leverages the capabilities of large language models (LLMs), in combination with sparse experimental data and enriched literature data, to derive an interpretable and generalizable equation for predicting humping initiation. By capturing key correlations among physical parameters, T2EGPT generates a compact and dimensionless expression that accurately predicts hump formation. The equation reveals that humping arises from the interplay between inertia-driven backward melt flow and capillary-driven surface stabilization, where inertial forces drive molten metal backward and capillary forces resist surface deformation. Furthermore, compared to traditional data-driven models, T2EGPT demonstrates enhanced predictive accuracy and cross-material transferability. More broadly, this study highlights the potential of LLMs to integrate textual information with data-driven discovery, enabling the extraction of physical laws in data-scarce scientific domains.

36 MATERIALS SCIENCE

Data mining the missing ordered phases of Li/Na metal oxides

Data-driven discovery of Li-ion and Na-ion battery materials has been pioneered by generic materials data platforms such as the Materials Project. After decades of progress, it is timely to ask whether there remain underexplored compositional spaces. Here, in this work, we present a systematic data-mining effort to uncover missing ordered binary, ternary and quaternary Li/Na-containing metal oxides using high-throughput density functional theory (DFT). Building on 19,120 stable and metastable oxides entries from the Materials Project, we performed 13,245 additional calculations through isovalent substitutions of known ground states, experimentally reported compounds, and specific prototype structures. Our study identifies 36 new ground states within the GGA/GGA + U convex hull and 45 within the r 2 SCAN convex hull. Additionally, we identified 840 metastable compounds from GGA/GGA + U and 979 from r 2 SCAN that are absent in the present Materials Project databases. Moreover, we have tripled the metastable materials in compositional spaces with a molar ratio of cation/anion >1, highlighting the overlooked opportunities in this compositional space.

25 ENERGY STORAGE

SODAs: sparse optimization for the discovery of differential and algebraic equations

Differential-algebraic equations (DAEs) integrate ordinary differential equations (ODEs) with algebraic constraints, providing a fundamental framework for developing models of dynamical systems characterized by time-scale separation, conservation laws and physical constraints. While sparse optimization has revolutionized model development by allowing data-driven discovery of parsimonious models from a library of possible equations, existing approaches for dynamical systems assume DAEs can be reduced to ODEs by eliminating variables before model discovery. This assumption limits the applicability of such methods for DAE systems with unknown constraints and time scales. We introduce sparse optimization for differential-algebraic systems (SODAs), a data-driven method for the identification of DAEs in their explicit form. By discovering the algebraic and dynamic components sequentially without prior identification of the algebraic variables, this approach leads to a sequence of convex optimization problems. It has the advantage of discovering interpretable models that preserve the structure of the underlying physical system. To this end, SODAs improves since SODAs is singular numerical stability when handling high correlations between library terms, caused by near-perfect algebraic relationships, by iteratively refining the conditioning of the candidate library. We demonstrate the performance of our method on biological, mechanical and electrical systems, showcasing its robustness to noise in both simulated time series and real-time experimental data.

DAE

Anomaly Detection and Approximate Similarity Searches of Transients in Real-time Data Streams

Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.

79 ASTRONOMY AND ASTROPHYSICS

Application of Machine Learning and Data Augmentation Algorithms in the Discovery of Metal Hydrides for Hydrogen Storage

The development of efficient and sustainable hydrogen storage materials is a key challenge for realizing hydrogen as a clean and flexible energy carrier. Among various options, metal hydrides offer high volumetric storage density and operational safety, yet their application is limited by thermodynamic, kinetic, and compositional constraints. In this work, we investigate the potential of machine learning (ML) to predict key thermodynamic properties—equilibrium plateau pressure, enthalpy, and entropy of hydride formation—based solely on alloy composition using Magpie-generated descriptors. We significantly expand an existing experimental dataset from ~400 to 806 entries and assess the impact of dataset size and data augmentation, using the PADRE algorithm, on model performance. Models including Support Vector Machines and Gradient Boosted Random Forests were trained and optimized via grid search and cross-validation. Results show a marked improvement in predictive accuracy with increased dataset size, while data augmentation benefits are limited to smaller datasets and do not improve accuracy in underrepresented pressure regimes. Furthermore, clustering and cross-validation analyses highlight the limited generalizability of models across different material classes, though high accuracy is achieved when training and testing within a single hydride family (e.g., AB2). The study demonstrates the viability and limitations of ML for accelerating hydride discovery, emphasizing the importance of dataset diversity and representation for robust property prediction.

augmentation

Machine-Learning-Driven Discovery of Water Splitting BaFe 2 O 4 and Human-in-the-Loop Improvement via Al-Substitution for Increased Thermal Stability

Thermochemical hydrogen (TCH) production offers a promising method for converting thermal energy into hydrogen fuel through heat-driven redox cycles of metal oxides. Here, in this work a defect graph neural network (dGNN) was used to predict oxygen vacancy formation energies ΔH V O combined with Materials Project predictions of oxygen chemical potential stability to screen candidate oxides via high-throughput database analysis. BaFe 2 O 4 was identified as a promising material for experimental validation based on its predicted ΔH V O , oxygen chemical potential stability range, and potential for tunable substitutions to improve thermal properties. Experimental validation using thermogravimetric analysis (TGA), stagnation flow reactor (SFR), X-ray diffraction (XRD), and electron microscopy confirmed positive water-splitting behavior but also revealed limitations in thermal stability under aggressive reduction conditions. To address this, a human-in-the-loop modification strategy was employed introducing Al substitution in BaFe 2–x Al x O 4 ; this modification improves thermal stability, alters the crystal structure and enhances overall performance. These results demonstrate a combined computational and experimental workflow in which machine learning accelerates identification of promising candidates, while targeted experimental design enables optimization of functional performance. This approach advances the development of robust, cost-effective TCH materials and highlights the importance of integrating data-driven discovery with human-guided materials design in paving the way for scalable hydrogen production technologies.

organic