Search NASASearch

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data-Driven Supervised Dimension Reduction for Scientific Discovery (LDRD QTI Report)

This report summarizes the findings of a four months FY24 Advanced Science & Technology (AS&T) LDRD Quick Targeted Investigation (QTI) project focused on the exploration of supervised dimension reduction approaches based on autoencoders. Autoencoders have been extensively employed in literature for unsupervised learning tasks, however, their use for supervised regression tasks, which are common within scientific applications, has been limited. Motivated by linear dimension reduction strategies like Active Subspaces and Adaptive Basis, we explored the possibility of employing autoencoders to discover a non-linear manifold able to represent the original function in fewer dimensions. In this report, we discuss a neural network architecture and we perform a numerical campaign on several problems ranging from simple two-dimensional functions to a model problem for magnetohydrodynamics in five dimensions. In our preliminary results, we show that the proposed approach is found to be superior to linear dimension reduction strategies in representing the target function even with a single latent variable.

97 MATHEMATICS AND COMPUTING

Enabling pan-repository reanalysis for big data science of public metabolomics data

Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, here we develop pan-repository universal identifiers and harmonized cross-repository metadata. This ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplified data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.

El Abiead, Yasin

Building the Eyes of Discovery

My presentation focuses on the development and enhancement of sensor modules that help particle detectors “see” particles invisible to the human eye, specifically the Compact Muon Shield detector, which is part of the High-Luminosity Large Hadron Collider. As a CCI intern at Fermilab’s Silicon Detector Facility, I worked on testing and visual inspection of different module parts before assembly and on testing the modules after assembly. My talk highlights how small sensors carry the capability to collect tiny particle signals within the detector and the reliability of control checks pre-assembly for scientists to produce reliable data for future discoveries.

Siddiqui, Hooriya [Fermilab] (ORCID:00090001510973

SEAFORML (Smart Exploration and Analysis For Optimal and Robust Machine Learning)

The poster discusses data analysis of the WAVgraph database and applied machine learning methods for it. The database is a long-term project that seeks to be a comprehensive repository of information on cyber threats and is updated regularly. It was previously unanalyzed and unexplored. The goal was to learn more about it and its contents in order to have a better understanding and enable better use. The data analysis and discovery enabled further exploration through natural language processing, similarity, and clustering methods. The poster shows some of the insights from the analysis and explains the methods used for the machine learning applications.

24 - POWER TRANSMISSION AND DISTRIBUTION

Changing effects of external forcing on Atlantic–Pacific interactions

Recent studies have highlighted the increasingly dominant role of external forcing in driving Atlantic and Pacific Ocean variability during the second half of the 20th century. This paper provides insights into the underlying mechanisms driving interactions between modes of variability over the two basins. We define a set of possible drivers of these interactions and apply causal discovery to reanalysis data, two ensembles of pacemaker simulations where sea surface temperatures in either the tropical Pacific or the North Atlantic are nudged to observations, and a pre-industrial control run. We also utilize large-ensemble means of historical simulations from the Coupled Model Intercomparison Project Phase 6 (CMIP6) to quantify the effect of external forcing and improve the understanding of its impact. A causal analysis of the historical time series between 1950 and 2014 identifies a regime switch in the interactions between major modes of Atlantic and Pacific climate variability in both reanalysis and pacemaker simulations. A sliding window causal analysis reveals a decaying El Niño–Southern Oscillation (ENSO) effect on the Atlantic as the North Atlantic fluctuates towards an anomalously warm state. The causal networks also demonstrate that external forcing contributed to strengthening the Atlantic's negative-sign effect on ENSO since the mid-1980s, where warming tropical Atlantic sea surface temperatures induce a La Niña-like cooling in the equatorial Pacific during the following season through an intensification of the Pacific Walker circulation. The strengthening of this effect is not detected when the historical external forcing signal is removed in the Pacific pacemaker ensemble. The analysis of the pre-industrial control run supports the notion that the Atlantic and Pacific modes of natural climate variability exert contrasting impacts on each other even in the absence of anthropogenic forcing. The interactions are shown to be modulated by the (multi)decadal states of temperature anomalies of both basins with stronger connections when these states are “out of phase”. We show that causal discovery can detect previously documented connections and provides important potential for a deeper understanding of the mechanisms driving changes in regional and global climate variability.

54 ENVIRONMENTAL SCIENCES

Commutative Algebra Modeling in Materials Science – A Case Study on Metal–Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry’s constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

Khaemba, Caleb S.

Accurate Dehydrogenation Enthalpies Dataset for Liquid Organic Hydrogen Carriers

This contribution presents a comprehensive extension of the QM9 dataset (originally at 133 K molecules) with the calculation of G4MP2 enthalpies for 9,841 molecules, featuring up to nine heavy atoms. We present QM9-LOHC, a (de)hydrogenation dataset of 10,373 reactions, including a minimum of 5.5% weight hydrogen storage capacity in line with the Department of Energy standards for Liquid Organic Hydrogen Carriers (LOHC). By utilizing the accurate quantum chemical method G4MP2 we expand the QM9 database and explore new avenues for the exploration of hydrogen storage technologies (electrochemical LOHCs, alkali metal-LOHCs, and mixtures of LOHCs). The QM9-LOHC dataset, with its focus on reactions that vary only by hydrogen saturation levels, provides a needed data resource for advancing the design and optimization of both conventional and innovative LOHC systems, and high-fidelity data for molecular discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat

In-pixel integration of signal processing and AI/ML based data filtering for particle tracking detectors

We present the first physical realization of in-pixel signal processing with integrated AI-based data filtering for particle tracking detectors. Building on prior work that demonstrated a physics-motivated edge-AI algorithm suitable for ASIC implementation, this work marks a significant milestone toward intelligent silicon trackers. Our prototype readout chip performs real-time data reduction at the sensor level while meeting stringent requirements on power, area, and latency. The chip is taped-out in 28nm TSMC CMOS bulk process, which has been shown to have sufficient radiation hardness for particle experiments. This development represents a key step toward enabling fully on-detector edge AI, with broad implications for data throughput and discovery potential in high-rate, high-radiation environments such as the High-Luminosity LHC.

Parpillon, Benjamin [Fermilab; Illinois U., Chicag

Protein Structure Inspired Discovery of a Novel Inducer of Anoikis in Human Melanoma

Drug discovery historically starts with an established function, either that of compounds or proteins. This can hamper discovery of novel therapeutics. As structure determines function, we hypothesized that unique 3D protein structures constitute primary data that can inform novel discovery. Using a computationally intensive physics-based analytical platform operating at supercomputing speeds, we probed a high-resolution protein X-ray crystallographic library developed by us. For each of the eight identified novel 3D structures, we analyzed binding of sixty million compounds. Top-ranking compounds were acquired and screened for efficacy against breast, prostate, colon, or lung cancer, and for toxicity on normal human bone marrow stem cells, both using eight-day colony formation assays. Effective and non-toxic compounds segregated to two pockets. One compound, Dxr2-017, exhibited selective anti-melanoma activity in the NCI-60 cell line screen. In eight-day assays, Dxr2-017 had an IC50 of 12 nM against melanoma cells, while concentrations over 2100-fold higher had minimal stem cell toxicity. Dxr2-017 induced anoikis, a unique form of programmed cell death in need of targeted therapeutics. Our findings demonstrate proof-of-concept that protein structures represent high-value primary data to support the discovery of novel acting therapeutics. This approach is widely applicable.

Oncology

JOINT APPOINTEE: Evolution of ferroelectric properties in SmxBi1-xFeO3 via automated Piezoresponse Force Microscopy across combinatorial spread libraries

Combinatorial spread libraries offer a innovative approach to explore the evolution of material properties over broad concentration, temperature, and growth parameter spaces. However, traditional limitation of this approach is the requirement for the read-out of functional properties across the library. Here we develop automated Piezoresponse Force Microscopy (PFM) for the exploration of combinatorial spread libraries and demonstrate its application in the SmxBi1-xFeO3 system with the ferroelectric-antiferroelectric morphotropic phase boundary. This approach relies on the synergy of the quantitative nature of PFM and the implementation of automated experiments that allow PFM-based sampling over macroscopic samples. The concentration dependence of pertinent ferroelectric parameters has been determined and used to develop the mathematical framework based on Ginzburg-Landau theory describing the evolution of these properties across the concentration space. We pose that a combination of automated scanning probe microscope and combinatorial spread library approach will emerge as an efficient research paradigm to close the characterization gap in the high-throughput materials discovery. We make the data sets open to the community and hope that this will stimulate other efforts to interpret and understand the physics of these systems.

Automated Microscopy, Combinatorial Library, Ferro

Radioisotope Identification with List-Mode Gamma-Ray Data

This work explores the potential of utilizing temporal data from gamma-ray detectors, known as list-mode data, to enhance radioisotope identification. Traditional identification methods, which rely on full gamma-ray spectrum analysis, often require long dwell times and struggle with spectra containing similarly spaced spectral peaks. We hypothesize that by leveraging the probabilistic nature of nuclear decay and the time-encoded information from decay sequences and interactions with surrounding materials, we can improve classification accuracy over static spectral analysis. This research examines the temporal content of list-mode data through exploratory data analysis via correlation discovery and qualitative distribution analysis. Additionally, we propose a probabilistic classification model that can utilize spectral data, temporal data, or both to determine if the incorporation of temporal information improves radioisotope identification. Our findings suggest that the temporal information present in list-mode gamma-ray data has merit and should be further investigated to develop more robust and optimal methods for utilizing this temporal information in applications requiring radioisotope identification.

List-mode data

Livewire Data Platform File Standards Version 1.0

This living document describes required and recommended standards for data additions to the Livewire Data Platform (https://livewire.energy.gov/). Adherence to the standards described enables development of automated analysis and discovery tools for Livewire data and will facilitate development of future capabilities for delivering data that can be tailored to meet user needs.

33 - ADVANCED PROPULSION SYSTEMS

Radioisotope Identification with List-Mode Gamma Ray Data: A rigorous assessment on the value of temporal information applied to radioisotope identification.

This work explores the potential of utilizing temporal data from gamma-ray detectors, known as list-mode data, to enhance radioisotope identification. Traditional identification methods, which rely on full gamma-ray spectrum analysis, often require long dwell times and struggle with “confuser” sources, or spectra with similarly spaced spectral peaks. We hypothesize that by leveraging the probabilistic nature of nuclear decay and the time-encoded information from decay sequences and interactions with surrounding materials, we can improve classification accuracy over static spectral analysis. This research rigorously examines the temporal content of list-mode data through exploratory data analysis via correlation discovery and information theory. We further propose a basic classification model that can utilize spectral or temporal data (or both) to determine if the incorporation of temporal information can improve radioisotope identification. The findings suggest that the temporal information present in list-mode gamma-ray data has merit and should be further investigated.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Ultra-faint Milky Way Satellites Discovered in Carina, Phoenix, and Telescopium with DELVE Data Release 3

We report the discovery of three Milky Way satellite candidates: Carina IV, Phoenix III, and DELVE 7, in the third data release of the DECam Local Volume Exploration survey (DELVE). The candidate systems were identified by cross-matching results from two independent search algorithms. All three are extremely faint systems composed of old, metal-poor stellar populations (τ ≳ 10 Gyr, [Fe/H] ≲−1.4). Carina IV (M V = −2.8; r 1/2 = 40 pc) and Phoenix III (M V = −1.2; r 1/2 = 19 pc) have half-light radii that are consistent with the known population of dwarf galaxies, while DELVE 7 (M V = 1.2; r 1/2 = 2 pc) is very compact and seems more likely to be a star cluster, though its nature remains ambiguous without spectroscopic follow-up. The Gaia proper motions of stars in Carina IV ($M_{\star} = 2250^{+1180}_{-830} M_⊙$) indicate that it is unlikely to be associated with the LMC, while DECam CaHK photometry confirms that its member stars are metal poor. Phoenix III ($M_{\star} = 520^{+660}_{-290} M_⊙$) is the faintest known satellite in the extreme outer stellar halo (D GC > 100 kpc), while DELVE 7 ($M_{\star} = 60^{+120}_{-40} M_⊙$) is the faintest known satellite with D GC > 20 kpc.

Tan, Chin Yi [Univ. of Chicago, IL (United States)

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING