Search NASASearch

SEARCH · Search NASA

Results for “functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Autonomous phase mapping of gold nanoparticles synthesis with differentiable models of spectral shape

Autonomous experimentation–or self-driving labs–offers a systematic approach to accelerate materials discovery by integrating automated synthesis, characterization, and data-driven decision-making. We present a closed-loop workflow for the on-demand synthesis and structural characterization of colloidal gold nanoparticles, enabling direct mapping from composition to nanoscale structure. Our framework leverages differentiable models of spectral shape to address two central tasks in self-driving labs: (a) phase mapping, or identifying compositional regions with distinct structural behavior; and (b) material retrosynthesis, or optimizing compositions for target structure. Using functional data analysis, we develop a data-driven model with generative pre-training, active learning, and high-throughput experiments to predict spectral responses across composition space. We demonstrate the approach on seed-mediated growth of gold nanoparticles, showcasing its ability to extract design rules, reveal secondary interactions, and efficiently navigate morphology space. Gradient-based optimization of the models enables inverse design, making this a unified platform.

36 MATERIALS SCIENCE

Data-driven analysis of dipole strength functions using artificial neural networks

Here, we present a data-driven analysis of dipole strength functions across the nuclear chart, employing an artificial neural network to model nuclear dipole responses. We train the network on a dataset of experimentally measured dipole strength functions for 216 different nuclei. To assess its predictive capability, we test the trained model on an additional set of 10 new nuclei, where experimental data exist. We demonstrate that the artificial neural network not only accurately reproduces known data but also identifies potential inconsistencies in experimental datasets, indicating which results may warrant further review or possible rejection. For nuclei where experimental data are sparse or unavailable, the network confirms theoretical calculations, reinforcing its utility as a predictive tool in nuclear physics. Finally, utilizing the predicted electric dipole polarizability, we extract the value of the symmetry energy at saturation density and find it consistent with results from the literature.

artificial neural networks

Elastic Bayesian Model Calibration

Functional data are ubiquitous in scientific modeling. For instance, quantities of interest are modeled as functions of time, space, energy, density, etc. Uncertainty quantification methods for computer models with functional response have resulted in tools for emulation, sensitivity analysis, and calibration that are widely used. However, many of these tools do not perform well when the computer model’s parameters control both the amplitude variation of the functional output and its alignment (or phase variation). This paper introduces a framework for Bayesian model calibration when the model responses are misaligned functional data. The approach generates two types of data out of the misaligned functional responses: (1) aligned functions so that the amplitude variation is isolated and (2) warping functions that isolate the phase variation. These two types of data are created for the computer simulation data (both of which may be emulated) and the experimental data. The calibration approach uses both types so that it seeks to match both the amplitude and phase of the experimental data. The framework is careful to respect constraints that arise, especially when modeling phase variation, and is framed in a way that it can be done with readily available calibration software. In conclusion, we demonstrate the techniques on two simulated data examples and on two dynamic material science problems: a strength model calibration using flyer plate experiments and an equation of state model calibration using experiments performed on the Sandia National Laboratories’ Z-machine.

97 MATHEMATICS AND COMPUTING

Uncertainty Quantification for Smooth Functional Data with Application to Material Properties

This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.

36 MATERIALS SCIENCE

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis

ampworks: Battery analysis tools in Python [SWR-25-39]

Ampworks is a collection of tools designed to process experimental battery data with a focus on model-relevant analyses. It currently provides functions for incremental capacity analysis and GITT data processing, helping extract key properties for life and physics-based models (e.g., SPM and P2D). Some tools, like the incremental capacity analysis module, also include graphical user interfaces for ease of use. https://github.com/NREL/ampworks/ https://pypi.org/project/ampworks/

Randall, Corey [National Renewable Energy Laborato

Pursuing Heteroleptic Ligand Design Principles for Photoactive Fe Complexes with Ultrafast X-ray Emission and Variable-Temperature Optical Spectroscopies

Understanding the key parameters that govern the photophysical and photochemical properties of transition metal complexes is essential for the development of efficient photosensitizers for photocatalytic applications. Achieving this objective necessitates clear and detailed investigations of their electronic excited states, for which time-resolved metal Kβ X-ray emission spectroscopy (XES) has proven highly effective. Here, we present a time-resolved Fe Kβ XES study of a heteroleptic Fe(II) polypyridyl carbene complex, [Fe(phen) 2 (C 4 H 10 N 4 )] 2+ (1; phen = 1,10-phenanthroline), utilizing both the valence-to-core and Kβ mainline spectral regions, complemented by variable-temperature transient optical absorption (VT-TA) spectroscopy. Detailed analysis of the time-resolved Kβ XES data, supported by density functional theory (DFT) calculations and an Eyring analysis of the VT-TA data, reveals parallel excited state relaxation dynamics that support an assignment of the long-lived excited state to a triplet metal-centered state. Placing these results in the context of prior studies of heteroleptic Fe(II) polypyridyl cyanide complexes motivated a series of DFT calculations to investigate the effects of ligand structural flexibility and arrangement. These calculations reinforce the experimentally derived conclusion that constraining structural flexibility with multidentate ligands significantly impacts the excited state relaxation dynamics. Furthermore, our study emphasizes that the arrangement of strong field ligands in heteroleptic complexes substantially affects the energy of Jahn–Teller active triplet metal-centered states in low-spin d 6 metal complexes. Together, these findings provide synthetic design principles for extending metal-to-ligand charge transfer excited state lifetimes of heteroleptic Fe complexes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology

Strangeness in the proton from $W+$ charm production and SIDIS data

We perform a global QCD analysis of unpolarized parton distribution functions (PDFs) in the proton, including new 𝑊+⁢ charm production data from 𝑝⁢𝑝 collisions at the LHC and semi-inclusive pion and kaon production data in lepton-nucleon deep-inelastic scattering, both of which have been suggested for constraining the strange quark PDF. Compared with a baseline global fit that does not include these datasets, the new analysis reduces the uncertainty on the strange quark distribution over the range 0.01 < 𝑥 < 0.3, and provides a consistent description of processes sensitive to strangeness in the proton. Including the new datasets, the ratio of strange to nonstrange sea quark distributions is $R_s = (s + \bar{s})/(\bar{u} +\bar{d})$ $=$ {$0.7⁢2^{+0.52}_{−0.34}, 0.4⁢6^{+0.30}_{−0.20}, 0.3⁢2^{+0.23}_{−0.15}$} for 𝑥 ={$0.01, 0.04, 0.1$} at 𝑄 2 $=$ 4 GeV 2 . The data place more stringent constraints on the strange asymmetry $(s - \bar{s})$, which is found to be consistent with zero in this range.

Anderson, Trey [College of William and Mary, Willi

Studying baryon acoustic oscillations using photometric redshifts from the DESI Legacy Imaging survey DR9

Context. The Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Survey DR9 (DR9 hereafter), with its extensive dataset of galaxy locations and photometric redshifts, presents an opportunity to study baryon acoustic oscillations (BAOs) in the region covered by the ongoing spectroscopic survey with DESI. Aims. We aim to investigate differences between different parts of the DR9 footprint. Furthermore, we want to measure the BAO scale for luminous red galaxies within them. Our selected redshift range of 0.6–0.8 corresponds to the bin in which a tension between DESI Y1 and eBOSS was found. Methods. We calculated the anisotropic two-point correlation function in a modified binning scheme to detect the BAOs in DR9 data. We then used template fits based on simulations to measure the BAO scale in the imaging data. Results. Our analysis reveals the expected correlation function shape in most of the footprint areas, showing a BAO scale consistent with Planck’s observations. Aside from identified mask-related data issues in the southern region of the South Galactic Cap, we find a notable variance between the different footprints. Conclusions. We find that this variance is consistent with the difference between the DESI Y1 and eBOSS data, and it supports the argument that that tension is caused by sample variance. Additionally, we also uncovered systematic biases not previously accounted for in photometric BAO studies. We emphasize the necessity of adjusting for the systematic shift in the BAO scale associated with typical photometric redshift uncertainties to ensure accurate measurements.

79 ASTRONOMY AND ASTROPHYSICS

Validation of the DESI DR2 measurements of baryon acoustic oscillations from galaxies and quasars

The Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) galaxy and quasar clustering data represents a significant expansion of data from Data Release 1 (DR1), providing improved statistical precision in baryon acoustic oscillation (BAO) constraints across multiple tracers, including bright galaxies, luminous red galaxies, emission line galaxies, and quasars. In this paper, we validate the BAO analysis of DR2. We present the results of robustness tests on the blinded DR2 data and, after unblinding, consistency checks on the unblinded DR2 data. All results are compared with those obtained from a suite of mock catalogs that replicate the selection and clustering properties of the DR2 sample. We confirm the consistency of DR2 BAO measurements with DR1 while achieving a reduction in statistical uncertainties due to the increased survey volume and completeness. The combined BAO precision, including both statistical and systematic errors, improves from ∼0.52% in DR1 to 0.30% in DR2—a factor of 1.7 gain. We assess the impact of analysis choices, including different data vectors (correlation function vs power spectrum), modeling approaches and systematics treatments, and an assumption of the Gaussian likelihood, finding that our BAO constraints are stable across these variations and assumptions with a few minor refinements to the baseline setup of the DR1 BAO analysis. We summarize a series of pre-unblinding tests that confirmed the readiness of our analysis pipeline, the final systematic errors, and the DR2 BAO analysis baseline. The successful completion of these tests led to the unblinding of the DR2 BAO measurements, ultimately leading to the DESI DR2 cosmological analysis, with their implications for the expansion history of the Universe and the nature of dark energy presented in the DESI key paper (companion paper).

79 ASTRONOMY AND ASTROPHYSICS

Results from a multi-laboratory ocean metaproteomic intercomparison: effects of LC-MS acquisition and data analysis procedures

Metaproteomics is an increasingly popular methodology that provides information regarding the metabolic functions of specific microbial taxa and has potential for contributing to ocean ecology and biogeochemical studies. A blinded multi-laboratory intercomparison was conducted to assess comparability and reproducibility of taxonomic and functional results and their sensitivity to methodological variables. Euphotic zone samples from the Bermuda Atlantic Time-series Study (BATS) in the North Atlantic Ocean collected by in situ pumps and the autonomous underwater vehicle (AUV) Clio were distributed with a paired metagenome, and one-dimensional (1D) liquid chromatographic data-dependent acquisition mass spectrometry analysis was stipulated. Analysis of mass spectra from seven laboratories through a common bioinformatic pipeline identified a shared set of 1056 proteins from 1395 shared peptide constituents. Quantitative analyses showed good reproducibility: pairwise regressions of spectral counts between laboratories yielded R 2 values averaged 0.62±0.11, and a Sørensen similarity analysis of the top 1000 proteins revealed 70 %–80 % similarity between laboratory groups. Taxonomic and functional assignments showed good coherence between technical replicates and different laboratories. A bioinformatic intercomparison study, involving 10 laboratories using eight software packages, successfully identified thousands of peptides within the complex metaproteomic datasets, demonstrating the utility of these software tools for ocean metaproteomic research. Lessons learned and potential improvements in methods were described. Future efforts could examine reproducibility in deeper metaproteomes, examine accuracy in targeted absolute quantitation analyses, and develop standards for data output formats to improve data interoperability. Together, these results demonstrate the reproducibility of metaproteomic analyses and their suitability for microbial oceanography research, including integration into global-scale ocean surveys and ocean biogeochemical models.

59 BASIC BIOLOGICAL SCIENCES

Accessing bands with extended quantum metric in kagome Cs 2 Ni 3 S 4 through soft chemical processing

Flat bands that do not merely arise from weak interactions can produce exotic physical properties, such as superconductivity or correlated many-body effects. The quantum metric can differentiate whether flat bands will result in correlated physics or are merely dangling bonds. A potential avenue for achieving correlated flat bands involves leveraging geometrical constraints within specific lattice structures, such as the kagome lattice; however, materials are often more complex. In these cases, quantum geometry becomes a powerful indicator of the nature of bands with small dispersions. We present a simple, soft-chemical processing route to access a flat band with an extended quantum metric below the Fermi level. By oxidizing Ni-kagome material Cs 2 Ni 3 S 4 to CsNi 3 S 4 , we see a two orders of magnitude drop in the room temperature resistance. However, CsNi 3 S 4 is still insulating, with no evidence of a phase transition. Using experimental data, density functional theory calculations, and symmetry analysis, our results suggest the emergence of a correlated insulating state of unknown origin.

Science & Technology - Other Topics

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Computational insights into hydrogen adsorption energies on medium-entropy oxides

High entropy oxides (HEOs) have emerged as promising catalysts for several important chemical transformations including alkane activation. Hydrogen adsorption energy (HAE) has been used as a key descriptor for many reactions including methane C–H activation and hydrogen evolution reactions. Hence, understanding the relationship between HAEs and the surface chemistry of HEO surfaces could lay the foundation for meaningful correlations among methane C–H activation, HAE, and the complex, local environment of HEO surfaces. Here, we used a medium-entropy oxide as a prototypical system – Mg 0.25 Ni 0.25 Cu 0.25 Zn 0.25 O with a rock-salt structure – to interrogate these relationships. We sampled 2000 different surfaces of its (100) plane and calculated the HAEs at randomly chosen surface O sites using density functional theory (DFT). Our analysis of the 2000 data points reveals that the HAEs at the surface O sites are significantly influenced by the local environment around the adsorption sites, particularly the nature of the metal atom directly below the surface O site where H adsorbs. After comparing several popular graph-neural-network-based machine learning models, we found that the DimeNet++ model performed best achieving satisfactory accuracy in predicting HAEs for both Mg 0.25 Ni 0.25 Cu 0.25 Zn 0.25 O and slightly varied compositions. Our work underscores the promise of such models and the need for further refinement to address the complexity of HEOs.

Song, Haohong [Vanderbilt Univ., Nashville, TN (Un

Window Observables for Benchmarking Parton Distribution Functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel “window observables” that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-𝑥, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different window observables that can be defined within a region of 𝑥 where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

lattice QCD

Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing

Quantum computing technology holds substantial promise as a reliable computational paradigm. However, current noisy intermediate scale quantum (NISQ) systems, are significantly impacted by noise originating from hardware inconsistencies. This noise causes errors and lowers output fidelity. So we must find which basis states cause errors. However, there are two main challenges in analyzing noise corresponding to basis states. First, the noise distribution data is high dimensional in nature, thereby making its analysis challenging. Second, although functional box plots have been used in the state of the art research to understand such a high dimensional data, they suffer from clutter and occlusion issues because of overplotting. In this study, we introduce an innovative visualization pipeline to address the aforementioned challenges to provide a clear depiction of noisy and less-noisy basis states. Specifically, our proposed visualization pipeline comprises three stages namely, low dimensional embedding, clustering, and violin plot visualization, to reduce visual clutter and effectively analyze high-dimensional noise distribution data. Our analysis uses quantum machine learning (QML) circuits as case study for drawing a distinction between noisy and less noisy basis states.

Senapati, Priyabrata [Kent State University]

Window observables for benchmarking parton distribution functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel ``window observables'' that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-x, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different ``window observables'' that can be defined within a region of x where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS