Search NASA⌕ Search

SEARCH · Search NASA

Results for “Functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Functional Data Analysis in Wearable Body Sensor Networks

Improving response time of indirect room-size calorimeters is still an outstanding problem in metabolic research. Accurate estimates of instantaneous rates of gaseous exchange require numerical differentiation of measured gaseousgas concentrations. We propose a new method to estimate the instantaneous gaseousgas exchange rates in indirect calorimetry. In contrast to the previously developed techniques, the method addresses the problem of differentiation of gaseous concentrations as an ill-posed problem. By applying the method of regularization, the problem of differentiation is converted into a well-posed problem resulting in smooth and consistent gaseous exchange rates. The validity of the method is tested on a large dataset of calorimeter experiments which included 313 human experiments along with 231 alcohol combustion experiments. It is demonstrated that the method is able to reliably differentiate between the “unphysiological” process of alcohol combustion and physiological variations produced by human metabolism. The method also allowed unraveling the previously unreported relative kinetics of O2 consumption and Respiratory Quotient (RQ) in humans. It was found that the kinetics of oxidative fuel selection lags behind the energy expenditure in humans exhibiting some sort of oxidative inertia. The time lag varies from 2-3 min up to 30 min, depending on particular individual. No such lag was found in alcohol combustion experiments. In addition to the relative kinetics of substrate oxidation, two statistical indexes reflecting variability of minute-by-minute RQ were estimated. The indexes were the RQ’s standard deviation and RQ’s first-order derivative. Both indexes showed statistically significant difference between human experiments and alcohol combustion experiments. We conclude that the proposed method can consistently extract physiologically-relevant information from noisy calorimetry data and the aforesaid information can provide additional insights into the mechanism of metabolic fuel selection in humans.

54 ENVIRONMENTAL SCIENCES↗

Functional Data Analysis in Wearable Body Sensor Networks

Improving response time of indirect room-size calorimeters is still an outstanding problem in metabolic research. Accurate estimates of instantaneous rates of gaseous exchange require numerical differentiation of measured gaseousgas concentrations. We propose a new method to estimate the instantaneous gaseousgas exchange rates in indirect calorimetry. In contrast to the previously developed techniques, the method addresses the problem of differentiation of gaseous concentrations as an ill-posed problem. By applying the method of regularization, the problem of differentiation is converted into a well-posed problem resulting in smooth and consistent gaseous exchange rates. The validity of the method is tested on a large dataset of calorimeter experiments which included 313 human experiments along with 231 alcohol combustion experiments. It is demonstrated that the method is able to reliably differentiate between the “unphysiological” process of alcohol combustion and physiological variations produced by human metabolism. The method also allowed unraveling the previously unreported relative kinetics of O2 consumption and Respiratory Quotient (RQ) in humans. It was found that the kinetics of oxidative fuel selection lags behind the energy expenditure in humans exhibiting some sort of oxidative inertia. The time lag varies from 2-3 min up to 30 min, depending on particular individual. No such lag was found in alcohol combustion experiments. In addition to the relative kinetics of substrate oxidation, two statistical indexes reflecting variability of minute-by-minute RQ were estimated. The indexes were the RQ’s standard deviation and RQ’s first-order derivative. Both indexes showed statistically significant difference between human experiments and alcohol combustion experiments. We conclude that the proposed method can consistently extract physiologically-relevant information from noisy calorimetry data and the aforesaid information can provide additional insights into the mechanism of metabolic fuel selection in humans.

60 - APPLIED LIFE SCIENCES↗

Inverse prediction of PuO2 processing conditions using Bayesian seemingly unrelated regression with functional data

Over the past decade, a variety of innovative methodologies have been developed to better characterize the relationships between processing conditions and the physical, morphological, and chemical features of special nuclear material (SNM). Different processing conditions generate SNM products with different features, which are known as “signatures” because they are indicative of the processing conditions used to produce the material. These signatures can potentially allow a forensic analyst to determine which processes were used to produce the SNM and make inferences about where the material originated. This article investigates a statistical technique for relating processing conditions to the morphological features of PuO 2 particles. We develop a Bayesian implementation of seemingly unrelated regression (SUR) to inverse-predict unknown PuO 2 processing conditions from known PuO 2 features. Model results from simulated data demonstrate the usefulness of the technique. Applied to empirical data from a bench-scale experiment specifically designed with inverse prediction in mind, our model successfully predicts nitric acid concentration, while results for Pu concentration and precipitation temperature were equivalent to a simple mean model. Our technique compliments other recent methodologies developed for forensic analysis of nuclear material and can be generalized across the field of chemometrics for application to other materials.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Autonomous phase mapping of gold nanoparticles synthesis with differentiable models of spectral shape

Autonomous experimentation–or self-driving labs–offers a systematic approach to accelerate materials discovery by integrating automated synthesis, characterization, and data-driven decision-making. We present a closed-loop workflow for the on-demand synthesis and structural characterization of colloidal gold nanoparticles, enabling direct mapping from composition to nanoscale structure. Our framework leverages differentiable models of spectral shape to address two central tasks in self-driving labs: (a) phase mapping, or identifying compositional regions with distinct structural behavior; and (b) material retrosynthesis, or optimizing compositions for target structure. Using functional data analysis, we develop a data-driven model with generative pre-training, active learning, and high-throughput experiments to predict spectral responses across composition space. We demonstrate the approach on seed-mediated growth of gold nanoparticles, showcasing its ability to extract design rules, reveal secondary interactions, and efficiently navigate morphology space. Gradient-based optimization of the models enables inverse design, making this a unified platform.

36 MATERIALS SCIENCE↗

Data-driven analysis of dipole strength functions using artificial neural networks

Here, we present a data-driven analysis of dipole strength functions across the nuclear chart, employing an artificial neural network to model nuclear dipole responses. We train the network on a dataset of experimentally measured dipole strength functions for 216 different nuclei. To assess its predictive capability, we test the trained model on an additional set of 10 new nuclei, where experimental data exist. We demonstrate that the artificial neural network not only accurately reproduces known data but also identifies potential inconsistencies in experimental datasets, indicating which results may warrant further review or possible rejection. For nuclei where experimental data are sparse or unavailable, the network confirms theoretical calculations, reinforcing its utility as a predictive tool in nuclear physics. Finally, utilizing the predicted electric dipole polarizability, we extract the value of the symmetry energy at saturation density and find it consistent with results from the literature.

artificial neural networks↗

Elastic Bayesian Model Calibration

Functional data are ubiquitous in scientific modeling. For instance, quantities of interest are modeled as functions of time, space, energy, density, etc. Uncertainty quantification methods for computer models with functional response have resulted in tools for emulation, sensitivity analysis, and calibration that are widely used. However, many of these tools do not perform well when the computer model’s parameters control both the amplitude variation of the functional output and its alignment (or phase variation). This paper introduces a framework for Bayesian model calibration when the model responses are misaligned functional data. The approach generates two types of data out of the misaligned functional responses: (1) aligned functions so that the amplitude variation is isolated and (2) warping functions that isolate the phase variation. These two types of data are created for the computer simulation data (both of which may be emulated) and the experimental data. The calibration approach uses both types so that it seeks to match both the amplitude and phase of the experimental data. The framework is careful to respect constraints that arise, especially when modeling phase variation, and is framed in a way that it can be done with readily available calibration software. In conclusion, we demonstrate the techniques on two simulated data examples and on two dynamic material science problems: a strength model calibration using flyer plate experiments and an equation of state model calibration using experiments performed on the Sandia National Laboratories’ Z-machine.

97 MATHEMATICS AND COMPUTING↗

Uncertainty Quantification for Smooth Functional Data with Application to Material Properties

This document outlines a method for processing functional output (i.e., curves) for the ultimate purpose of sampling curves under specified input conditions for use in modeling and simulation uncertainty quantification (UQ) studies. A set of benchmark curves sufficiently representative of the relevant scenario(s) being simulated are provided to the process and formatted as described in Section 1. Principal Component Analysis (PCA) is utilized to discover the components of uncertainty in the benchmark curves and is outlined in Section 2. Section 3 describes the application of uncertainty quantification to the PCA results for the purpose of sampling curves to be used in UQ analysis. Section 4 applies these techniques to an example benchmark dataset. Concluding remarks are provided in the final section.

36 MATERIALS SCIENCE↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Pursuing Heteroleptic Ligand Design Principles for Photoactive Fe Complexes with Ultrafast X-ray Emission and Variable-Temperature Optical Spectroscopies

Understanding the key parameters that govern the photophysical and photochemical properties of transition metal complexes is essential for the development of efficient photosensitizers for photocatalytic applications. Achieving this objective necessitates clear and detailed investigations of their electronic excited states, for which time-resolved metal Kβ X-ray emission spectroscopy (XES) has proven highly effective. Here, we present a time-resolved Fe Kβ XES study of a heteroleptic Fe(II) polypyridyl carbene complex, [Fe(phen) 2 (C 4 H 10 N 4 )] 2+ (1; phen = 1,10-phenanthroline), utilizing both the valence-to-core and Kβ mainline spectral regions, complemented by variable-temperature transient optical absorption (VT-TA) spectroscopy. Detailed analysis of the time-resolved Kβ XES data, supported by density functional theory (DFT) calculations and an Eyring analysis of the VT-TA data, reveals parallel excited state relaxation dynamics that support an assignment of the long-lived excited state to a triplet metal-centered state. Placing these results in the context of prior studies of heteroleptic Fe(II) polypyridyl cyanide complexes motivated a series of DFT calculations to investigate the effects of ligand structural flexibility and arrangement. These calculations reinforce the experimentally derived conclusion that constraining structural flexibility with multidentate ligands significantly impacts the excited state relaxation dynamics. Furthermore, our study emphasizes that the arrangement of strong field ligands in heteroleptic complexes substantially affects the energy of Jahn–Teller active triplet metal-centered states in low-spin d 6 metal complexes. Together, these findings provide synthetic design principles for extending metal-to-ligand charge transfer excited state lifetimes of heteroleptic Fe complexes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗

Strangeness in the proton from $W+$ charm production and SIDIS data

We perform a global QCD analysis of unpolarized parton distribution functions (PDFs) in the proton, including new 𝑊+⁢ charm production data from 𝑝⁢𝑝 collisions at the LHC and semi-inclusive pion and kaon production data in lepton-nucleon deep-inelastic scattering, both of which have been suggested for constraining the strange quark PDF. Compared with a baseline global fit that does not include these datasets, the new analysis reduces the uncertainty on the strange quark distribution over the range 0.01 < 𝑥 < 0.3, and provides a consistent description of processes sensitive to strangeness in the proton. Including the new datasets, the ratio of strange to nonstrange sea quark distributions is $R_s = (s + \bar{s})/(\bar{u} +\bar{d})$ $=$ {$0.7⁢2^{+0.52}_{−0.34}, 0.4⁢6^{+0.30}_{−0.20}, 0.3⁢2^{+0.23}_{−0.15}$} for 𝑥 ={$0.01, 0.04, 0.1$} at 𝑄 2 $=$ 4 GeV 2 . The data place more stringent constraints on the strange asymmetry $(s - \bar{s})$, which is found to be consistent with zero in this range.

Anderson, Trey [College of William and Mary, Willi↗

Studying baryon acoustic oscillations using photometric redshifts from the DESI Legacy Imaging survey DR9

Context. The Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Survey DR9 (DR9 hereafter), with its extensive dataset of galaxy locations and photometric redshifts, presents an opportunity to study baryon acoustic oscillations (BAOs) in the region covered by the ongoing spectroscopic survey with DESI. Aims. We aim to investigate differences between different parts of the DR9 footprint. Furthermore, we want to measure the BAO scale for luminous red galaxies within them. Our selected redshift range of 0.6–0.8 corresponds to the bin in which a tension between DESI Y1 and eBOSS was found. Methods. We calculated the anisotropic two-point correlation function in a modified binning scheme to detect the BAOs in DR9 data. We then used template fits based on simulations to measure the BAO scale in the imaging data. Results. Our analysis reveals the expected correlation function shape in most of the footprint areas, showing a BAO scale consistent with Planck’s observations. Aside from identified mask-related data issues in the southern region of the South Galactic Cap, we find a notable variance between the different footprints. Conclusions. We find that this variance is consistent with the difference between the DESI Y1 and eBOSS data, and it supports the argument that that tension is caused by sample variance. Additionally, we also uncovered systematic biases not previously accounted for in photometric BAO studies. We emphasize the necessity of adjusting for the systematic shift in the BAO scale associated with typical photometric redshift uncertainties to ensure accurate measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

Validation of the DESI DR2 measurements of baryon acoustic oscillations from galaxies and quasars

The Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) galaxy and quasar clustering data represents a significant expansion of data from Data Release 1 (DR1), providing improved statistical precision in baryon acoustic oscillation (BAO) constraints across multiple tracers, including bright galaxies, luminous red galaxies, emission line galaxies, and quasars. In this paper, we validate the BAO analysis of DR2. We present the results of robustness tests on the blinded DR2 data and, after unblinding, consistency checks on the unblinded DR2 data. All results are compared with those obtained from a suite of mock catalogs that replicate the selection and clustering properties of the DR2 sample. We confirm the consistency of DR2 BAO measurements with DR1 while achieving a reduction in statistical uncertainties due to the increased survey volume and completeness. The combined BAO precision, including both statistical and systematic errors, improves from ∼0.52% in DR1 to 0.30% in DR2—a factor of 1.7 gain. We assess the impact of analysis choices, including different data vectors (correlation function vs power spectrum), modeling approaches and systematics treatments, and an assumption of the Gaussian likelihood, finding that our BAO constraints are stable across these variations and assumptions with a few minor refinements to the baseline setup of the DR1 BAO analysis. We summarize a series of pre-unblinding tests that confirmed the readiness of our analysis pipeline, the final systematic errors, and the DR2 BAO analysis baseline. The successful completion of these tests led to the unblinding of the DR2 BAO measurements, ultimately leading to the DESI DR2 cosmological analysis, with their implications for the expansion history of the Universe and the nature of dark energy presented in the DESI key paper (companion paper).

79 ASTRONOMY AND ASTROPHYSICS↗

Results from a multi-laboratory ocean metaproteomic intercomparison: effects of LC-MS acquisition and data analysis procedures

Metaproteomics is an increasingly popular methodology that provides information regarding the metabolic functions of specific microbial taxa and has potential for contributing to ocean ecology and biogeochemical studies. A blinded multi-laboratory intercomparison was conducted to assess comparability and reproducibility of taxonomic and functional results and their sensitivity to methodological variables. Euphotic zone samples from the Bermuda Atlantic Time-series Study (BATS) in the North Atlantic Ocean collected by in situ pumps and the autonomous underwater vehicle (AUV) Clio were distributed with a paired metagenome, and one-dimensional (1D) liquid chromatographic data-dependent acquisition mass spectrometry analysis was stipulated. Analysis of mass spectra from seven laboratories through a common bioinformatic pipeline identified a shared set of 1056 proteins from 1395 shared peptide constituents. Quantitative analyses showed good reproducibility: pairwise regressions of spectral counts between laboratories yielded R 2 values averaged 0.62±0.11, and a Sørensen similarity analysis of the top 1000 proteins revealed 70 %–80 % similarity between laboratory groups. Taxonomic and functional assignments showed good coherence between technical replicates and different laboratories. A bioinformatic intercomparison study, involving 10 laboratories using eight software packages, successfully identified thousands of peptides within the complex metaproteomic datasets, demonstrating the utility of these software tools for ocean metaproteomic research. Lessons learned and potential improvements in methods were described. Future efforts could examine reproducibility in deeper metaproteomes, examine accuracy in targeted absolute quantitation analyses, and develop standards for data output formats to improve data interoperability. Together, these results demonstrate the reproducibility of metaproteomic analyses and their suitability for microbial oceanography research, including integration into global-scale ocean surveys and ocean biogeochemical models.

59 BASIC BIOLOGICAL SCIENCES↗

Elastic Changepoint Detection for Globally-indexed Functional Time Series Data with Climate Applications

Changepoint detection is a vital tool in the application of climate data analysis. Numerous types of climate observation data are most properly represented by functional time series, implying a need for accurate changepoint detection methods applicable to functional time series data. Such data taken at a global scale often contain both spatial heterogeneity and dependence as well as phase (time) misalignment. In this report, we present methods which can detect spatially-dependent changepoints while allowing different estimates of change time and change strength depending on location. Additionally, we provide extensions to this spatially-predicted model which controls for phase variability among observations. Our methods provide the ability to detect a single change, or control for epidemic changes (where a “return-to-normal” change is more likely to be detected than the initial change). We showcase results analyzing the June 1991 eruption of Mt. Pinatubo, where our methods demonstrate the ability to accurately detect both single and epidemic changepoints even in the presence of strong seasonal variability. We find that our spatially-predicted model improves the detection of relevant changepoints versus methods which do not take spatial information into account, and we find that controlling for phase variability helps to control the false discovery rate during the detection process.

54 ENVIRONMENTAL SCIENCES↗

Analysis of parton distributions in a pion with Bézier parametrizations

We explore the role of parametrizations for nonperturbative QCD functions in global analyses, with a specific application to extending a phenomenological analysis of the parton distribution functions (PDFs) in the charged pion realized in the xFitter fitting framework. The parametrization dependence of PDFs in our pion fits substantially enlarges the uncertainties from the experimental sources estimated in the previous analyses. We systematically explore the parametrization dependence by employing a novel technique to automate generation of polynomial parametrizations for PDFs that makes use of Bézier curves. This technique is implemented in a ++ module that is included in the xFitter program. Our analysis reveals that the sea and gluon distributions in the pion are not well disentangled, even when considering measurements in leading-neutron deep inelastic scattering. For example, the pion PDF solutions with a vanishing gluon and large quark sea are still experimentally allowed, which elevates the importance of ongoing lattice and nonperturbative QCD calculations, together with the planned pion scattering experiments, for conclusive studies of the pion structure. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Structure and phase transitions in niobium and tantalum derived nanoscale transition metal perovskites, Ba(Ti,MV)O3, M=Nb,Ta

The prospect of creating ferroelectric or high permittivity nanomaterials provides motivation for investigating complex transition metal oxides of the form Ba(Ti, MV)O3, where M = Nb or Ta. Solid state processing typically produces mixtures of crystalline phases, rarely beyond minimally doped Nb/Ta. Using a modified sol-gel method, we prepared single phase nanocrystals of Ba(Ti, M)O3. Compositional and elemental analysis puts the empirical formulas close to BaTi0.5Nb0.5O3−δ and BaTi0.5Ta0.5O3−δ. For both materials, a reversible temperature dependent phase transition (non-centrosymmetric to symmetric) is observed in the Raman spectrum in the region 533–583 K (260–310 °C); for Ba(Ti, Nb)O3, the onset is at 543 K (270 °C); and for Ba(Ti, Ta)O3, the onset is at 533 K (260 °C), which are comparable with 390–393 K (117–120 °C) for bulk BaTiO3. The crystal structure was resolved by examination of the powder x-ray diffraction and atomic pair distribution function (PDF) analysis of synchrotron total scattering data. It was postulated whether the structure adopted at the nanoscale was single or double perovskite. Double perovskites (A2B′B″O6) are characterized by the type and extent of cation ordering, which gives rise to higher symmetry crystal structures. PDF analysis was used to examine all likely candidate structures and to look for evidence of higher symmetry. The feasible phase space that evolves includes the ordered double perovskite structure Ba2(Ti, MV)O6 (M = Nb, Ta) Fm-3m, a disordered cubic structure, as a suitable high temperature analog, Ba(Ti, MV)O3Pm-3m, and an orthorhombic Ba(Ti, MV)O3Amm2, a room temperature structure that presents an unusually high level of lattice displacement, possibly due to octahedral tilting, and indication of a highly polarized crystal.

Chemistry↗