Search NASASearch

SEARCH · Search NASA

Results for “functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Window observables for benchmarking parton distribution functions

Global analysis of collider and fixed-target experimental data and calculations from lattice quantum chromodynamics (QCD) are used to gain complementary information on the structure of hadrons. We propose novel ``window observables'' that allow for higher precision cross-validation between the different approaches, a critical step for studies that wish to combine the datasets. Global analyses are limited by the kinematic regions accessible to experiment, particularly in a range of Bjorken-x, and lattice QCD calculations also have limitations requiring extrapolations to obtain the parton distributions. We provide two different ``window observables'' that can be defined within a region of x where extrapolations and interpolations in global analyses remain reliable and where lattice QCD results retain sensitivity and precision.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Uncertainty Visualization of Critical Points of 2D Scalar Fields for Parametric and Nonparametric Probabilistic Models

This paper presents a novel end-to-end framework for closed-form computation and visualization of critical point uncertainty in 2D uncertain scalar fields. Critical points are fundamental topological descriptors used in the visualization and analysis of scalar fields. The uncertainty inherent in data (e.g., observational and experimental data, approximations in simulations, and compression), however, creates uncertainty regarding critical point positions. Uncertainty in critical point positions, therefore, cannot be ignored, given their impact on downstream data analysis tasks. Here, in this work, we study uncertainty in critical points as a function of uncertainty in data modeled with probability distributions. Although Monte Carlo (MC) sampling techniques have been used in prior studies to quantify critical point uncertainty, they are often expensive and are infrequently used in production-quality visualization software. We, therefore, propose a new end-to-end framework to address these challenges that comprises a threefold contribution. First, we derive the critical point uncertainty in closed form, which is more accurate and efficient than the conventional MC sampling methods. Specifically, we provide the closed-form and semianalytical (a mix of closed-form and MC methods) solutions for parametric (e.g., uniform, Epanechnikov) and nonparametric models (e.g., histograms) with finite support. Second, we accelerate critical point probability computations using a parallel implementation with the VTK-m library, which is platform portable. Finally, we demonstrate the integration of our implementation with the ParaView software system to demonstrate near-real-time results for real datasets.

97 MATHEMATICS AND COMPUTING

DESI DR1 Ly$α$ forest: 3D full-shape analysis and cosmological constraints

We perform an analysis of the full shapes of Lyman-$α$ (Ly$α$) forest correlation functions measured from the first data release (DR1) of the Dark Energy Spectroscopic Instrument (DESI). Our analysis focuses on measuring the Alcock-Paczynski (AP) effect and the cosmic growth rate times the amplitude of matter fluctuations in spheres of $8$$h^{-1}\text{Mpc}$, $fσ_8$. We validate our measurements using two different sets of mocks, a series of data splits, and a large set of analysis variations, which were first performed blinded. Our analysis constrains the ratio $D_M/D_H(z_\mathrm{eff})=4.525\pm0.071$, where $D_H=c/H(z)$ is the Hubble distance, $D_M$ is the transverse comoving distance, and the effective redshift is $z_\mathrm{eff}=2.33$. This is a factor of $2.4$ tighter than the Baryon Acoustic Oscillation (BAO) constraint from the same data. When combining with Ly$α$ BAO constraints from DESI DR2, we obtain the ratios $D_H(z_\mathrm{eff})/r_d=8.646\pm0.077$ and $D_M(z_\mathrm{eff})/r_d=38.90\pm0.38$, where $r_d$ is the sound horizon at the drag epoch. We also measure $fσ_8(z_\mathrm{eff}) = 0.37\; ^{+0.055}_{-0.065} \,(\mathrm{stat})\, \pm 0.033 \,(\mathrm{sys})$, but we do not use it for cosmological inference due to difficulties in its validation with mocks. In $Λ$CDM, our measurements are consistent with both cosmic microwave background (CMB) and galaxy clustering constraints. Using a nucleosynthesis prior but no CMB anisotropy information, we measure the Hubble constant to be $H_0 = 68.3\pm 1.6\;\,{\rm km\,s^{-1}\,Mpc^{-1}}$ within $Λ$CDM. Finally, we show that Ly$α$ forest AP measurements can help improve constraints on the dark energy equation of state, and are expected to play an important role in upcoming DESI analyses.

Cuceu, Andrei [LBL, Berkeley; Chicago U., KICP] (O

Extending quantum-mechanical benchmark accuracy to biological ligand-pocket interactions

Predicting the binding affinity of ligands to protein pockets is key in the drug design pipeline. The flexibility of ligand-pocket motifs arises from a range of attractive and repulsive electronic interactions during binding. Accurately accounting for all interactions requires robust quantum-mechanical (QM) benchmarks, which are scarce for ligand-pocket systems. Additionally, disagreement between “gold standard” Coupled Cluster (CC) and Quantum Monte Carlo (QMC) methods casts doubt on many benchmarks for larger non-covalent systems. We introduce the “QUantum Interacting Dimer” (QUID) benchmark framework containing 170 non-covalent (non-)equilibrium systems modeling chemically and structurally diverse ligand-pocket motifs. Symmetry-adapted perturbation theory shows that QUID broadly covers non-covalent binding motifs and energetic contributions. Robust binding energies are obtained using complementary CC and QMC methods, achieving agreement of 0.5 kcal/mol. The benchmark data analysis reveals that several dispersion-inclusive density functional approximations provide accurate energy predictions, though their atomic van der Waals forces differ in magnitude and orientation. Contrarily, semiempirical methods and empirical force fields require improvements in capturing non-covalent interactions (NCIs) for out-of-equilibrium geometries. The wide span of NCIs, highly accurate interaction energies, and analysis of molecular properties take QUID beyond the “gold standard” for QM benchmarks of ligand-protein systems.

Puleva, Mirela [University of Luxembourg, Luxembou

Deviations from the Porter-Thomas Distribution due to Nonstatistical 𝛾 Decay below the 150 Nd Neutron Separation Threshold

We introduce a new method for the study of fluctuations of partial transition widths based on nuclear resonance fluorescence experiments with quasimonochromatic linearly polarized photon beams below particle separation thresholds. It is based on the average branching of decays of 𝐽=1 states of an even-even nucleus to the 2$^{+}_{1}$ state in comparison to the ground state. Between 5 and 7 MeV, a constant average branching ratio for 𝛾 decays from 1 − states of 0.490(16) is observed for the nuclide 150 Nd. Assuming 𝜒 2 -distributed partial transition widths, this average branching ratio is related to a degree of freedom of 𝜈 = 1.93⁢(12), rejecting the validity of the Porter-Thomas distribution, requiring 𝜈 = 1. The observed deviation can be explained by nonstatistical effects in the 𝛾-decay behavior with contributions in the range of 9.4(10)% up to 94(10)%.

150 ≤ A ≤ 189

Strong coupling from hadronic τ -decay data including τ → π − π 0 ν τ from Belle

In previous work we have combined the π − π 0 , 2 π − π + π 0 , and π − 3 π 0 spectral data obtained from hadronic τ decays measured by the ALEPH and OPAL experiments, together with electroproduction data for several of the subleading hadronic modes and data for the K K ¯ mode to construct an inclusive nonstrange vector spectral function entirely based on experimental data, with no Monte-Carlo generated input. In this paper, we include, for the first time, the Belle τ → π − π 0 ν τ high-statistics decay data to construct a new inclusive nonstrange vector spectral function that combines more of the world’s available data. As no Belle data are at present available for the two 4 π modes, this requires a revised data analysis in comparison with our previous work. From the resulting new spectral function, we obtain a new determination of the strong coupling, α s , using our previously developed strategy based on finite-energy sum rules. We find, at the Z mass scale, α s ( m Z 2 ) = 0.1159 ( 14 ) . We discuss the smaller central value and larger error of our new result compared to our previous result, showing the shifts to be due mainly to significant changes in updated HFLAV results for the π − 3 π 0 decay mode. Published by the American Physical Society 2025

Boito, Diogo (ORCID:0000000244267984)

Data for Spatial Analysis of Cell Patterning to Aid Genetic and Phenotypic Understanding of Grass Stomatal Density: A Case Study in Maize

Biological processes involve complex hierarchies where composite traits result from multiple component traits. However, holistically understanding of how sets of component traits interact to underpin genotype-to-phenotype relationships is generally lacking. Stomatal density (SD) is a tractable model system for exploring how high-throughput phenotyping (HTP) data could be exploited by a new spatial analysis approach to better understand a developmentally and functionally important trait. SD is a composite trait, resulting from various components related to cell identity and size, which are themselves governed by a series of spatio-developmental processes. Data from 192 recombinant inbred lines of maize [Zea mays (L.)] were analyzed by a new stomatal patterning phenotype (SPP) to (1) describe the average spatial probability distribution of the nearest neighboring stomata; (2) derive a core set of component traits related to cell size, cell packing, and positional probabilities; (3) build a structural equation model of component traits underlying SD; and (4) identify stomatal patterning quantitative trait loci (QTL). The core set of SPP-derived traits explained 74% of the variation in SD. Analyzing SPP component traits allowed some loci previously identified as generic SD QTL to be recognized as specific to lateral versus longitudinal elements of stomatal patterning. Therefore, this study highlights how novel insights can be gained by decomposing a composite trait (e.g., SD) into a set of component traits that were present in HTP data but not previously exploited.

AI/ML

Short and medium range structure in elastic deformation of metallic and covalent glasses

Here, we present a concise methodology to analyze structural response to the applied stress in amorphous solids, including metallic glasses (MG), glassy selenium, silica and polycarbonate, using high energy x-ray diffraction and atomic pair distribution function (PDF) analysis. To assess the structural anisotropy induced by applied axial stress, diffraction data were expanded into spherical harmonics. Using Bessel transformation, components of the structure function were converted into isotropic and anisotropic PDFs. The PDFs were compared to the expected model behavior for ideal elastic deformation to separate homogeneous affine strain from local non-affine strains. In metallic glass the range of non-affine deformation is limited to the nearest neighbor shell, suggesting local strain relaxation under stress that occurs even in the elastic regime. Beyond the second atomic shell strain is uniform. However, in glassy silica, polycarbonate and selenium strong local bonding inhibits local displacements and strain in short range order is accommodated by rotation of local units. Interestingly, beyond a molecular unit, deformation in covalent systems is similar to MG, and response of the medium range order scales with the macroscopic stress.

glassy structure

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES

A Life Cycle Analysis Framework for Point Source Capture Systems

NETL studies the costs and benefits of PSC for electricity, industry, and mobile applications. Mobile point source capture (MPSC) and storage applied to freight modes captures emissions directly from exhaust. This poster presents a framework for conducting LCA of PSC systems applied to heavy-duty trucks, freight trains, and marine vessels. The framework defines a wheels-to-storage (gate-to-grave) boundary, including energy demands (electricity, heat, and cooling requirements), solvent use and cycling, onboard system components, carbon storage in a saline aquifer, and upstream manufacturing impacts for equipment, with a suggested functional unit of 1 tonne-km. Potential data sources for analysis include material, energy, and operational data from Oak Ridge National Laboratory, GREET (Greenhouse gases, Regulated Emissions, and Energy use in Technologies) model, and scientific literature. The suggested analytical approach includes comparison to publicly available business-as-usual systems without capture across all modes of transportation, sensitivity to composition of the capture solvent, and sensitivity to capture rate variation, all of which would support a wholistic PSC business case analysis. For future consideration, analysis can be augmented with consideration of different sources of electricity (e.g., nuclear), fuel substitution, deploying supportive infrastructure such as pipeline offloading points, and downstream applications like enhanced oil recovery (EOR).

life cycle analysis (LCA)

Spin structure of the proton from global QCD analysis

In this talk we review recent results for spin-dependent parton distribution functions extracted in global QCD analysis of high energy scattering data by the JAM collaboration, including inclusive and semi-inclusive deep-inelastic scattering, jet and weak boson production in polarised hadron-hadron collisions. In particular, we focus on the determination of the gluon polarisation in the proton, whose sign and magnitude have been the subject of debate recently.

Melnitchouk, Wally [Thomas Jefferson National Acce

A road map to cosmological parameter analysis with third-order shear statistics: III. Efficient estimation of third-order shear correlation functions and an application to the KiDS-1000 data

Context. Third-order lensing statistics contain a wealth of cosmological information that is not captured by second-order statistics. However, the computational effort it takes to estimate such statistics in forthcoming stage IV surveys is prohibitively expensive. Aims. We derive and validate an efficient estimation procedure for the three-point correlation function (3PCF) of polar fields such as weak lensing shear. We then use our approach to measure the shear 3PCF and the third-order aperture mass statistics on the KiDS-1000 survey. Methods We constructed an efficient estimator for third-order shear statistics that builds on the multipole decomposition of the 3PCF. We then validated our estimator on mock ellipticity catalogs obtained from N -body simulations. Finally, we applied our estimator to the KiDS-1000 data and presented a measurement of the third-order aperture statistics in a tomographic setup. Results. Our estimator provides a speedup of a factor of ∼100–1000 compared to the state-of-the-art estimation procedures. It is also able to provide accurate measurements for squeezed and folded triangle configurations without additional computational effort. We report a significant detection of tomographic third-order aperture mass statistics in the KiDS-1000 data (S/N = 6.69). Conclusions. Our estimator will make it computationally feasible to measure third-order shear statistics in forthcoming stage IV surveys. Furthermore, it can be used to construct empirical covariance matrices for such statistics.

Astronomy & Astrophysics

Isospin dependence of the nuclear EMC effect from a global QCD analysis

We perform a new global QCD analysis of unpolarized parton distribution functions (PDFs) in the nucleon from proton, deuteron, and A = 3 data, including recent measurements of He 3 / D and H 3 / D cross section ratios from the MARATHON experiment at Jefferson Lab. Simultaneously inferring the PDFs and nucleon off-shell corrections allows both to be determined consistently, without theoretical assumptions about the isospin dependence of nuclear effects. The analysis provides strong evidence for the need of nucleon off-shell corrections to describe the A = 3 data, with large isoscalar and a suggestion of nonzero isovector contributions in A ≤ 3 nuclei. We find that the extracted EMC ratios of nuclear to nucleon structure functions for A = 2 and 3 differ from those naively extrapolated from heavy nuclei down to low A .

Cocuzza, C. [William & Mary] (ORCID:00000003492292

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio

Effect of solvothermal synthesis parameters on the crystallite size and atomic structure of cobalt iron oxide nanoparticles

We here investigate how the synthesis method affects the crystallite size and atomic structure of cobalt iron oxide nanoparticles. By using a simple solvothermal method, we first synthesized cobalt ferrite nanoparticles of ca. 2 and 7 nm, characterized by Transmission Electron Microscopy (TEM), Small Angle X-ray scattering (SAXS), X-ray and neutron total scattering. The smallest particle size corresponds to only a few spinel unit cells. Nevertheless, Pair Distribution Function (PDF) analysis of X-ray and neutron total scattering data shows that the atomic structure, even in the smallest nanoparticles, is well described by the spinel structure, although with significant disorder and a contraction of the unit cell parameter. These effects can be explained by the surface oxidation of the small nanoparticles, which is confirmed by X-ray near edge absorption spectroscopy (XANES). Neutron total scattering data and PDF analysis reveal a higher degree of inversion in the spinel structure of the smallest nanoparticles. Neutron total scattering data also allow magnetic PDF (mPDF) analysis, which shows that the ferrimagnetic domains correspond to ca. 80% of the crystallite size in the larger particles. A similar but less well-defined magnetic ordering was observed for the smallest nanoparticles. Finally, we used a co-precipitation synthesis method at room temperature to synthesize ferrite nanoparticles similar in size to the smallest crystallites synthesized by the solvothermal method. Structural analysis with PDF demonstrates that the ferrite nanoparticles synthesized via this method exhibit a significantly more defective structure compared to those synthesized via a solvothermal method.

77 NANOSCIENCE AND NANOTECHNOLOGY

Asymptotic inconsistency of the cumulative algorithm for laser-induced damage probability analysis

The “cumulative algorithm” is a data analysis method that has been proposed to provide an objective, nonparametric determination of laser-induced damage probability as a function of fluence from experimental data that contain both damaged sites and undamaged sites (i.e., 1-on-1 or S-on-1 testing protocols). In this work, the limitations of this approach are explored by considering the asymptotic limit of a large number of test sites. It is shown that the cumulative algorithm does not converge to the true probability distribution and significantly underestimates the damage probability near the damage onset. Here, based on the results of this work, the cumulative algorithm is not recommended for accurate estimation of damage probability.

Computational methods

Machine-learning-enabled on-the-fly analysis of RHEED patterns during thin film deposition by molecular beam epitaxy

Thin film deposition is a fundamental technology for the discovery, optimization, and manufacturing of functional materials. Deposition by molecular beam epitaxy (MBE) typically employs reflection high-energy electron diffraction (RHEED) as a real-time in situ probe of the growing film. However, the state-of-the-art for RHEED analysis during deposition requires human observation. Here, we present an approach using machine learning (ML) methods to monitor, analyze, and interpret RHEED images on-the-fly during thin film deposition. In the analysis workflow, RHEED pattern images are collected at one frame per second and featurized using a pretrained deep convolutional neural network. The feature vectors are then statistically analyzed to identify changepoints; these changepoints can be related to changes in the deposition mode from initial film nucleation to a transition regime, smooth film deposition, and in some cases, an additional transition to a rough, islanded deposition regime. The feature vectors are additionally analyzed via graph analysis and community classification. The graph is quantified as a stabilization plot, and we show that inflection points in the stabilization plot correspond to changes in the growth regime. The full RHEED analysis workflow is termed RHAAPsody and includes data transfer and output to a visual dashboard. We demonstrate the functionality of RHAAPsody by analyzing the precaptured RHEED images from epitaxial depositions of anatase TiO2 on SrTiO3(001) and show that the analysis workflow can be executed in less than 1 s. Our approach shows promise as one component of ML-enabled real-time feedback control of the MBE deposition process.

36 MATERIALS SCIENCE

Spatial analysis of cell patterning to aid genetic and phenotypic understanding of grass stomatal density: A case study in maize

Biological processes involve complex hierarchies where composite traits result from multiple component traits. However, holistically understanding of how sets of component traits interact to underpin genotype-to-phenotype relationships is generally lacking. Stomatal density (SD) is a tractable model system for exploring how high-throughput phenotyping (HTP) data could be exploited by a new spatial analysis approach to better understand a developmentally and functionally important trait. SD is a composite trait, resulting from various components related to cell identity and size, which are themselves governed by a series of spatio-developmental processes. Data from 192 recombinant inbred lines of maize [Zea mays (L.)] were analyzed by a new stomatal patterning phenotype (SPP) to (1) describe the average spatial probability distribution of the nearest neighboring stomata; (2) derive a core set of component traits related to cell size, cell packing, and positional probabilities; (3) build a structural equation model of component traits underlying SD; and (4) identify stomatal patterning quantitative trait loci (QTL). The core set of SPP-derived traits explained 74% of the variation in SD. Analyzing SPP component traits allowed some loci previously identified as generic SD QTL to be recognized as specific to lateral versus longitudinal elements of stomatal patterning. Therefore, this study highlights how novel insights can be gained by decomposing a composite trait (e.g., SD) into a set of component traits that were present in HTP data but not previously exploited.

59 BASIC BIOLOGICAL SCIENCES