Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A Science Gateway for the Repeatable Analysis of Machine Learning Predicted Gravity Anomalies

In recent years, deep learning has become an increasingly popular alternative for modeling in geoscience applications due to its scalability and efficiency. However, the interpretability, compute, data volume, and hyperparameter tuning requirements of deep learning models make development and monitoring difficult. Furthermore, model explainability and communicating results obtained by these models to users or domain experts is a challenge, as domain experts in geoscience also need to have a deep understanding of how those models function in order to support their scientific works. Here, we describe a science gateway and machine learning pipeline for predicting gravity anomalies from geophysical data. The gateway, built on open-source technologies, provides a holistic view of the pipeline through interactive visualizations aimed at enabling efficient exploratory data analysis. The repeatability, reproducibility, and monitoring capabilities of this overall system allow us to iterate and analyze at scale. Using this pipeline and gateway, we can repeatedly produce accurate high-resolution gravity anomaly datasets. By describing the underlying technologies, implementation, and results, here we provide a foundation for the broader adoption of science gateways into cross-cutting geoscience and machine learning research projects as a means to improve the scientific discovery and collaboration in the geophysics and computational sciences community.

58 GEOSCIENCES↗

RWRtoolkit: multi-omic network analysis using random walks on multiplex networks in any species

Abstract We introduce RWRtoolkit, a multiplex generation, exploration, and statistical package built for R and command-line users. RWRtoolkit enables the efficient exploration of large and highly complex biological networks generated from custom experimental data and/or from publicly available datasets, and is species agnostic. A range of functions can be used to find topological distances between biological entities, determine relationships within sets of interest, search for topological context around sets of interest, and statistically evaluate the strength of relationships within and between sets. The command-line interface is designed for parallelization on high-performance cluster systems, which enables high-throughput analysis such as permutation testing. Several tools in the package have also been made available for use in reproducible workflows via the KBase web application.

Kainer, David (ORCID:0000000172714676)↗

The landscape of regulatory element evolution in a C4 perennial grass

Gene regulatory evolution is a well-known source of phenotypic diversity and adaptive evolution. Although cis-regulatory elements (CREs) play a vital role in gene expression evolution, the molecular evolution of CREs remains mostly unknown due to the difficulty in identifying and characterizing these functional elements. Comparative genomic analyses of noncoding DNA can be leveraged to identify conserved noncoding sequences (CNS), many of which may harbor functional CREs conserved by purifying selection. However, purely computational inference of CREs from putative CNS can be erroneous due to the complex genomic architecture in plants. One promising experimental approach to identify CREs is by profiling accessible chromatin regions (ACRs) that are often associated with the location of CREs. In this study, we use comparative genomics along with the profiling of ACRs to study the molecular evolution of putative functional noncoding regulatory regions in Panicoid grasses. We identified sets of CNS that varied in relationship to the degree of evolutionary divergence among the studied taxa, including identifying core-Panicoid-CNS. We augmented this analysis by profiling ACRs in Panicum hallii ecotypes using ATAC-seq. ACRs had low SNP density at the summit, harbored a high frequency of core-Panicoid-CNS, and were enriched with expression QTL. These data help to annotate the P. hallii genome for putative functional elements and suggest that a large proportion of these ACRs are evolving under purifying selection. Turnover in CNS and ACR between ecotypes of P. hallii identifies a small set of putatively divergent CREs that may underlie differences in gene regulation between genotypes from inland and coastal habitats. In summary, we profiled ACRs in Panicoid grasses and integrated this data with our putative CNS prediction framework, which provides unique insight into patterns of polymorphism and divergence in CREs in C4 perennial grasses.

59 BASIC BIOLOGICAL SCIENCES↗

Constructing Data-Driven Predictions at the Far Detector for NOvA's Neutrino Oscillation Analysis.

NOvA, is a two-detector, long-baseline neutrino oscillation experiment located at Fermilab, Batavia, IL, USA. It is designed primarily to constrain neutrino oscillation parameters using $\nu_\mu \ (\bar{\nu}_\mu)$ disappearance and $\nu_e \ (\bar{\nu}_e)$ appearance data. The Neutrinos at Main Injector (NuMI) beamline at Fermilab provides a high purity 900 KW intense beam of neutrinos and anti-neutrinos to NOvA. The NOvA Near Detector, located 100m underground and 1km away from the beam source, observes the un-oscillated $\nu_\mu \ (\bar{\nu}_\mu)$ and beam $\nu_e \ (\bar{\nu}_e)$ event spectrum. The Far Detector, located in Ash River, MN, USA, is 809 km from the ND and records the oscillated $\nu_e \ (\bar{\nu}_e)$ and the un-oscillated $\nu_\mu \ (\bar{\nu}_\mu)$ event spectrum. NOvA uses a data-driven technique called extrapolation to predict the expected number of $\nu_\mu \ (\bar{\nu}_\mu)$ and $\nu_e \ (\bar{\nu}_e)$ events at the Far Detector using the Near Detector data. The use of data from a functionally equivalent Near Detector provides a powerful constraint on the systematic uncertainties in NOvA neutrino oscillation analyses. As NOvA continues to add data statistics, a robust constraint on systematics becomes more crucial for neutrino oscillation analysis. The details of the NOvA neutrino oscillation analysis framework and how it constrains dominant systematic uncertainties using the Near Detector data will be discussed in this poster.

43 PARTICLE ACCELERATORS↗

Glauber-Theory Calculations of High-Energy Nuclear Scattering Observables Using Variational Monte Carlo Wave Functions

Experiments using intermediate- to high-energy radioactive nuclear beams present numerous findings. Extracting important properties of physical observables relies on a firm theoretical analysis. Though Glauber theory is believed to work well, no convincing calculation has so far been done. Here, we perform ab initio Glauber theory calculations of both elastic differential cross sections and total reaction cross sections for p+ 12 C, 12 C+ 12 C, and 6 He+ 12 C systems. The wave functions of both 6 He and 12 C are generated by variational Monte Carlo calculations with spatial and spin-isospin correlations induced by realistic two- and three-nucleon potentials. Glauber’s phase-shift function is computed by Monte Carlo integration up to all orders of nucleon-nucleon multiple scatterings. We show an excellent performance of the Glauber description to the selected data on the above systems. We also find that the cumulant expansion of the phase-shift function converges rapidly up to the second order for the above systems. This finding will open up interesting applications for the analysis of high-energy nuclear experiments.

Horiuchi, W. [Osaka Metropolitan University (Japan↗

Analysis of Tar and Oil Derived from Pyrolysis and Copyrolysis of Waste Plastics and Biomass

Pyrolysis has been proposed as a potential technology for managing the growing volume of plastic waste generated worldwide. Co-pyrolysis of plastic waste with biomass is a promising technology for generating fuel and chemical products. However, this process generates tar as a waste product. The chemical properties of this tar have yet to be thoroughly analyzed. Further, this study presents the results of gas chromatography–mass spectrometry (GC–MS), Fourier-transform infrared spectroscopy (FTIR), and thermogravimetric analysis (TGA) of oil and tar obtained from the pyrolysis of pure plastics including high-density polyethylene (HDPE), low-density polyethylene (LDPE), polyethylene (PE), polystyrene (PS), and plastic-biomass mixtures. GC–MS analysis revealed the presence of C 7 –C 37 carbon-containing hydrocarbons, which include alkanes and alkenes as the dominant products. FTIR data revealed the presence of various functional groups, including alcohols, aldehydes, ketones, and carboxylic acids, indicating the complexity of the pyrolysis and copyrolysis oil obtained from waste plastics and biomass. TGA data show that tar from all four plastics has a higher decomposition rate, suggesting the presence of heavier hydrocarbons compared with their corresponding oils. This research will be of interest to researchers looking to advance the study of plastic and biomass waste management.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Single-cell proteomics of Arabidopsis leaf mesophyll reveals dynamic protein responses to water-deficit stress

Background The application of single-cell omics tools to biological systems can provide unique insights into diverse cellular populations and their heterogeneous responses to internal and external perturbations. Thus far, most single-cell studies in plant systems have been limited to RNA-sequencing approaches, which only provide indirect readouts of cellular functions. Results Here, we present a single-cell proteomics workflow for plant cells that integrates tape-sandwich protoplasting, piezoelectric cell sorting, nanoPOTS sample preparation, and ion mobility-based MS data acquisition method for label-free single-cell proteomics analysis of Arabidopsis leaf mesophyll cells. From a single leaf protoplast, over 3,000 proteins were quantified with high precision. The workflow is demonstrated to identify stress associated changes in protein abundance by analyzing 117 protoplasts from well-watered and water-deficit stressed plants. Additionally, we describe a new approach for constructing covarying protein networks at the single-cell level and demonstrate how single-cell protein covariation analysis can reveal previously unrecognized protein functions while also capturing stress-induced changes in protein–protein dynamics. Conclusions The label-free scProteomic approach presented here represents a significant advance through the demonstration of a facile protoplast isolation method combined with deep and precise proteomic coverage of Arabidopsis leaf mesophyll cell types. We believe this study will serve as an informative reference to future plant scProteomic investigations.

Arabidopsis↗

ASCR Workshop Position Paper: Challenges and Opportunities in High Energy Physics

High energy particle physics and cosmology concern themselves with estimating fundamental parameters of nature, such as the masses and interactions of fundamental particles like the Higgs boson and the rate of expansion of the universe. In doing so, they analyze exabyte-scale datasets, some of the largest in all of science, and face many challenges in subsequent data analysis. These challenges are shared between the two disciplines, but we focus on particle physics to highlight one specific domain. In particle physics, the standard method for estimating parameters involves performing Monte Carlo (MC) integration as a function of both parameters of interest and nuisance parameters using an expensive simulator, counting the number of observed collision events (i.i.d. samples) from an experiment in the corresponding integration domains, and forming a Poisson likelihood function. This likelihood function is then used in a Frequentist manner to construct a maximum likelihood point estimate (MLE) and confidence set for the parameters. To sufficiently populate the high-dimensional integration domains, simulators consume billions of CPU-hours annually and produce hundreds of petabytes of intermediate output data. Several techniques have been developed to: optimize definitions of the integration domains so as to be maximally sensitive to a particular subset of parameters, efficiently estimate the integrals, and build robust surrogate models by interpolating between integral evaluations at different parameter points. One can view this whole endeavor as classical Simulation-Based Inference (SBI).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Meta-analysis of North American Arctic and boreal aboveground biomass datasets: assessing accuracy, dynamics, and similarities

The North American arctic and boreal regions (ABRs) are rapidly warming and experiencing intensifying disturbances. Accurately quantifying aboveground biomass (AGB) is critical for understanding the impacts of these changes on the carbon cycle and for designing climate change mitigation strategies. Several AGB maps have been developed for the North American ABRs, including recent contributions from National Aeronautics and Space Administration’s Arctic-Boreal Vulnerability Experiment (ABoVE) campaign. However, these maps differ widely in training data, methodology, and resulting AGB density estimates. Presently, a comprehensive comparative evaluation is lacking, making it difficult for users to select datasets suited to their research or management needs. Here, in this study, we conducted a comparative analysis of nine AGB density datasets across North American ABRs, specifically for Alaska and Canada. We (1) summarized AGB by ecoregion and Canadian provinces, (2) evaluated their accuracy against field-based measurements, (3) analyzed spatial and temporal similarities among datasets, and (4) assessed their ability to capture disturbance (fire and harvest) impacts on AGB. We found substantial variation in regional and local AGB estimates across datasets, with overall accuracy ranging from R 2 = 0.25–0.62 and Bias% from −47.8% to 69.9% when validated against field plots. Despite these differences, most datasets have comparatively consistent spatial patterns in AGB (r > 0.8 for most cases). In contrast, agreement on the temporal patterns of AGB change is generally low. We found datasets with spatial resolutions ⩽300 m are capable of capturing disturbance impacts on AGB dynamics, though sensitivity varies across products. Our findings and dataset summary provide guidance for selecting appropriate AGB datasets for different applications within our study area. Our analysis also highlights the need to decrease map bias and increase capability to detect temporal change to decrease uncertainty of AGB datasets potentially by using training data which is representative of major plant functional types within the mapped area.

ABoVE↗

Investigating Quantum Materials with Half-Polarized Diffraction and magnetic PDF analysis at the HB-2A Neutron Powder Diffractometer

Local magnetic ordering and anisotropy is often central to the emergent behavior and subsequent functional properties in quantum materials and beyond. Neutron powder diffraction provides a straightforward yet extremely powerful technique for quantitative measurements of microscopic magnetic properties. The HB-2A powder diffractometer located at the High Flux Isotope Reactor in ORNL is traditionally utilized for long-range magnetic structure determination. Recently these capabilities have been extended to include methods aimed at accessing local magnetism: Half- polarized neutron powder diffraction (pNPD) and magnetic pair distribution function (mPDF) analysis. These two distinct techniques are possible on HB-2A due to the versatility of the instrument’s reciprocal space coverage, resolution and novel ultra-low temperature multi-sample changers that operate down to dilution refrigerator temperatures. This provides unique capabilities not found on any powder diffraction instrument and is particularly well suited to investigations of magnetic quantum materials. The development and implementation of these techniques will be discussed with a series of science case examples ranging from geometric frustrated magnets to magnetic metal-organic frameworks. Data reduction and analysis tools will be presented that enable the extraction of the local site susceptibility tensor and local spin-spin correlations in real space. Finally, potential combinations of these techniques in the form of half-polarized magnetic pair distribution function (pmPDF) analysis will be considered. Looking forward, HB-2A is undergoing a detector upgrade that will be in the user program by 2026. This will offer an order of magnitude increase in count rates to further aid the development of these often low signal measurements and provide new scientific capabilities.

Neutron Scattering↗

Dark Energy Survey Year 3 Results: Cosmological constraints from second- and third-order shear statistics

Here, we present a cosmological analysis of the third-order aperture mass statistic using Dark Energy Survey Year 3 (DES Y3) data. We perform a complete tomographic measurement of the three-point correlation function of the Y3 weak lensing shape catalog with the four fiducial source redshift bins. Building upon our companion methodology paper, we apply a pipeline that combines the two-point function ξ ± with the mass aperture skewness statistic ⟨ M ap 3 ⟩ , which is an efficient compression of the full shear three-point function. We use a suite of simulated shear maps to obtain a joint covariance matrix. By jointly analyzing ξ ± and ⟨ M ap 3 ⟩ measured from DES Y3 data with a Λ CDM model, we find S 8 = 0.780 ± 0.015 and Ω m = 0.26 6 - 0.040 + 0.039 , yielding 111% of figure-of-merit improvement in the Ω m - S 8 plane relative to ξ ± alone, consistent with expectations from simulated likelihood analyses. With a w CDM model, we find S 8 = 0.74 9 - 0.026 + 0.027 and w 0 = - 1.39 ± 0.31 , which gives an improvement of 22% on the joint S 8 - w 0 constraint. Our results are consistent with w 0 = - 1 . Our new constraints are compared to CMB data from the Planck satellite, and we find that with the inclusion of ⟨ M ap 3 ⟩ the existing tension between the datasets is at the level of 2.3 σ . We show that the third-order statistic enables us to self-calibrate the mean photometric redshift uncertainty parameter of the highest redshift bin with little degradation in the figure of merit. Our results demonstrate the constraining power of higher-order lensing statistics and establish ⟨ M ap 3 ⟩ as a practical observable for joint analyses in current and future surveys.

Gomes, R. C. H. [University of Pennsylvania] (ORCI↗

Lost and Found: Rediscovering Microbiome-Associated Phenotypes that Reshape Agricultural Sustainability

Overview Code and data repository for NIL Manuscript. Documentation includes sequence processing examples and data analysis. Supplemental sequence processing and R statistical analysis for publication, which compares the microbiome of teosinte-B73 Near Isogenic Lines. Sample Data Amplicon sequence data for 16S rRNA genes, the fungal ITS2 region, and nitrogen-cycling functional genes are available through the NCBI Sequence Read Archive (SRA) under accession number PRJNA1042643(https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1042643). Raw metabolomic data are available on Metabolomics Workbench, Project ID: PR002654. This study is available at the NIH Common Fund's National Metabolomics Data Repository (NMDR) website, the Metabolomics Workbench, https://www.metabolomicsworkbench.org where it has been assigned Study ID ST004211. The data can be accessed directly via its Project DOI: http://dx.doi.org/10.21228/M8KV8T.

Near Isogeneic Lines↗

How does drought affect residential water demand and price elasticity?

Urban water scarcity is an important social and economic concern, particularly as the intensity, duration, and frequency of droughts is increasing in many regions. We consider whether drought induces changes to water demand and the price elasticity of demand for water that may last beyond a drought’s official end date. If drought shocks prompt long-term changes in water demand behavior, and these changes occur at broad geographic scale, they could have important implications for modeling adaptive responses to water scarcity. We assemble a novel dataset on residential water demand and pricing in the western United States to test empirically for effects of drought on water demand and price elasticity. We perform our analysis with aggregate quantity, price, and drought data, accounting for endogenous prices under increasing-block water tariffs and using both average and marginal water fees in estimating water demand functions. Results are consistent with the hypothesis that households may become less price-sensitive after exposure to drought. However, we find no systematic evidence of long-run, drought-related reductions in water demand, itself.

demand hardening↗

Data and scripts associated with a manuscript analyzing ELM-FATES parameter sensitivity under pre-fire and postfire scenarios using machine learning

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Fire Severity-Dependent Shifts in Vegetation Parameter Sensitivity: A Pre- and Post-Fire Analysis Using ELM-FATES and Explainable AI” submitted to Journal of Advances in Modeling Earth Systems (Zahura et al. 2026). The study examines vegetation physiological parameters controlling pre-fire and post-fire vegetation dynamics. To support this analysis, 73 vegetation parameters in Functionally Assembled Terrestrial Ecosystem Simulator (FATES) (Fisher et al., 2018) , which is coupled with E3SM (Energy Exascale Earth System Model) land model (ELM, ELM-FATES), were perturbed using a Sobol sequence to generate 1,024 ensemble members for two plant functional types: needleleaf evergreen extratropical trees (NEET) and C3 grass. Simulations were conducted for the pre-fire period (2016) and post-fire period (2018–2023). Burn severity was represented by modifying the Nesterov index in FATES to 75,000, 150,000, and 300,000 for low, moderate, and high severity, respectively. A no-fire scenario was also included. Simulations were performed for 16 grid cells in the American River Watershed across different burn severities and plant functional types. XGBoost (eXtreme Gradient Boosting) models were trained using the parameter ensembles and ELM-FATES-simulated outputs, including leaf area index (LAI), gross primary productivity (GPP), aboveground biomass, vegetation evaporation, transpiration, and soil evaporation. Models were trained separately for each year and burn severity, followed by SHAP (SHapley Additive exPlanations) analysis to identify changes in dominant parameters after fire disturbance. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package contains the ELM-FATES simulation data. The scripts and data related to the analysis will be added later. The inputs and outputs from ELM-FATES are inside the “FATES” folder. “FATES_domain_surface” contains the domain and surface netcdfs that were used to run ELM-FATES in the study area. “FATES_parameters” contains the 1024 ensembles that were generated using Sobol sequence. “FATES_outputs” folder contains ELM-FATES simulated variables. All files are .csv and .nc (NetCDF).

Aboveground biomass↗

Probing the limits of cosmological information from the Lyman- α forest 2-point correlation functions

The standard cosmological analysis with the Lyα forest relies on a continuum fitting procedure that suppresses information on large scales and distorts the three-dimensional correlation function on all scales. In this work, we present the first cosmological forecasts without continuum fitting distortion in the Lyα forest, focusing on the recovery of large-scale information. Using idealized synthetic data, we compare the constraining power of the full shape of the Lyα forest auto-correlation and its cross-correlation with quasars using the baseline continuum fitting analysis versus the true continuum. We find that knowledge of the true continuum enables a ∼ 10% reduction in uncertainties on the Alcock-Paczyński (AP) parameter and the matter density, Ω m . We also explore the impact of large-scale information by extending the analysis up to separations of 240 h -1 Mpc along and across the line of sight. The combination of these analysis choices can recover significant large-scale information, yielding up to a ∼ 15% improvement in AP constraints. This improvement is analogous to extending the Lyα forest survey area by ∼ 40%.

Lyman alpha forest↗

Transient Catalytic Reaction Analysis Through Signal Defragmentation

The Temporal Analysis of Products (TAP) pulse response technique provides valuable insights into catalytic function and reaction kinetics. However, complex fragmentation patterns in the TAP mass spectrometry signals can complicate precise quantification, particularly when analyzing transient gas flux data typical of TAP experiments. This work demonstrates a standard defragmentation method that deconvolves transient TAP signals while maintaining the temporal resolution of the experiment. First, the integrals of calibration gas fluxes are used to determine the fingerprint fragmentation pattern and construct a fragmentation matrix. This matrix is then used to defragment experimental flux data at each recorded time point via a non-negative least squares regression. The effectiveness of this method is demonstrated using virtual data and control experiments with a TAP reactor system. The defragmentation is then applied to the more complex propane dehydrogenation reaction on a chromia/alumina catalyst, which can contain up to ten significant gas species in the reactor outlet. Initial propane pulsing reveals an induction period during which propane is fully oxidized to CO2, followed by partial reduction to CO. Afterwards, there is a transition in chemistries towards coking and propylene production. Our example illustrates a practical method for the accurate determination of the time-dependent reactant/product concentrations and rates for a thorough analysis of the propane dehydrogenation kinetics. This approach can be broadly applied to any transient mass spectrometry experiment for a better understanding of catalyst-reaction dynamics.

36 - MATERIALS SCIENCE↗

Study of Nuclear Structure Functions at Jefferson Lab

This thesis presents a study of nuclear modification of quark distributions (the EMC effect) using inclusive electron-scattering data from Jefferson Lab Hall C, with emphasis on experiment E12-10-008 (XEM2). The analysis spans nuclei from light to heavy targets and combines measurements at multiple spectrometer settings to constrain EMC ratios over a broad range in Bjorken-x and Q2. A complete analysis framework was developed to extract charge-normalized and efficiency-corrected yields, including detector calibrations, beam-current calibration, density-loss corrections for cryogenic targets, background subtraction, radiative and Coulomb corrections, and systematic studies. Instrumental cross-checks, including HMS–SHMS comparisons and reconstruction studies, show that residual spectrometer differences are predominantly multiplicative in the kinematic region relevant to the EMC analysis. The extracted EMC ratios are broadly consistent with previous measurements while extending coverage across many nuclei. The EMC slope increases from light to heavy nuclei and shows signs of saturation at large A. After accounting for the dominant A dependence, the results are consistent with no isospin dependence, but also consistent with the predicted modification from the isovector model.

Sharda, Abhyuday [Univ. of Tennessee, Knoxville, T↗

Accurate and Data‐Efficient Micro X‐ray Diffraction Phase Identification Using Multitask Learning: Application to Hydrothermal Fluids

Traditional analysis of highly distorted micro X‐ray diffraction (μ‐XRD) patterns from hydrothermal fluid environments is a time‐consuming process, often requiring substantial data preprocessing and labeled experimental data. Herein, the potential of deep learning with a multitask learning (MTL) architecture to overcome these limitations is demonstrated. MTL models are trained to identify phase information in μ‐XRD patterns, minimizing the need for labeled experimental data and masking preprocessing steps. Notably, MTL models show superior accuracy compared to binary classification convolutional neural networks. Additionally, introducing a tailored cross‐entropy loss function improves MTL model performance. Most significantly, MTL models tuned to analyze raw and unmasked XRD patterns achieve close performance to models analyzing preprocessed data, with minimal accuracy differences. This work indicates that advanced deep learning architectures like MTL can automate arduous data handling tasks, streamline the analysis of distorted XRD patterns, and reduce the reliance on labor‐intensive experimental datasets.

97 MATHEMATICS AND COMPUTING↗