Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Data-Efficient Dimensionality Reduction and Surrogate Modeling of High-Dimensional Stress Fields

Tensor datatypes representing field variables like stress, displacement, velocity, etc., have increasingly become a common occurrence in data-driven modeling and analysis of simulations. Numerous methods [such as convolutional neural networks (CNNs)] exist to address the meta-modeling of field data from simulations. As the complexity of the simulation increases, so does the cost of acquisition, leading to limited data scenarios. Modeling of tensor datatypes under limited data scenarios remains a hindrance for engineering applications. Here, in this article, we introduce a direct image-to-image modeling framework of convolutional autoencoders enhanced by information bottleneck loss function to tackle the tensor data types with limited data. The information bottleneck method penalizes the nuisance information in the latent space while maximizing relevant information making it robust for limited data scenarios. The entire neural network framework is further combined with robust hyperparameter optimization. We perform numerical studies to compare the predictive performance of the proposed method with a dimensionality reduction-based surrogate modeling framework on a representative linear elastic ellipsoidal void problem with uniaxial loading. The data structure focuses on the low-data regime (fewer than 100 data points) and includes the parameterized geometry of the ellipsoidal void as the input and the predicted stress field as the output. The results of the numerical studies show that the information bottleneck approach yields improved overall accuracy and more precise prediction of the extremes of the stress field. Additionally, an in-depth analysis is carried out to elucidate the information compression behavior of the proposed framework.

artificial intelligence↗

DiffLense: a conditional diffusion model for super-resolution of gravitational lensing data

Abstract Gravitational lensing data is frequently collected at low resolution due to instrumental limitations and observing conditions. Machine learning-based super-resolution techniques offer a method to enhance the resolution of these images, enabling more precise measurements of lensing effects and a better understanding of the matter distribution in the lensing system. This enhancement can significantly improve our knowledge of the distribution of mass within the lensing galaxy and its environment, as well as the properties of the background source being lensed. Traditional super-resolution techniques typically learn a mapping function from lower-resolution to higher-resolution samples. However, these methods are often constrained by their dependence on optimizing a fixed distance function, which can result in the loss of intricate details crucial for astrophysical analysis. In this work, we introduce DiffLense , a novel super-resolution pipeline based on a conditional diffusion model specifically designed to enhance the resolution of gravitational lensing images obtained from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). Our approach adopts a generative model, leveraging the detailed structural information present in Hubble space telescope (HST) counterparts. The diffusion model, trained to generate HST data, is conditioned on HSC data pre-processed with denoising techniques and thresholding to significantly reduce noise and background interference. This process leads to a more distinct and less overlapping conditional distribution during the model’s training phase. We demonstrate that DiffLense outperforms existing state-of-the-art single-image super-resolution techniques, particularly in retaining the fine details necessary for astrophysical analyses.

Computer Science↗

Optimizing and Exploring Untapped Micro-Hydro Hybrid Systems: a Multi-Objective Approach for Crystal Lake as a Large-Scale Energy Storage Solution

Increasing electricity demand and concerns about climate change and fossil fuel consumption have highlighted the importance of renewable energy resources and storage systems. This paper proposes a method for exploring untapped pumped hydro storage potentials to accommodate intermittent renewable energy generation profiles. Hourly measured data from 2022 in Benzie County, Michigan, United States, were gathered for system sizing and a thorough, realistic analysis. By employing the multi-objective grey wolf optimization algorithm, we formulated optimal sizing and energy-management strategies for three different scenarios. Unlike similar studies, the 3rd with triple objective functions (OFs) scenario aims to maximize both reliability and ecological OFs while minimizing the cost OF. It has shown promising results with multiple solutions, considering economic, environmental, and reliability factors. A case study conducted in Crystal Lake, Michigan, revealed that although Crystal Lake would function only as a micro-hydro power facility, it is a promising and huge storage unit with a substantial storage capacity of around 14.9734GWh. The system investigated is significant in the USA due to its rapid deployment capabilities, minimal construction requirements, and ease of integration with the distribution grid. The fuzzy logic method was employed to identify the best non-dominant solution among the other solutions. Furthermore, these outcomes include a notably low levelized cost of energy at 0.046147$/kWh, a robust index of reliability of 99.705%, and a significant reduction in CO₂ emissions amounting to 7.9142×10 3 tons/year, when considering the triple OFs. The paper’s methodology provides valuable insights for regions aiming to utilize renewable energy from untapped storage sources.

13 HYDRO ENERGY↗

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie↗

Phosphoproteomics Modifications in Women with Rheumatoid Arthritis─Application of Web-Based Software to Enhance Data Visualization

Individuals with rheumatoid arthritis (RA) are at increased risk of functional disability, cardiovascular disease, and obesity, all of which are influenced by dysregulated skeletal muscle. Here, this pilot study aims to identify phosphoproteomics changes in RA skeletal muscle and visualize modifications through development of a web-based app designed to promote user-friendly data interpretation and visualization. NanoLC–MS/MS analysis was performed on vastus lateralis biopsies from three women with RA and matched healthy controls. Differential analysis was performed using the Limma R package. Kinase substrate enrichment analysis (KSEA) predicted changes in kinase activity. RA muscle displayed 35 upregulated and 60 downregulated phosphosites, including the cytoskeletal proteins TTN (Ser33201, Ser33013, Ser20925), NEB (Ser2219, Thr254, Ser33013, Ser20925), FLNA (Ser1459), and LASP1 (Ser146). Compared to healthy controls, KSEA predicted decreased activity of several kinases in RA muscle, including PRKACA and CDKs. All such changes were visualized by use of our web-based app. Overall, phosphoproteome analysis reveals signaling alterations in RA skeletal muscle linked to cytoskeletal proteins, representing candidate disease biomarkers; these modifications can be explored through use of our web-based software.

phosphoproteomics↗

Single-channel and single-energy partial-wave analysis with continuity improved through minimal phase constraints

Single-energy partial-wave analysis has often been applied as a way to fit data with minimal model dependence. However, remaining unconstrained, partial waves at neighboring energies will vary discontinuously because the overall amplitude phase cannot be determined through single-channel measurements. This problem can be mitigated through the use of a constraining penalty function based on an associated energy-dependent fit. However, the weight given to this constraint results in a biased fit to the data. In this paper, for the first time, we explore a constraining function which does not influence the fit to data. The constraint comes from the overall phase found in multichannel fits which, in the present study, are the Bonn-Gatchina and Jülich-Bonn multichannel analyses. The data are well reproduced and weighting of the penalty function does not influence the result. The method is applied to K⁢Λ photoproduction data and all observables can be maximally well reproduced. While the employed multichannel analyses display very different multipole amplitudes, we show that the major difference between two sets of multipoles can be related to the different overall phases.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Rancor Integrated Procedure System (RIPS): A Computer-Based Procedure Platform for Advanced Reactor Research

The Rancor Microworld Simulator is a simplified, pressurized water, small modular reactor simulator that includes a multi-unit plant model server, an advanced digital human-machine control interface, and the Rancor Integrated Procedure System (RIPS). Rancor provides a research and development tool that can be used for collecting operator performance data and for prototyping concepts of operations (ConOps) for advanced reactor development. RIPS is meant as a research tool and includes many unique features: (1) RIPS has a robust procedure authoring system. (2) RIPS has the capability to run any of the three IEEE-Std-1786 computer-based procedure types. (3) RIPS can be configured to take on the look and feel of different vendors’ computer-based procedure systems for the purpose of developing and evaluating different ConOps for plant upgrades or new builds. (4) RIPS includes the capability for logging operator procedure use, including integrating procedure logs with Rancor simulator logs, thereby allowing automated data collection of operator scenario runs. (5) RIPS integrates with the Human Unimodel for Nuclear Technology to Enhance Reliability (HUNTER), a dynamic human reliability analysis environment that creates a digital human twin or virtual operator to mimic reactor operator performance. (6) RIPS includes support for automation of plant monitoring and control functions. While RIPS is explicitly built into Rancor, it may also be used with full-scope training simulators. This functionality allows RIPS to be used for existing plants and advanced reactors under development.

99 - GENERAL AND MISCELLANEOUS↗

Design Optimization of a Criticality Experiment for the Molten Chloride Reactor Experiment Facility

Neutronics simulations of Molten Chloride Fast Reactors have quantifiable biases that arise from nuclear data, modeling choices, or numerical methods. The multiphysics nature of molten salt reactors makes it challenging to disentangle neutronics modeling biases from biases originating from other physical phenomena. In comparison to a mock-up reactor, criticality experiments can specifically assess the neutronics modeling bias while limiting multiphysics effects. The criticality experiment must be neutronically representative of the full-scale reactor to be valuable. Here, in this paper, we describe the design of a criticality experiment to validate only the neutronics of TerraPower’s Molten Chloride Reactor Experiment (MCRE) and its criticality safety upset scenarios. The proposed experiment uses different chlorine-containing materials to maximize its similarity to the MCRE. The design process uses a constrained Bayesian optimization algorithm to investigate different objective functions that use covariance information for 35 Cl nuclear data. The experiments could reduce the nuclear data–induced uncertainty in k eff of the MCRE from 2161 to 886 pcm. They would also increase the upper subcritical limit of the MCRE criticality safety upset scenario from 0.94101 to 0.94476 when using the WHISPER analysis framework.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Using Machine Learning to Understand Electric and Hybrid Vehicles Ownership in Burdened and Nonburdened Communities

Transitioning to electric and hybrid vehicles (EHVs) for all communities is a pivotal step toward sustainable transportation and environmental conservation. This paper aims to understand the adoption of EHVs, focusing on burdened communities (BCs) in the United States. The EHV ownership-based analysis combines two datasets—behavioral data from the Puget Sound Regional Travel Survey integrated with BCs (Justice40) data covering transportation insecurity, environmental burden, social vulnerability, health vulnerability, and climate and disaster risk burden. After creating this unique database, descriptive analysis and modeling are used to analyze the data and predict EHV ownership in the future. Specifically, we use a new method that combines particle swarm optimization (PSO) with a stacking model named PSO-Stacking, which incorporates heterogeneous base learners of machine learning and deep learning. PSO applies a customized objective function to select the optimal hyperparameters for heterogeneous learners within the stacking model, effectively addressing challenges such as multicollinearity, data imbalance, nonlinearity, and overfitting. The proposed solution covers more accurate results than standard benchmark models for EHV ownership in BCs and non-BCs. In addition, the results of the PSO-Stacking method are explained using the local interpretable model-agnostic explanations technique. Results show a negative correlation between the BCs indicators, that is, higher transportation insecurity associated with lower EHV ownership. Furthermore, BCs have higher future climate risk scores, diesel particulate matter levels, and PM2.5 in the air than non-BCs because of higher conventional vehicle ownership. These communities are at higher risk and can benefit from electrification, EV infrastructure, and EV policies to address environmental challenges.

Aslam, Zeeshan [ORNL]↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

The DESI DR1 peculiar velocity survey: growth rate measurements from the maximum likelihood fields method

We present the constraint on the growth rate of structure from the combination of DESI DR1 BGS sample, Fundamental Plane, and Tully-Fisher peculiar velocity catalogues using the maximum likelihood fields method. The combined catalogue contains 415,523 galaxy redshifts and 76,616 peculiar velocity measurements. To handle the large amount of data in the DESI DR1 peculiar velocity catalogue, we significantly improve the computational efficiency by rewriting the algorithm with JAX. After removing outliers and Tully-Fisher galaxies that are affected by systematics, we find fσ 8 = 0.483 -0.043 +0.080 (stat) ± 0.018(sys), consistent within 1σ with the power spectrum and correlation function analysis using the same dataset. Combining all three measurements with appropriate correlations, the consensus measurement is fσ 8 (z eff = 0.07) = 0.450±0.055, consistent with Planck +ΛCDM cosmology (fσ 8 = 0.449±0.008). Combining with the high redshift growth rate of structure measurements from DESI ShapeFit, the constraint on the growth index is γ = 0.58±0.11, consistent with GR.

cosmic flows↗

Deep-learning methods for contrast enhancement and artifact reduction in cryo-electron tomography: a systematic analysis of the state of the art and proposed improvements

Cryo-electron tomography (cryo-ET) has emerged as the preferred technique for visualizing the organization of macromolecular complexes in situ and resolving their structures at subnanometre resolution [Tegunov et al. (2021)View full citation, Nat. Methods, 18, 186–193]. Despite improvements in data quality as a result of advances in detector technology, microscope stability and stage precision, the analysis and interpretation of tomograms remains challenging due to a low signal-to-noise ratio and reconstruction artifacts stemming from experimental constraints in specimen tilt during data collection resulting in a missing wedge in the Fourier space. Recently, self-supervised deep-learning methods have been proposed for contrast enhancement and reduction of resolution anisotropy in reconstructed tomograms. Here, we evaluate several state-of-the-art deep-learning methods which aim to improve the interpretability of cryo-ET reconstructions, with a focus on their performance on downstream tasks of template matching, sub­tomogram averaging and segmentation. We propose new training architectures and a loss function based on Fourier shell correlation that show improved performance over the standard U-Net with L1/L2 losses. We demonstrate our analysis on four diverse experimental datasets: purified 80S ribosomes, in situ Chlamydomonas reinhardtii, immature HIV-1 virus-like particles and INS-1E cells.

contrast enhancement↗

Achieving Multimodal and Multicolor Luminescence in LaAlO 3 :Pr 3+ , Gd 3+ via Trap Engineering and Energy Transfer

Achieving multimodal luminescence within a single phosphor is vital for multifunctional applications but remains challenging due to complex color tuning and trap engineering. In this study, we report Pr 3+ and Gd 3+ co‐doped LaAlO 3 (LAO:PG) phosphors, designed through careful modulation of multilevel traps and Pr 3+ → Gd 3+ energy transfer dynamics. These materials exhibit diverse luminescence modes, including down‐conversion luminescence (DCL), up‐conversion luminescence (UCL), persistent luminescence (PersL), optically stimulated luminescence (OSL), and thermally stimulated luminescence (TSL) across a wide spectral range. Unlike previously studied Pr 3+ ‐doped LAO, the co‐doped LAO:PG shows DCL in both UV‐visible and NIR regions and displays ultraviolet‐C UCL under visible excitation. Notably, we observe, for the first time, PersL lasting several minutes in these phosphors—an improvement over the non‐PersL behavior of Pr 3+ ‐only doped LAO. Additionally, the LAO:PG phosphors exhibit strong OSL response. TSL analysis reveals five distinct trap levels linked to these properties. Density functional theory calculations further correlate intrinsic defects to these traps, supporting a proposed mechanism for the observed multimodal luminescence. These findings highlight LAO:PG as a promising platform for developing advanced phosphors with integrated luminescence modes, paving the way for future applications in data storage, phototherapy, and anti‐counterfeiting technologies.

Chemistry↗

Characterizing Defect Dynamics in Silicon Carbide Using Symmetry-Adapted Collective Variables and Machine Learning Interatomic Potentials

Silicon carbide (SiC) divacancies are attractive candidates for spin-defect qubits possessing long coherence times and optical addressability. The high activation barriers associated with SiC defect formation and motion pose challenges for their study by first-principles molecular dynamics. In this work, we develop and deploy machine learning interatomic potentials (MLIPs) to accelerate defect dynamics simulations while retaining ab initio accuracy. We employ an active learning strategy comprising symmetry-adapted collective variable discovery and enhanced sampling to compile configurationally diverse training data, calculation of energies and forces using density functional theory (DFT), and training of an E(3)-equivariant MLIP based on the Allegro model. Here, the trained MLIP reproduces DFT-level accuracy in defect transition activation free energy barriers, enables the efficient and stable simulation of multidefect 216-atom supercells, and permits an analysis of the temperature dependence of defect thermodynamic stability and formation/annihilation kinetics to propose an optimal annealing temperature to maximally stabilize VV divacancies.

Computer simulations↗

Examination of Replicate Syntheses of Metal Organic Frameworks as a Window into Reproducibility in Materials Chemistry

Replicate experiments are a useful tool in understanding the repeatability of scientific measurements. In 2019, a systematic search for replicate syntheses of a collection of 130 metal–organic frameworks (MOFs) found that 89% of these materials had no reported replicate syntheses apart from the original publications identifying the material (Agrawal, M. Proc. Natl. Acad. Sci. U.S.A. 2020, 117, 877−88210.1073/pnas.1918484117). A potential weakness of that search was that only 5–11 years had elapsed since the original publication of each material. Here, this analysis is extended to all publications 11–17 years after the original publication. Although this extended time period identifies more repeat syntheses, 83% of the materials still have no reported replicate syntheses. We also consider how appropriately selected Density Functional Theory (DFT) calculations can provide corroboration for the experimentally reported crystal structures. By using data from previous high-throughput DFT studies, corroborating evidence from DFT was available for 17% of the 130 structures for which no replicate syntheses are available. In total, approximately 1/3 of the 130 MOFs have data associated with replicate synthesis experiments and/or directly corroborating DFT calculations.

Sholl, David S. [Oak Ridge National Laboratory (OR↗

Wavelet flow for extragalactic foreground simulations

Extragalactic foregrounds in cosmic microwave background (CMB) observations are both a source of cosmological and astrophysical information and a nuisance to the CMB. Effective field-level modeling that captures their non-Gaussian statistical distributions is increasingly important for optimal information extraction, particularly given the low-noise observations from current and upcoming experiments. Here, we explore the use of Wavelet Flow (WF) models to tackle the novel task of modeling the field-level probability distributions of multi-component CMB secondaries and foregrounds. Specifically, we jointly train correlated CMB lensing convergence (κ) and cosmic infrared background (CIB) maps with a WF model and obtain a network that statistically recovers the input to high accuracy — the trained network generates samples of κ and CIB fields whose average power spectra are within a few percent of the inputs across all scales, and whose Minkowski functionals are similarly accurate compared to the inputs. Leveraging the multiscale architecture of these models, we fine-tune both the model parameters and the priors at each scale independently, optimizing performance across different resolutions. These results demonstrate that WF models can accurately simulate correlated components of CMB secondaries, supporting improved analysis of cosmological data. Our code and trained models can be found on this GitHub repo.

cosmological simulations↗

X-ray scattering based scanning tomography for imaging and structural characterization of cellulose in plants

X-ray and neutron scattering have long been used for structural characterization of cellulose in plants. Due to averaging over the illuminated sample volume, these measurements traditionally overlooked the compositional and morphological heterogeneity within the sample. Here, a scanning tomographic imaging method is described, using contrast derived from the X-ray scattering intensity, for virtually sectioning the sample to reveal its internal structure at a resolution of a few micrometres. This method provides a means for retrieving the local scattering signal that corresponds to any voxel within the virtual section, enabling characterization of the local structure using traditional data-analysis methods. This is accomplished through tomographic reconstruction of the spatial distribution of a handful of mathematical components identified by non-negative matrix factorization from the large dataset of X-ray scattering intensity. Joint analysis of multiple datasets, to find similarity between voxels by clustering of the decomposed data, could help elucidate systematic differences between samples, such as those expected from genetic modifications, chemical treatments or fungal decay. The spatial distribution of the microfibril angle can also be analyzed, based on the tomographically reconstructed scattering intensity as a function of the azimuthal angle.

36 MATERIALS SCIENCE↗

An efficient hybrid downscaling framework to estimate high-resolution river hydrodynamics

Flow depth and velocity are the most important hydrodynamic variables that govern various river functions, including water resources, navigation, sediment transport, and biogeochemical cycling. Existing high-resolution flow depth simulations rely on either computationally expensive river hydrodynamic models (RHMs) or data-driven models with formidable training costs, whereas data-driven modeling of flow velocity has rarely been explored. Here, using the hybrid Low-fidelity, Spatial analysis, and Gaussian process learning (LSG) model, we developed a downscaling approach to construct high-resolution flow depth and velocity from a two-dimensional (2-D) RHM simulation at coarse resolution. The LSG models were trained and tested in an urban watershed in Houston using two different hurricane-driven flood events. The high-resolution (as fine as 30 m resolution) and low-resolution (mostly 1000 m resolution) meshes include 664 724 and 14 536 grid cells, respectively. The results showed that through downscaling, the simulation errors were reduced to less than one-fourth and one-third of the errors of the low-resolution 2-D RHM for flow depth and velocity, respectively. Our analysis further revealed that the dominant uncertainty sources of the downscaled hydrodynamics are different, with flow velocity dominated by the dimensionality reduction error, which we reduced by using a regionalized training procedure. The downscaling approach achieves an 84-fold acceleration in computational time compared to the high-resolution 2-D RHM, making high-fidelity ensemble flood modeling feasible. More importantly, the developed method provides an opportunity to couple large-scale hydrodynamical processes with local physical, chemical, and biological processes in river models.

Tan, Zeli [Pacific Northwest National Laboratory (↗