Search NASA⌕ Search

SEARCH · Search NASA

Results for “latent variable”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Variational Autoencoder Model Toward Molecular Structure Representation Learning of Fuels

Here, in this work, a Variational Autoencoder (VAE)-based data-driven modeling framework is developed with the overarching goal of enabling fuel design. The VAE model is trained on a large dataset with several chemical species to learn a compressed latent space molecular representation. Chemical structure in the form of Simplified Molecular Input Line Entry System (SMILES) string is fed as input, encoded into the VAE latent space, and decoded back to the SMILES string using Long Short-Term Memory (LSTM) networks. Complexities of the VAE training loss function are thoroughly examined by varying the weightage (beta (𝜷) parameter) of the latent space regularization term, thereby assessing the balance between reconstruction accuracy and validity, and focusing on both accurate molecular structure reconstruction and latent space consistency. Two different strategies for 𝜷 variation are evaluated: linear annealing and cyclic annealing. In addition, the impact of total correlation adjustment and hierarchical priors is also studied with regard to the balance between reconstruction fidelity and latent space regularization, and potential issues such as posterior collapse, over-regularization, and poor disentanglement of latent variables. Overall, the best performance of the model is achieved with hierarchical priors and incrementally increasing 𝜷 from 0 to a threshold value of 0.25 over 75 epochs. The generative VAE model can be readily coupled with Quantitative Structure–Property Relationship (QSPR) analysis to develop an integrated end-to-end framework for fuel-property prediction and molecular design of novel promising fuels.

fuel design↗

Data for Influence of Particle Size on NIR Spectroscopic Characterization of Sorghum Biomass for the Biofuel Industry

NIR spectroscopy is a rapid and accurate green technology for high-throughput biomass characterization, including sorghum ( Sorghum bicolor ), a promising energy crop for the biofuel industry. This study assessed the influence of particle size on NIR spectroscopic analysis (wavelength range: 867–2535 nm) of sorghum biomass composition. Grown under field conditions, a total of 113 types of genetically diverse sorghum accessions were dried, ground, and sieved (<250, 250–600, 600–850, and > 850 µm particle size) for developing partial least square regression (PLSR) prediction models for moisture, ash, extractive, glucan, xylan, acid-soluble lignin (ASL), acid-insoluble lignin (AIL), and total lignin (ASL + AIL). Overall, smaller particle sizes provided better model performance, while no single particle size provided the best performance for all the selected components. With only 9 selected bands and 4 latent variables (LVs), the best PLSR model was obtained for moisture with particle size of 600–850 µm with the square root of the coefficient of determination (R) of 0.85, the ratio of prediction to deviation (RPD) of 2.2, and the root mean square error (RMSE) of 0.46 % in external validation. Similar model performances were also obtained for ash, extractive, glucan, and xylan. This study showed that size reduction could effectively improve NIR spectroscopic analysis for lipid-producing sorghum biomass for the biofuel industry.

Biomass Analytics↗

Efficient Dimension Reduction of Complex Three-dimensional CO2 Saturation using Deep Learning Models

In the domain of deep learning (DL), dimension reduction is crucial for enhancing training efficiency and mitigating overfitting, particularly when managing complex data such as three-dimensional (3D) saturation data. The 3D saturation data in the context of geological carbon storage (GCS) presents unique challenges due to its inherent sparsity and the abrupt transitions at plume boundaries, known as shock fronts. To address the challenges, we proposed a novel DL framework that integrates dimension reduction with advanced 3D reconstruction techniques. Our model leveraged latent variables derived from 2D average saturation data, offering a robust and efficient solution tailored to the intricate dynamics of 3D saturation fields. The proposed framework can extract the critical features of the high-dimensional data while reducing the variable numbers, which is more tractable for DL models and enhances the model robustness and accuracy. Therefore, it provides a novel approach for modeling and analyses in complex geological scenarios, which finds great potential applications in environmental monitoring and energy storage.

Wang, Hongsheng↗

Data-Driven Supervised Dimension Reduction for Scientific Discovery (LDRD QTI Report)

This report summarizes the findings of a four months FY24 Advanced Science & Technology (AS&T) LDRD Quick Targeted Investigation (QTI) project focused on the exploration of supervised dimension reduction approaches based on autoencoders. Autoencoders have been extensively employed in literature for unsupervised learning tasks, however, their use for supervised regression tasks, which are common within scientific applications, has been limited. Motivated by linear dimension reduction strategies like Active Subspaces and Adaptive Basis, we explored the possibility of employing autoencoders to discover a non-linear manifold able to represent the original function in fewer dimensions. In this report, we discuss a neural network architecture and we perform a numerical campaign on several problems ranging from simple two-dimensional functions to a model problem for magnetohydrodynamics in five dimensions. In our preliminary results, we show that the proposed approach is found to be superior to linear dimension reduction strategies in representing the target function even with a single latent variable.

97 MATHEMATICS AND COMPUTING↗

Deep Learning-based Parameterization of Complex 3D CO2 Saturation Data in Large-scale Geological Carbon Storage

In deep learning (DL), dimension reduction plays a pivotal role in improving training efficiency and minimizing overfitting, especially when working with complex datasets like three-dimensional (3D) saturation data. In the context of geological carbon storage (GCS), 3D saturation data introduces unique challenges due to its sparse nature and sharp transitions at plume boundaries, known as shock fronts. To tackle these challenges, we developed a novel DL framework that combines dimension reduction with advanced 3D reconstruction techniques. Our approach utilizes latent variables derived from 2D average saturation fields to efficiently capture the essential features of high-dimensional data while reducing the number of variables. This enhances both the robustness and accuracy of DL models, making the framework more practical for real-world applications. By offering a tailored solution for modeling complex 3D saturation dynamics, this framework holds significant potential for environmental monitoring, energy storage, and other geological applications.

Wang, Hongsheng [University of Texas at Austin]↗

Uncertainty in Tropical Ocean Latent Heat Flux Variability During the Last 25 Years

When averaged over the tropical oceans (30deg N/S), latent heat flux anomalies derived from passive microwave satellite measurements as well as reanalyses and climate models driven with specified seal-surface temperatures show considerable disagreement in their decadal trends. These estimates range from virtually no trend to values over 8.4 W/sq m decade. Satellite estimates also tend to have a larger interannual signal related to El Nino/Southern Oscillation (ENSO) events than do reanalyses or model simulations. An analysis of wind speed and humidity going into bulk aerodynamic calculations used to derive these fluxes reveals several error sources. Among these are apparent remaining intercalibration issues affecting passive microwave satellite 10 m wind speeds and systematic biases in retrieval of near-surface humidity. Likewise, reanalyses suffer from discontinuities in availability of assimilated data that affect near surface meteorological variables. The results strongly suggest that current latent heat flux trends are overestimated.

Robertson, F. R.↗

Moisture and latent heat flux variabilities in the tropical Pacific derived from satellite data

This paper describes a method of determining latent heat flux and the ocean-atmosphere moisture from sea surface temperature, precipitable water, and surface wind speed data derived from 1980-1983 observations of SMMR aboard Nimbus 7 above tropical Pacific. The observation period included a very intense El Nino-Southern Oscillation (ENSO) episode. It was found that, during the early phase of the 1982-1983 ENSO, a surface convergence center moved east leading the anomalous equatorial westerlies. At this center, the low wind and high humidity caused negative (low) latent heat flux anomalies, despite anomalously high sea surface temperatures. Latent heat flux was found to play an important role in the seasonal cooling of the upper ocean, except in areas covered by major surface convergence zones and in areas of ocean upwelling.

Liu, W. Timothy↗

Interannual and Decadal Variability of Ocean Surface Latent Heat Flux as Seen from Passive Microwave Satellite Algorithms

Ocean surface turbulent fluxes are critical links in the climate system since they mediate energy exchange between the two fluid systems (ocean and atmosphere) whose combined heat transport determines the basic character of Earth's climate. Deriving physically-based latent and sensible heat fluxes from satellite is dependent on inferences of near surface moisture and temperature from coarser layer retrievals or satellite radiances. Uncertainties in these "retrievals" propagate through bulk aerodynamic algorithms, interacting as well with error properties of surface wind speed, also provided by satellite. By systematically evaluating an array of passive microwave satellite algorithms, the SEAFLUX project is providing improved understanding of these errors and finding pathways for reducing or eliminating them. In this study we focus on evaluating the interannual variability of several passive microwave-based estimates of latent heat flux starting from monthly mean gridded data. The algorithms considered range from those based essentially on SSM/I (e.g. HOAPS) to newer approaches that consider additional moisture information from SSM/T-2 or AMSU-B and lower tropospheric temperature data from AMSU-A. On interannual scales, variability arising from ENSO events and time-lagged responses of ocean turbulent and radiative fluxes in other ocean basins (as well as the extratropical Pacific) is widely recognized, but still not well quantified. Locally, these flux anomalies are of order 10-20 W/sq m and present a relevant "target" with which to verify algorithm performance in a climate context. On decadal time scales there is some evidence from reanalyses and remotely-sensed fluxes alike that tropical ocean-averaged latent heat fluxes have increased 5-10 W/sq m since the early 1990s. However, significant uncertainty surrounds this estimate. Our work addresses the origin of these uncertainties and provides statistics on time series of tropical ocean averages, regional space / time correlation analysis, and separation of contributions by variations in wind and near surface humidity deficit. Comparison to variations in reanalysis data sets is also provided for reference.

Robertson, Franklin R.↗

Month-to-month variability of ocean-atmosphere latent heat flux as observed from the Nimbus microwave radiometer

Comparison with in situ measurements shows that the Nimbus/Scanning Multichannel Microwave Radiometer is useful in describing the month-to-month variability of the latent heat flux 'LE' and related parameters during the 1982-1983 El Nino event. The spaceborne measured monthly mean LE was found to be within 30 W/sq m of those derived from ship reports. Absolute accuracy could not be determined, though satellite measurements could extrapolate information on the LE both in space and in time.

Liu, W. Timothy↗

A statistical model for interpreting computerized dynamic posturography data

Computerized dynamic posturography (CDP) is widely used for assessment of altered balance control. CDP trials are quantified using the equilibrium score (ES), which ranges from zero to 100, as a decreasing function of peak sway angle. The problem of how best to model and analyze ESs from a controlled study is considered. The ES often exhibits a skewed distribution in repeated trials, which can lead to incorrect inference when applying standard regression or analysis of variance models. Furthermore, CDP trials are terminated when a patient loses balance. In these situations, the ES is not observable, but is assigned the lowest possible score--zero. As a result, the response variable has a mixed discrete-continuous distribution, further compromising inference obtained by standard statistical methods. Here, we develop alternative methodology for analyzing ESs under a stochastic model extending the ES to a continuous latent random variable that always exists, but is unobserved in the event of a fall. Loss of balance occurs conditionally, with probability depending on the realized latent ES. After fitting the model by a form of quasi-maximum-likelihood, one may perform statistical inference to assess the effects of explanatory variables. An example is provided, using data from the NIH/NIA Baltimore Longitudinal Study on Aging.

NASA Discipline Neuroscience↗

W2VPCA: A Machine Learning Method for Measuring Attitudes With Natural Language

Company strategy influences many decisions in freight transportation. Behavioral models of company decision-making therefore could benefit from including strategy variables. However, strategy is difficult to observe and quantify. Attitudinal surveys of company executives can be used to collect measurements of latent strategy to use in quantitative models. However, surveys are costly and burdensome. Text mining methods to collect measurements overcome these issues somewhat, but typically require manual intervention and ignore the context of words, which can be problematic. This study introduces a new machine learning method to generate strategy measurement data from existing big text data. The new method, called W2VPCA, combines Natural Language Processing and Principal Components Analysis. W2VPCA produces measurement data that serve as quantitative indicators of latent strategy in behavioral models. W2VPCA is unsupervised, data-driven, and uses information on word context. We apply W2VPCA to generate measurements of latent strategies using readily available, large-scale text data: annual company reports. The empirical measurements are used successfully to associate two latent strategies, one focusing on distribution and the other on products, with truck fleet and distribution center outsourcing decisions. The main empirical outcome is that the W2VPCA measurements outperform Bag-of-Words measurements in a psychometric analysis of latent firm strategies. While this study focuses on freight behavioral models, W2VPCA may also have applications in behavioral modeling in other domains.

97 MATHEMATICS AND COMPUTING↗

Joint Modeling of Quasar Variability and Accretion Disk Reprocessing Using Latent Stochastic Differential Equations

Quasars are bright active galactic nuclei powered by the accretion of matter around supermassive black holes at the center of galaxies. Their stochastic brightness variability depends on the physical properties of the accretion disk and black hole. The upcoming Rubin Observatory Legacy Survey of Space and Time (LSST) is expected to observe tens of millions of quasars, so there is a need for efficient techniques like machine learning that can handle the large volume of data. Quasar variability is believed to be driven by an X-ray corona, which is reprocessed by the accretion disk and emitted as UV/optical variability. We are the first to introduce an auto-differentiable simulation of the accretion disk and reprocessing. We use the simulation as a direct component of our neural network to jointly model the driving variability and reprocessing, trained with supervised learning on simulated LSST-like 10 yr quasar light curves. We encode the light curves using a transformer encoder, and the driving variability is reconstructed using latent stochastic differential equations, a physically motivated generative deep learning method that can model continuous-time stochastic dynamics. By embedding the physical processes of the driving signal and reprocessing into our network, we achieve a model that is more robust and interpretable. We demonstrate that our model outperforms a Gaussian process regression baseline and can infer accretion disk parameters and time delays between wave bands, even for out-of-distribution driving signals. Our approach provides a powerful framework that can be adapted to solve other inverse problems in multivariate time series.

Fagin, Joshua [City Univ. of New York (CUNY), NY (↗

Bayesian reduced-order deep learning surrogate model for dynamic systems described by partial differential equations

We propose a reduced-order deep-learning surrogate model for dynamic systems described by time-dependent partial differential equations. This method employs space–time Karhunen–Loève expansions (KLEs) of the state variables and space-dependent KLEs of space-varying parameters to identify the reduced (latent) dimensions. Subsequently, a deep neural network (DNN) is used to map the parameter latent space to the state variable latent space. An approximate Bayesian method is developed for uncertainty quantification (UQ) in the proposed KL-DNN surrogate model. The KL-DNN method is tested for the linear advection–diffusion and nonlinear diffusion equations, and the Bayesian approach for UQ is compared with the deep ensembling (DE) approach, commonly used for quantifying uncertainty in DNN models. It was found that the approximate Bayesian method provides a more informative distribution of the PDE solutions in terms of the coverage of the reference PDE solutions (the percentage of nodes where the reference solution is within the confidence interval predicted by the UQ methods) and log predictive probability. The DE method is found to underestimate uncertainty and introduce bias. For the nonlinear diffusion equation, we compare the KL-DNN method with the Fourier Neural Operator (FNO) method and find that KL-DNN is 10% more accurate and needs less training time than the FNO method.

97 MATHEMATICS AND COMPUTING↗

Automating ridehailing services would reduce pooling, especially among women

Here, this study investigates how autonomous vehicles (AVs) could transform pooled (shared) ridehailing services, focusing on the impacts of fare reductions, the absence of drivers/staff, and psychological attributes such as trust in other passengers and privacy concerns. We distinguish between the automation of driving tasks and the removal of human driver/staff from the vehicle, providing novel insights into the factors influencing AV ridehailing adoption. Using a national survey with stated preference (SP) choice experiments and psychometric questions, we analyze the complex interactions of ridehailing fare, pooled ridehailing service quality, and latent attitudes on ridehailing choices. Our findings suggest that the elimination of drivers/staff from fully autonomous ridehailing could lead to a shift from pooled to solo rides, particularly among female travelers who may have greater concerns about trust and safety in unstaffed AVs. This study highlights the importance of addressing trust and comfort beyond fare discounts to ensure the inclusivity and widespread adoption of pooled AV ridehailing. These insights underscore the need for ridehailing providers and policymakers to prioritize trust-building measures, user-centered AV design that offers greater privacy, and dynamic pricing strategies, to ensure inclusive and widespread adoption of pooled AV services.

Autonomous vehicle↗

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING↗

BLOOD-BASED MULTI-SCALE MODEL FOR CANCER RISK FROM GCR IN GENETICALLY DIVERSE POPULATIONS

OBJECTIVES AND METHODS This project addresses the challenge of understanding and predicting individual radiation sensitivity by integrating genetics, demographics and biomarker characteristics across species (mice and humans). We hypothesize that ex vivo DNA repair response to GCR components is a central determinant of cancer risk from space radiation and can serve as a biomarker of radiation risk in combination with genetics. Automated image quantification of 53BP1+ radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET in non-immortalized primary skin fibroblasts derived from 76 mice across 15 strains (5 inbred reference strains and 10 collaborative-cross strains) exposed to X rays (0.1, 1 and 4 Gy), 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100μm2), as well as in peripheral blood mononuclear cells (PBMCs) from 768 healthy donors (matched ethnicity, 50/50 male/female, 18-70 years old) exposed to gamma rays (0.1 and 1 Gy), 350 MeV/n 28Si, 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100μm2). QUANTIFICATION OF 53BP1+ FOCI IN VITRO AND ASSOCIATIONS TO IN VIVO RADIATION SUSCEPTIBILITY IN 15 MOUSE STRAINS We reported in vitro repair kinetic and repairable fractions of RIF for the 15 mouse strains and introduced a mathematical model for RIF as a function of time, dose and LET. We noted that the metabolic activity of cells modulates the RIF response, and we introduced the open access tool terRIFic (Tool for Enhanced Results of RIF In Cells, https://radbiolab.shinyapps.io/terrific/) to correct for such bias using confluence level. Notably, at 4h post-irradiation, RIF/Gy decreased with dose or LET: as the dose or LET increases, so does the proximity of DNA double-strand-breaks (DSB) and our data suggest that proximal DSBs are brought together inside isolated RIF for repair. The RIF/Gy trend was inverted at 24h, suggesting RIF with high DSB content are more difficult to repair. We showed that in vitro metrics correlate with in vivo measurements in the same 15 mouse strains, such as survival levels of immune cells or spontaneous cancer incidence, suggesting a relationship between the efficiency of DSB repair and cancer risk or radiation toxicity. In addition to the efficiency of repair and persistent RIF, the amount of spontaneous foci before irradiation was also found to be strain dependent. Finally, we performed genome-wide association study in the same 15 mouse strains using all RIF phenotypes measured in vitro, identifying genes of interest and validating RIF as an ideal biomarker for individual radiation sensitivity. BASELINE 53BP1+ FOCI PREDICTS INDIVIDUAL HUMAN RESPONSE TO GCR COMPONENTS Based on the analysis of radiation responses of 576 donor PBMCs (using quantification of 53BP1+ foci, oxidative stress and cell death), we observed a wide variability of subject- and LET-dependent radiation responses, with radiation-induced DNA repair foci increasing with LET, though oxidative stress being notably reduced by high-LET irradiation, potentially due to a switch between hydrogen peroxide and oxygen radical-based mechanisms. We identified a relationship between few spontaneous DNA foci at baseline and increased DNA repair after irradiation, accompanied by an alteration in immunoregulatory cytokine secretion, which might be adapted as biomarkers to predict ionizing radiation sensitivity. Among demographic variables, only latent cytomegalovirus infection and age were predictive of high baseline foci formation. Finally, we have performed low-throughput whole genome sequencing of all samples and are currently in the process of identifying the genes and pathways associated with low and high-LET ionizing radiation sensitivity in humans.

53BP1↗

Latent Stochastic Differential Equations for Modeling Quasar Variability and Inferring Black Hole Properties

Quasars are bright and unobscured active galactic nuclei (AGN) thought to be powered by the accretion of matter around supermassive black holes at the centers of galaxies. The temporal variability of a quasar’s brightness contains valuable information about its physical properties. The UV/optical variability is thought to be a stochastic process, often represented as a damped random walk described by a stochastic differential equation (SDE). Upcoming wide-field telescopes such as the Rubin Observatory Legacy Survey of Space and Time (LSST) are expected to observe tens of millions of AGN in multiple filters over a ten year period, so there is a need for efficient and automated modeling techniques that can handle the large volume of data. Latent SDEs are machine learning models well suited for modeling quasar variability, as they can explicitly capture the underlying stochastic dynamics. In this work, we adapt latent SDEs to jointly reconstruct multivariate quasar light curves and infer their physical properties such as the black hole mass, inclination angle, and temperature slope. Our model is trained on realistic simulations of LSST ten year quasar light curves, and we demonstrate its ability to reconstruct quasar light curves even in the presence of long seasonal gaps and irregular sampling across different bands, outperforming a multioutput Gaussian process regression baseline. Our method has the potential to provide a deeper understanding of the physical properties of quasars and is applicable to a wide range of other multivariate time series with missing data and irregular sampling.

79 ASTRONOMY AND ASTROPHYSICS↗

Constraining Arctic Climate Projections of Wintertime Warming With Surface Turbulent Flux Observations and Representation of Surface-Atmosphere Coupling

The drivers of rapid Arctic climate change—record sea ice loss, warming SSTs, and a lengthening of the sea ice melt season—compel us to understand how this complex system operates and use this knowledge to enhance Arctic predictability. Changing energy flows sparked by sea ice decline, spotlight atmosphere-surface coupling processes as central to Arctic system function and its climate change response. Despite this, the representation of surface turbulent flux parameterizations in models has not kept pace with our understanding. The large uncertainty in Arctic climate change projections, the central role of atmosphere-surface coupling, and the large discrepancy in model representation of surface turbulent fluxes indicates that these processes may serve as useful observational constraints on projected Arctic climate change. This possibility requires an evaluation of surface turbulent fluxes and their sensitivity to controlling factors (surface-air temperature and moisture differences, sea ice, and winds) within contemporary climate models (here Coupled Model Intercomparison Project 6). The influence of individual controlling factors and their interactions is diagnosed using a multi-linear regression approach. This evaluation is done for four sea ice loss regimes, determined from observational sea ice loss trends, to control for the confounding effects of natural variability between models and observations. The comparisons between satellite- and model-derived surface turbulent fluxes illustrate that while models capture the general sensitivity of surface turbulent fluxes to declining sea ice and to surface-air gradients of temperature and moisture, substantial mean state biases exist. Specifically, the central Arctic is too weak of a heat sink to the winter atmosphere compared to observations, with implications to the simulated atmospheric circulation variability and thermodynamic profiles. Models were found to be about 50% more efficient at turning an air-sea temperature gradient anomaly into a sensible heat flux anomaly relative to observations. Further, the influence of sea ice concentration on the sensible heat flux is underestimated in models compared to observations. The opposite is found for the latent heat flux variability in models; where the latent heat flux is too sensitive to a sea ice concentration anomaly. Lastly, the results suggest that present-day trends in sea ice retreat regions may serve as suitable observational constraints of projected Arctic warming.

turbulent fluxes↗