Search NASA⌕ Search

SEARCH · Search NASA

Results for “data inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Applying Gaussian Process Machine Learning and Modern Probabilistic Programming to Satellite Data to Infer CO 2 Emissions

Satellite data provides essential insights into the spatiotemporal distribution of CO 2 concentrations. However, many atmospheric inverse models fail to adequately incorporate the spatial and temporal correlations inherent in satellite observations and often lack rigorous methods for estimating parameters like spatial length scales. We introduce an inference model that processes the spatiotemporal covariance in satellite data and estimates hyperparameters such as covariance length scales. Our approach uses the Gaussian process (GP) machine learning (ML) and modern probabilistic programming languages (PPLs) to perform atmospheric inversions of emissions from satellite data. We develop a GP ML inversion system based on modern PPLs and the GEOS-Chem chemical transport model, simulating atmospheric CO 2 concentrations corresponding to the Orbiting Carbon Observatory-2/3 (OCO-2/3) data for July 2020. In our supervised learning framework, we treat the GEOS-Chem simulated data set as the target, with predictors derived by scaling the target with sector-specific factors hidden from the GP machine. Our results show that the GP model, combined with GPU-enabled PPLs, effectively retrieves true emission scaling factors and infers noise levels concealed within the data. This suggests that our method could be applied over larger areas with more complex covariance structures, enabling comprehensive analysis of the spatiotemporal patterns observed in OCO-2/3 and similar satellite data sets.

54 ENVIRONMENTAL SCIENCES↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗

The Swan: Data-driven Inference of Stellar Surface Gravities for Cool Stars from Photometric Light Curves

Stellar light curves are well known to encode physical stellar properties. Precise, automated, and computationally inexpensive methods to derive physical parameters from light curves are needed to cope with the large influx of these data from space-based missions such as Kepler and TESS. Here we present a new methodology that we call “The Swan,” a fast, generalizable, and effective approach for deriving stellar surface gravity (logg) for main-sequence, subgiant, and red giant stars from Kepler light curves using local linear regression on the full frequency content of Kepler long-cadence power spectra. With this inexpensive data-driven approach, we recover logg to a precision of ~0.02 dex for 13,822 stars with seismic logg values between 0.2 and 4.4 dex and ~0.11 dex for 4646 stars with Gaia-derived logg values between 2.3 and 4.6 dex. We further develop a signal-to-noise metric and find that granulation is difficult to detect in many cool main-sequence stars (T {sub eff} ≲ 5500 K), in particular K dwarfs. By combining our logg measurements with Gaia radii, we derive empirical masses for 4646 subgiant and main-sequence stars with a median precision of ~7%. Finally, we demonstrate that our method can be used to recover logg to a similar mean absolute deviation precision for a TESS baseline of 27 days. Our methodology can be readily applied to photometric time series observations to infer stellar surface gravities to high precision across evolutionary states.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Leveraging Hydropower Multi-Sensor Data for Inference and Age-Informed Modeling

Increased demand of operational flexibility such as faster ramp up/down in generation, and more frequent start/stops are putting hydropower plants and their associated components in unprecedented stress. Consequently, these plants are at the high risk of extended and more frequent outage to accommodate unscheduled, and unexpected maintenance. Therefore, hydropower plants are in critical need of data driven and age-informed analysis for their regular and unscheduled operation. Yet not all hydropower plants are exhaustively equipped with sensors and/or measurement streams for their respective components – demanding solutions on how to detect, identify, and locate the cause of any event from the unobservable. Idaho National Laboratory (INL) analyzed the anonymized measurements and event records from the Hydropower Research Institute (HRI) to address this issue, as part of the Water Power Technologies Office (WPTO) funded one year multi-lab project. First, we investigated how time series of multiple sensor measurements can be leveraged to identify an event “root cause” as well as to develop an inference (i.e., estimate the unobservable) problem. INL also investigated how individual hydropower components’ reaction or response times vary across the pre-event, during event, and post-event conditions – enabling the hydropower dynamic models to be age-informed. Finally, the impact of clustering multi-sensor time series on short-term vibration prediction is analyzed. INL will present key findings from these analyses and recommend next steps for stakeholder adoption.

13 HYDRO ENERGY↗

Observational process data analytics using causal inference

Voluminous process data are available with the paradigm shift toward smart manufacturing. However, most historical data are observational, containing noncausal correlations due to confounders and mediators. Estimating causal effects from observational data remains a bottleneck in leveraging them for active applications such as optimization and control. Further, this work aims to introduce a causal modeling framework for analyzing observational process data and extracting quantitative causal information. We demonstrate a real-world application in steel manufacturing where causal inference is used to analyze observational production data and improve the steelmaking process. Additionally, we propose a novel formulation for identifying critical process parameters from observational data, where causal inference is combined with variance-based methods to estimate corresponding risks of interventions to the manufacturing system. The proposed methods are compared with statistical ones to illustrate that causally interpreting statistical correlation leads to problematic results, while the provided workflow generates satisfactory strategies for process improvement.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Qualitative Strategy for Fusion of Physics into Empirical Models for Process Anomaly Detection

To facilitate the automated online monitoring of power plants, a systematic and qualitative strategy for anomaly detection is presented. This strategy is essential to provide credible reasoning on why and when an empirical versus hybrid (i.e., physics-supported) approach should be used and to determine the ideal mix of these two approaches for a defined anomaly detection scope. Empirical methods are usually based on pattern, statistical, and causal inference. Hybrid methods include the use of physics models to train and test data methods, reduce data dimensionality, reduce data-model complexity, augment data, and reduce empirical uncertainty; hybrid methods also include the use of data to tune physics models. The presented strategy is driven by key decision points related to data relevance, simple modeling feasibility, data inference, physics-modeling value, data dimensionality, physics knowledge, method of validation, performance, data availability, and suitability for training and testing, cause-effect, entropy inference, and model fitting. The strategy is demonstrated through a pilot use case for the application of anomaly detection to capture a valve packing leak at the high-pressure coolant injection system of a nuclear power plant.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data Center High-Temperature Liquid Cooling and Heat Reuse Techno-Economic Study: Preprint

Data centers are energy-intensive facilities with growing demands for efficiency and cost-effective operations. Smaller, more distributed edge inference data centers are expected to proliferate as AI applications require low latency closer to the user of AI tools, which presents a growing opportunity to explore the systems implications of liquid cooling on water and energy use. This study analyzes the implementation of high-temperature liquid cooling systems in a prototypical inference 1-MW data center and explores the potential for heat reuse across varying climates with a goal to optimize energy efficiency, reduce capital and operational costs, and identify opportunities for high-performance cooling and water use reduction infrastructure. This analysis evaluated configurations utilizing a peak day hourly sizing and systems performance spreadsheet to evaluate design and operational conditions from which component sizes, installed cost, operational cost, and performance metrics were determined for the Base case and the Elevated case. The techno-economic analysis included heat reuse applications across a range of heat recovery temperatures and heat rejection options. The analysis shows that high-temperature liquid cooling allows for improved energy efficiency, lower water consumption, and lower capital costs compared to traditional cooling approaches. Transitioning to elevated water inlet/outlet temperatures (50 degrees C/60 degrees C) eliminates the need for chillers, cooling towers, and heat recovery equipment in many scenarios across three distinct climate zones. This results in up to 75% capital cost savings for the cooling and heat recovery equipment, and with significantly reduced water consumption, especially in non-heat reuse applications. Heat generated from data centers can also be repurposed for space heating, domestic hot water, and other applications, and is most cost-effective when data center outlet temperatures exceed 55-60 degrees C.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

Leveraging Application Data Constraints to Optimize Database-Backed Web Applications

Exploiting the relationships among data is a classical query optimization technique. As persistent data is increasingly being created and maintained programmatically, prior work that infers data relationships from data statistics misses an important opportunity. We present Coco, the first tool that identifies data relationships by analyzing database-backed applications. Once identified, Coco leverages the constraints to optimize the application's physical design and query execution. Instead of developing a fixed set of predefined rewriting rules, Coco employs an enumerate-test-verify technique to automatically exploit the discovered data constraints to improve query execution. Each resulting rewrite is provably equivalent to the original query. Using 14 real-world web applications, our experiments show that Coco can discover numerous data constraints from code analysis and improve real-world application performance significantly.

Computer Science↗

Physics Informed Neural Networks as Computational Physics Emulators

This report is a brief overview and evaluation of Physics Informed Neural Networks (PINNs). Karniadakis and co-workers, e.g., Karniadakis et al. (2021) assert that the PINNs approach integrates seamlessly both data and mathematical physics models, even in partially understood, uncertain and high-dimensional contexts. They further claim that PINNs are effective and efficient for ill-posed and inverse problems, and when combined with domain decomposition, are scalable to large problems and a tool to discover hidden physics. While they demonstrate the capabilities in specific academic instances, their overarching claims about PINNs seem to be an overstatement, at least at the current time. We have briefly considered a few of the limitations of PINNs in this investigation. It is not clear to us if the PINNs approach can ever be competitive with approaches that use specialized algorithms to achieve high-accuracy solutions of governing equations and other techniques that can combine observational data with such solutions. As an example of the latter, consistent with the principles of Bayesian inference, data assimilation, or more generally data-model fusion, is a process that fuses observational data typically with a computational model that respects certain constraints such as conservation laws. For example, improvements in observational network combined with data assimilation have been key in improving weather predictions over the past four decades Kalnay (2003).

97 MATHEMATICS AND COMPUTING↗

Bayesian inference of multi-messenger astrophysical data: Joint and coherent inference of gravitational waves and kilonovae

Multi-messenger observations of binary neutron star mergers can provide information on the neutron star’s equation of state (EOS) above the nuclear saturation density by directly constraining the mass-radius diagram. We present a Bayesian framework for joint and coherent analyses of multi-messenger binary neutron star signals. As a first application, we analyze the gravitational-wave GW170817 and the kilonova (kN) AT2017gfo data. These results are then combined with the most recent X-ray pulsar analyses of PSR J0030+0451 and PSR J0740+6620 to obtain new EOS constraints.We extend the bajes infrastructure with a joint likelihood for multiple datasets, support for various semi-analytical kN models, and numerical-relativity (NR)-informed relations for the mass ejecta, as well as a technique to include and marginalize over modeling uncertainties. The analysis of GW170817 used the TEOBResumS effective-one-body waveform template to model the gravitational-wave signal. The analysis of AT2017gfo used a baseline multicomponent spherically symmetric model for the kN light curves. Various constraints on the mass-radius diagram and neutron star properties were then obtained by resampling over a set of ten million parameterized EOSs, which was built under minimal assumptions (general relativity and causality).

79 ASTRONOMY AND ASTROPHYSICS↗

Morphological Characters Can Strongly Influence Early Animal Relationships Inferred from Phylogenomic Data Sets

There are considerable phylogenetic incongruencies between morphological and phylogenomic data for the deep evolution of animals. This has contributed to a heated debate over the earliest-branching lineage of the animal kingdom: the sister to all other Metazoa (SOM). Here, we use published phylogenomic data sets ($\sim $45,000–400,000 characters in size with $\sim $15–100 taxa) that focus on early metazoan phylogeny to evaluate the impact of incorporating morphological data sets ($\sim $15–275 characters). We additionally use small exemplar data sets to quantify how increased taxon sampling can help stabilize phylogenetic inferences. We apply a plethora of common methods, that is, likelihood models and their “equivalent” under parsimony: character weighting schemes. Our results are at odds with the typical view of phylogenomics, that is, that genomic-scale data sets will swamp out inferences from morphological data. Instead, weighting morphological data 2–10$\times $ in both likelihood and parsimony can in some cases “flip” which phylum is inferred to be the SOM. This typically results in the molecular hypothesis of Ctenophora as the SOM flipping to Porifera (or occasionally Placozoa). However, greater taxon sampling improves phylogenetic stability, with some of the larger molecular data sets ($>$200,000 characters and up to $\sim $100 taxa) showing node stability even with $\geqq100\times $ upweighting of morphological data. Accordingly, our analyses have three strong messages. 1) The assumption that genomic data will automatically “swamp out” morphological data is not always true for the SOM question. Morphological data have a strong influence in our analyses of combined data sets, even when outnumbered thousands of times by molecular data. Morphology therefore should not be counted out a priori. 2) We here quantify for the first time how the stability of the SOM node improves for several genomic data sets when the taxon sampling is increased. 3) The patterns of “flipping points” (i.e., the weighting of morphological data it takes to change the inferred SOM) carry information about the phylogenetic stability of matrices. The weighting space is an innovative way to assess comparability of data sets that could be developed into a new sensitivity analysis tool.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying impacts of contact tracing on HIV epidemiological inference from phylogenetic data

Abstract Robust sampling methods are foundational to inferences using phylogenies. Yet the impact of using contact tracing, a type of non-uniform sampling used in public health applications such as infectious disease outbreak investigations, has not been investigated in the molecular epidemiology field. To understand how contact tracing influences a recovered phylogeny, we developed a new simulation tool called SEEPS (Sequence Evolution and Epidemiological Process Simulator) that allows for the simulation of contact tracing and the resulting transmission tree, pathogen phylogeny, and corresponding virus genetic sequences. Importantly, SEEPS takes within-host evolution into account when generating pathogen phylogenies and sequences from transmission histories. Using SEEPS, we demonstrate that contact tracing can significantly impact the structure of the resulting tree, as described by popular tree statistics. Contact tracing generates phylogenies that are less balanced than the underlying transmission process, less representative of the larger epidemiological process, and affects the internal/external branch length ratios that characterize specific epidemiological scenarios. We also examined real data from a 2007–2008 Swedish HIV-1 outbreak and the broader 1998–2010 European HIV-1 epidemic to highlight the differences in contact tracing and expected phylogenies. Aided by SEEPS, we show that the data collection of the Swedish outbreak was strongly influenced by contact tracing even after downsampling, while the broader European Union epidemic showed little evidence of universal contact tracing, agreeing with the known epidemiological information about sampling and spread. Overall, our results highlight the importance of including possible non-uniform sampling schemes when examining phylogenetic trees. For that, SEEPS serves as a useful tool to evaluate such impacts, thereby facilitating better phylogenetic inferences of the characteristics of a disease outbreak. SEEPS is available at https://github.com/MolEvolEpid/SEEPS.

Virology↗

Synthesizing realistic sand assemblies with denoising diffusion in latent space

Abstract The shapes and morphological features of grains in sand assemblies have far‐reaching implications in many engineering applications, such as geotechnical engineering, computer animations, petroleum engineering, and concentrated solar power. Yet, our understanding of the influence of grain geometries on macroscopic response is often only qualitative, due to the limited availability of high‐quality 3D grain geometry data. In this paper, we introduce a denoising diffusion algorithm that uses a set of point clouds collected from the surface of individual sand grains to generate grains in the latent space. By employing a point cloud autoencoder, the three‐dimensional point cloud structures of sand grains are first encoded into a lower‐dimensional latent space. A generative denoising diffusion probabilistic model is trained to produce synthetic sand that maximizes the log‐likelihood of the generated samples belonging to the original data distribution measured by a Kullback‐Leibler divergence. Numerical experiments suggest that the proposed method is capable of generating realistic grains with morphology, shapes and sizes consistent with the training data inferred from an F50 sand database. We then use a rigid contact dynamic simulator to pour the synthetic sand in a confined volume to form granular assemblies in a static equilibrium state with targeted distribution properties. To ensure third‐party validation, 50,000 synthetic sand grains and the 1542 real synchrotron microcomputed tomography (SMT) scans of the F50 sand, as well as the granular assemblies composed of synthetic sand grains are made available in an open‐source repository.

Vlassis, Nikolaos N.↗