Search NASASearch

SEARCH · Search NASA

Results for “Causal inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A GPU‐Accelerated Generative Adversarial Model for Causal Inference

We develop a GPU-accelerated machine learning generative adversarial model designed to facilitate causal inferences from observational data. Our model's theoretical framework is conceptualized in a manner that is amenable to being operable and scalable for high-performance computing platforms. We leverage GPU acceleration to develop a parallel evolutionary algorithm to achieve large-scale parallel computation of the model within a now widely accessible computing platform. This capability both enhances computational speedup and efficiency and also extends the use of the model to a broader range of substantive research domains while maintaining the underlying theoretical properties of the model.

GPU

Granger causal inference for climate change attribution

Abstract Climate change detection and attribution (D&A) is concerned with determining the extent to which anthropogenic activities have influenced specific aspects of the global climate system. D&A fits within the broader field of causal inference, the collection of statistical methods that identify cause and effect relationships. There are a wide variety of methods for making attribution statements, each of which require different types of input data and focus on different types of weather and climate events and each of which are conditional to varying extents. Some methods are based on Pearl causality (direct experimental interference) while others leverage Granger (predictive) causality, and the causal framing provides important context for how the resulting attribution conclusion should be interpreted. However, while Granger-causal attribution analyses have become more common, there is no clear statement of their strengths and weaknesses relative to Pearl-causal attribution and no clear consensus on where and when Granger-causal perspectives are appropriate. In this prospective paper, we provide a formal definition for Granger-based approaches to trend and event attribution and a clear comparison with more traditional methods for assessing the human influence on extreme weather and climate events. Broadly speaking, Granger-causal attribution statements can be constructed quickly from observations and do not require computationally-intesive dynamical experiments. These analyses also enable rapid attribution, which is useful in the aftermath of a severe weather event, and provide multiple lines of evidence for anthropogenic climate change when paired with Pearl-causal attribution. Confidence in attribution statements is increased when different methodologies arrive at similar conclusions. Moving forward, we encourage the D&A community to embrace hybrid approaches to climate change attribution that leverage the strengths of both Granger and Pearl causality.

Risser, Mark D. (ORCID:0000000319561783)

Learning genetic perturbation effects with variational causal inference

Advances in sequencing technologies have enhanced the understanding of gene regulation in cells. In particular, Perturb-seq has enabled high-resolution profiling of the transcriptomic response to genetic perturbations at the single-cell level. This understanding has implications in functional genomics and potentially for identifying therapeutic targets. Various computational models have been developed to predict perturbational effects. While deep learning models excel at interpolating observed perturbational data, they tend to overfit in the lack of enough data and may not generalize well to unseen perturbations. In contrast, mechanistic models, such as linear causal models based on gene regulatory networks, hold greater potential for extrapolation, as they encapsulate regulatory information that can predict responses to unseen perturbations. However, their application has been limited to small studies due to overly simplistic assumptions, making them less effective in handling noisy, large-scale single-cell data. We propose a hybrid approach that combines a mechanistic causal model with variational deep learning, termed Single Cell Causal Variational Autoencoder (SCCVAE). The mechanistic model employs a learned regulatory network to represent perturbational changes as shift interventions that propagate through the learned network. SCCVAE integrates this mechanistic causal model into a variational autoencoder, generating rich, comprehensive transcriptomic responses. Our results indicate that SCCVAE exhibits superior performance over current state-of-the-art baselines for extrapolating to predict unseen perturbational responses. Additionally, for the observed perturbations, the latent space learned by SCCVAE allows for the identification of functional perturbation modules and simulation of single-gene knockdown experiments of varying penetrance, presenting a robust tool for interpreting and interpolating perturbational responses at the single-cell level.

59 BASIC BIOLOGICAL SCIENCES

Incorporating Biological Knowledge into Evaluation of Casual Regulatory Hypothesis

Biological data can be scarce and costly to obtain. The small number of samples available typically limits statistical power and makes reliable inference of causal relations extremely difficult. However, we argue that statistical power can be increased substantially by incorporating prior knowledge and data from diverse sources. We present a Bayesian framework that combines information from different sources and we show empirically that this lets one make correct causal inferences with small sample sizes that otherwise would be impossible.

Chrisman, Lonnie

Decomposing causality into its synergistic, unique, and redundant components

Causality lies at the heart of scientific inquiry, serving as the fundamental basis for understanding interactions among variables in physical systems. Despite its central role, current methods for causal inference face significant challenges due to nonlinear dependencies, stochastic interactions, self-causation, collider effects, and influences from exogenous factors, among others. While existing methods can effectively address some of these challenges, no single approach has successfully integrated all these aspects. Here, we address these challenges with SURD: Synergistic-Unique-Redundant Decomposition of causality. SURD quantifies causality as the increments of redundant, unique, and synergistic information gained about future events from past observations. The formulation is non-intrusive and applicable to both computational and experimental investigations, even when samples are scarce. We benchmark SURD in scenarios that pose significant challenges for causal inference and demonstrate that it offers a more reliable quantification of causality compared to previous methods.

applied mathematics

Assessing the cumulative effects of nearshore habitat restoration actions for multiple populations of juvenile salmon in Whidbey Basin, Washington: foundation and approach for synthesis and evaluation

Ecosystem restoration is a common tool for re-establishing ecosystem processes, structures, and functions to improve biodiversity and services in coastal and estuarine ecosystems. In the Salish Sea, salmon habitats have been fragmented, reduced in size, and diminished in quality, and the ecosystem processes that form and sustain these habitats have been degraded and disrupted as well. This loss is especially prevalent in estuaries, where up to 90% of former salmon habitat has been lost or compromised. Salmon species are integral to the identities and cultures of people in the Pacific Northwest, yet salmon abundances remain at historic lows, especially in urbanized areas. Recent investments in restoration are creating rearing habitat and repairing lost ecosystem function. However, restoration efforts in this region have largely proceeded at the site scale, with less attention to big-picture thinking regarding how restoration will effectively recover degraded or lost habitats for target species. As a result, no landscape-scale evaluation program exists, and the cumulative benefits of multiple interventions are unknown. We describe innovative methods for science synthesis related to the evaluation of cumulative effects of ecosystem restoration for Pacific salmon, using years of existing, but disparate data. Building from previous work on cumulative effects evaluation and incorporating a hierarchy of hypotheses approach, we propose using causal inference across numerous hypotheses in a framework to assess the cumulative benefits to Pacific salmon from multiple estuarine restoration projects. We present the framework as a method that can be used to address many complex questions and provide examples from the Salish Sea where the approach is being implemented. The framework draws on science synthesis from numerous fields and uses a hierarchy of hypotheses, causal analysis at multiple scales, and a new hierarchy of synthesis for assessing multiple lines of evidence documenting restoration effects on Pacific salmon. We propose causal inference to synthesize dissimilar data streams, in our case, to identify various manifestations of cumulative effects of restoration and benefits to salmon, and to further inform restoration and recovery planning. A unifying framework would allow for the detection of thresholds at which restoration provides measurable improvement and would greatly advance understanding of the effects of restoration on ecosystems.

59 BASIC BIOLOGICAL SCIENCES

Deep Koopman operators for causal discovery

Causal discovery aims to identify cause-effect mechanisms for better scientific understanding, explainable decision-making, and more accurate modeling. Standard statistical frameworks, such as Granger causality, lack the ability to quantify causal relationships in nonlinear dynamics due to the presence of complex feedback mechanisms, timescale mixing, and nonstationarity. Thus, applying these methods to study causal dynamics in real-world systems, such as the Earth, is a major challenge. Addressing this shortcoming, we leverage deep learning and a Koopman operator-theoretic formalism to present a class of causal discovery algorithms. Kausal uses deep Koopman operator methods to approximate nonlinear dynamics in a linearized vector space in which traditional causal inference methods such as Granger causality can be more easily applied. Our idealized experiments demonstrate Kausal’s superior ability in discovering and characterizing causal signals compared to existing deep learning and non-deep learning state-of-the-art approaches. Finally, the successful identification of major El Niño and La Niña events in observations showcases Kausal’s skill to handle real-world applications.

54 ENVIRONMENTAL SCIENCES

Measuring impacts of California agri-environmental programs using field-scale satellite data

In the past decade, California has invested over $\$$200 million in direct grants to growers to support the adoption of agricultural practices that save water and/or improve soil health while also reducing greenhouse gas emissions. Ex-post evaluation of agri-environmental outcomes of these grant programs, however, is limited. We use satellite data to monitor changes in field-level consumptive water use and greenness (i.e. normalized difference vegetation index), a proxy for agricultural productivity, for the most frequently funded crop-types (almonds, grapes, and walnuts) in two California Department of Food and Agriculture programs. Nearly 600 fields receiving funding during the 2014–2022 period were analyzed using two causal inference methods. Fields that received grants to both upgrade irrigation systems and install irrigation water management sensors showed reduced consumptive water use and greenness by an average of 3.5% and 4.2%, respectively (significant at the 10% level). In contrast, we find that the adoption of only irrigation water management sensors, which are designed to inform irrigation scheduling and management, resulted in an average increase of 4.1% and 4.8% in consumptive water use and greenness respectively (significant at the 5% level). We find negligible effects for either consumptive water use or greenness when both pump efficiency upgrades and sensors were implemented. We further find that grants for compost addition and cover cropping led to small greenness increases of 1.7% and 2.8% respectively (significant at the 10% level) and had insignificant effects on consumptive water use. Our analysis of five agri-environmental program interventions reveals that several practice outcomes may be at odds with stated program goals of reducing water use while maintaining or improving agricultural productivity.

agriculture

CRISPR-CARB/nocap

Network Optimization and Causal Analysis of Perturb-seq (NOCAP) is a software package for causal inference of gene regulation networks using data from perturb-seq.

George, August

Computational tools and data integration to accelerate vaccine development: challenges, opportunities, and future directions

The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.

60 APPLIED LIFE SCIENCES

Prediction and causal reasoning in planning

Nonlinear planners are often touted as having an efficiency advantage over linear planners. The reason usually given is that nonlinear planners, unlike their linear counterparts, are not forced to make arbitrary commitments to the order in which actions are to be performed. This ability to delay commitment enables nonlinear planners to solve certain problems with far less effort than would be required of linear planners. Here, it is argued that this advantage is bought with a significant reduction in the ability of a nonlinear planner to accurately predict the consequences of actions. Unfortunately, the general problem of predicting the consequences of a partially ordered set of actions is intractable. In gaining the predictive power of linear planners, nonlinear planners sacrifice their efficiency advantage. There are, however, other advantages to nonlinear planning (e.g., the ability to reason about partial orders and incomplete information) that make it well worth the effort needed to extend nonlinear methods. A framework is supplied for causal inference that supports reasoning about partially ordered events and actions whose effects depend upon the context in which they are executed. As an alternative to a complete but potentially exponential-time algorithm, researchers provide a provably sound polynomial-time algorithm for predicting the consequences of partially ordered events.

Dean, T.

Genetic Network Inference: From Co-Expression Clustering to Reverse Engineering

Advances in molecular biological, analytical, and computational technologies are enabling us to systematically investigate the complex molecular processes underlying biological systems. In particular, using high-throughput gene expression assays, we are able to measure the output of the gene regulatory network. We aim here to review datamining and modeling approaches for conceptualizing and unraveling the functional relationships implicit in these datasets. Clustering of co-expression profiles allows us to infer shared regulatory inputs and functional pathways. We discuss various aspects of clustering, ranging from distance measures to clustering algorithms and multiple-duster memberships. More advanced analysis aims to infer causal connections between genes directly, i.e., who is regulating whom and how. We discuss several approaches to the problem of reverse engineering of genetic networks, from discrete Boolean networks, to continuous linear and non-linear models. We conclude that the combination of predictive modeling with systematic experimental verification will be required to gain a deeper insight into living organisms, therapeutic targeting, and bioengineering.

Dhaeseleer, Patrik

Causal machine learning uncovers conditions for convective intensification driven by organic and sulfate aerosols

Aerosols are often hypothesized to invigorate deep convective clouds (DCCs), but observational evidence remains limited and inconclusive. Clarifying this hypothesis is critical for regions vulnerable to thunderstorms and flooding, particularly highly polluted coastal cities. Leveraging a novel causal discovery–inference pipeline and high-resolution observations near Houston, TX, we identify multiple causal pathways among aerosols (mostly organic and sulfate), DCCs, and meteorological factors. However, a direct causal link from aerosols to DCCs is found to be uncommon, occurring in less than 35% of analyzed scenarios, and is characterized by strong conditionality and nonlinearity. When aerosol impacts on DCCs do occur, they can be substantial, enhancing DCC core heights by approximately 1.7 km, with 92% of this effect concentrated in warmer-phase cloud regions. Notably, the presence of sea breezes and the inclusion of all measured aerosol particles each enhance DCCs in over 95% of aerosol-sensitive cases.

54 ENVIRONMENTAL SCIENCES

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference

Mapping causal patterns in crystalline solids

The evolution of the atomic structures of the combinatorial library of Sm-substituted thin film BiFeO 3 along the phase transition boundary from the ferroelectric rhombohedral phase to the non-ferroelectric orthorhombic phase is explored using scanning transmission electron microscopy. Localized properties, including polarization, lattice parameter, and chemical composition, are parameterized from atomic-scale imaging, and their causal relationships are reconstructed using a linear non-Gaussian acyclic model. This approach is further extended to explore the spatial variability of the causal coupling using the sliding window transform method, which revealed that new causal relationships emerged at both the expected locations, such as domain walls and interfaces, and at additional regions forming clusters in the vicinity of the walls or spatially distributed features. While the exact physical origins of these relationships are unclear, they likely represent nanophase-separated regions in the morphotropic phase boundaries. Overall, we posit that an in-depth understanding of complex disordered materials away from thermodynamic equilibrium necessitates understanding not only the generative processes that can lead to observed microscopic states but also the causal links between multiple interacting subsystems.

Causal inference

Causal Directions Matter: How Environmental Factors Drive Convective Cloud Detrainment Heights

This study investigates how environmental factors influence the level of maximum detrainment (LMD) in deep convective clouds. Through a novel application of the Linear Non‐Gaussian Acyclic Model (LiNGAM), we discover causal structures between environmental variables and LMD, observed at six tropical sites operated by the Atmospheric Radiation Measurement (ARM) user facility. LiNGAM effectively identifies causal directions among variables of interest, revealing robust relationships such as those among the lifting condensation level (LCL), level of free convection (LFC), and convective inhibition (CIN), aligning with prior knowledge. Relative humidity is shown to directly influence LMD; however, this relationship exhibits strong nonlinearity and becomes difficult to detect when the contrast between oceanic and continental environments is excluded from the analysis. This study highlights the importance of establishing causal relationships before performing statistical inference.

54 ENVIRONMENTAL SCIENCES

Robust Strategy for Rocket Engine Health Monitoring

Monitoring the health of rocket engine systems is essentially a two-phase process. The acquisition phase involves sensing physical conditions at selected locations, converting physical inputs to electrical signals, conditioning the signals as appropriate to establish scale or filter interference, and recording results in a form that is easy to interpret. The inference phase involves analysis of results from the acquisition phase, comparison of analysis results to established health measures, and assessment of health indications. A variety of analytical tools may be employed in the inference phase of health monitoring. These tools can be separated into three broad categories: statistical, rule based, and model based. Statistical methods can provide excellent comparative measures of engine operating health. They require well-characterized data from an ensemble of "typical" engines, or "golden" data from a specific test assumed to define the operating norm in order to establish reliable comparative measures. Statistical methods are generally suitable for real-time health monitoring because they do not deal with the physical complexities of engine operation. The utility of statistical methods in rocket engine health monitoring is hindered by practical limits on the quantity and quality of available data. This is due to the difficulty and high cost of data acquisition, the limited number of available test engines, and the problem of simulating flight conditions in ground test facilities. In addition, statistical methods incur a penalty for disregarding flow complexity and are therefore limited in their ability to define performance shift causality. Rule based methods infer the health state of the engine system based on comparison of individual measurements or combinations of measurements with defined health norms or rules. This does not mean that rule based methods are necessarily simple. Although binary yes-no health assessment can sometimes be established by relatively simple rules, the causality assignment needed for refined health monitoring often requires an exceptionally complex rule base involving complicated logical maps. Structuring the rule system to be clear and unambiguous can be difficult, and the expert input required to maintain a large logic network and associated rule base can be prohibitive.

Santi, L. Michael