Search NASA⌕ Search

SEARCH · Search NASA

Results for “Causal Discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Deep Koopman operators for causal discovery

Causal discovery aims to identify cause-effect mechanisms for better scientific understanding, explainable decision-making, and more accurate modeling. Standard statistical frameworks, such as Granger causality, lack the ability to quantify causal relationships in nonlinear dynamics due to the presence of complex feedback mechanisms, timescale mixing, and nonstationarity. Thus, applying these methods to study causal dynamics in real-world systems, such as the Earth, is a major challenge. Addressing this shortcoming, we leverage deep learning and a Koopman operator-theoretic formalism to present a class of causal discovery algorithms. Kausal uses deep Koopman operator methods to approximate nonlinear dynamics in a linearized vector space in which traditional causal inference methods such as Granger causality can be more easily applied. Our idealized experiments demonstrate Kausal’s superior ability in discovering and characterizing causal signals compared to existing deep learning and non-deep learning state-of-the-art approaches. Finally, the successful identification of major El Niño and La Niña events in observations showcases Kausal’s skill to handle real-world applications.

54 ENVIRONMENTAL SCIENCES↗

Space‐Time Causal Discovery in Earth System Science: A Local Stencil Learning Approach

Causal discovery tools enable scientists to infer meaningful relationships from observational data, spurring advances in fields as diverse as biology, economics, and climate science. Despite these successes, the application of causal discovery to space-time systems remains immensely challenging due to the high-dimensional nature of the data. For example, in climate sciences, modern observational temperature records over the past few decades regularly measure thousands of locations around the globe. To address these challenges, we introduce Causal Space-Time Stencil Learning (CaStLe), a novel meta-algorithm for discovering causal structures in complex space-time systems. CaStLe leverages regularities in local space-time dependencies to learn governing global dynamics. This local perspective eliminates spurious confounding and drastically reduces sample complexity, making space-time causal discovery practical and effective. For causal discovery, CaStLe flexibly accepts any appropriately adapted time series causal discovery algorithm to recover local causal structures. These advances enable causal discovery of geophysical phenomena that were previously unapproachable, including non-periodic, transient phenomena such as volcanic eruption plumes. Regularities in local space-time dependencies are transformed into informative spatial replicates, which actually improve CaStLe's performance when applied to ever-larger spatial grids. We successfully apply CaStLe to discover the atmospheric dynamics governing the climate response to the 1991 Mount Pinatubo volcanic eruption. We provide validation experiments to demonstrate the effectiveness of CaStLe over existing causal-discovery frameworks on a range of geophysics-inspired benchmarks while identifying the method's limitations and domains where its assumptions may not hold.

Nichol, J. Jake [Univ. of New Mexico, Albuquerque,↗

Methods for Causal Discovery

SAND2025-11742O Methods for Causal Discovery is a software tool that is used for causal discovery from data, including predicting and visualizing directed acyclic graphs from data using traditional machine learning techniques. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference↗

Causal machine learning uncovers conditions for convective intensification driven by organic and sulfate aerosols

Aerosols are often hypothesized to invigorate deep convective clouds (DCCs), but observational evidence remains limited and inconclusive. Clarifying this hypothesis is critical for regions vulnerable to thunderstorms and flooding, particularly highly polluted coastal cities. Leveraging a novel causal discovery–inference pipeline and high-resolution observations near Houston, TX, we identify multiple causal pathways among aerosols (mostly organic and sulfate), DCCs, and meteorological factors. However, a direct causal link from aerosols to DCCs is found to be uncommon, occurring in less than 35% of analyzed scenarios, and is characterized by strong conditionality and nonlinearity. When aerosol impacts on DCCs do occur, they can be substantial, enhancing DCC core heights by approximately 1.7 km, with 92% of this effect concentrated in warmer-phase cloud regions. Notably, the presence of sea breezes and the inclusion of all measured aerosol particles each enhance DCCs in over 95% of aerosol-sensitive cases.

54 ENVIRONMENTAL SCIENCES↗

A Causal Approach to Model Validation and Calibration

This poster presents a novel method for validation and verification that focuses on identifying causal relationships between data elements, moving beyond traditional statistical and machine learning approaches. These methods employ causal discovery techniques to reveal the underlying mechanisms of data generation. The research utilizes structural causal models and directed acyclic graphs to depict causal relationships. This approach assists in achieving alignment between simulation models and reality.

97 MATHEMATICS AND COMPUTING↗

Changing effects of external forcing on Atlantic–Pacific interactions

Recent studies have highlighted the increasingly dominant role of external forcing in driving Atlantic and Pacific Ocean variability during the second half of the 20th century. This paper provides insights into the underlying mechanisms driving interactions between modes of variability over the two basins. We define a set of possible drivers of these interactions and apply causal discovery to reanalysis data, two ensembles of pacemaker simulations where sea surface temperatures in either the tropical Pacific or the North Atlantic are nudged to observations, and a pre-industrial control run. We also utilize large-ensemble means of historical simulations from the Coupled Model Intercomparison Project Phase 6 (CMIP6) to quantify the effect of external forcing and improve the understanding of its impact. A causal analysis of the historical time series between 1950 and 2014 identifies a regime switch in the interactions between major modes of Atlantic and Pacific climate variability in both reanalysis and pacemaker simulations. A sliding window causal analysis reveals a decaying El Niño–Southern Oscillation (ENSO) effect on the Atlantic as the North Atlantic fluctuates towards an anomalously warm state. The causal networks also demonstrate that external forcing contributed to strengthening the Atlantic's negative-sign effect on ENSO since the mid-1980s, where warming tropical Atlantic sea surface temperatures induce a La Niña-like cooling in the equatorial Pacific during the following season through an intensification of the Pacific Walker circulation. The strengthening of this effect is not detected when the historical external forcing signal is removed in the Pacific pacemaker ensemble. The analysis of the pre-industrial control run supports the notion that the Atlantic and Pacific modes of natural climate variability exert contrasting impacts on each other even in the absence of anthropogenic forcing. The interactions are shown to be modulated by the (multi)decadal states of temperature anomalies of both basins with stronger connections when these states are “out of phase”. We show that causal discovery can detect previously documented connections and provides important potential for a deeper understanding of the mechanisms driving changes in regional and global climate variability.

54 ENVIRONMENTAL SCIENCES↗

A generative machine learning model for designing metal hydrides applied to hydrogen storage

Developing new metal hydrides is a critical step toward efficient hydrogen storage in carbon-neutral energy systems. However, existing materials databases, such as the Materials Project, contain a limited number of well-characterized hydrides, which constrains the discovery of optimal candidates. This work presents a framework that integrates causal discovery with a lightweight generative machine learning model to generate novel metal hydride candidates that may not exist in current databases. Using a dataset of 450 samples (270 training, 90 validation, and 90 testing), the model generates 1000 candidates. After ranking and filtering, six previously unreported chemical formulas and crystal structures are identified, four of which are validated by density functional theory simulations and show strong potential for future experimental investigation. Overall, the proposed framework provides a scalable and time-efficient approach for expanding hydrogen storage datasets and accelerating materials discovery.

generative model↗

Causal Directions Matter: How Environmental Factors Drive Convective Cloud Detrainment Heights

This study investigates how environmental factors influence the level of maximum detrainment (LMD) in deep convective clouds. Through a novel application of the Linear Non‐Gaussian Acyclic Model (LiNGAM), we discover causal structures between environmental variables and LMD, observed at six tropical sites operated by the Atmospheric Radiation Measurement (ARM) user facility. LiNGAM effectively identifies causal directions among variables of interest, revealing robust relationships such as those among the lifting condensation level (LCL), level of free convection (LFC), and convective inhibition (CIN), aligning with prior knowledge. Relative humidity is shown to directly influence LMD; however, this relationship exhibits strong nonlinearity and becomes difficult to detect when the contrast between oceanic and continental environments is excluded from the analysis. This study highlights the importance of establishing causal relationships before performing statistical inference.

54 ENVIRONMENTAL SCIENCES↗

Application of advanced causal analyses to identify processes governing secondary organic aerosols

Abstract Understanding how different physical and chemical atmospheric processes affect the formation of fine particles has been a persistent challenge. Inferring causal relations between the various measured features affecting the formation of secondary organic aerosol (SOA) particles is complicated since correlations between variables do not necessarily imply causality. Here, we apply a state-of-the-art information transfer measure coupled with the Koopman operator framework to infer causal relations between isoprene epoxydiol SOA (IEPOX-SOA) and different chemistry and meteorological variables derived from detailed regional model predictions over the Amazon rainforest. IEPOX-SOA represents one of the most complex SOA formation pathways and is formed by the interactions between natural biogenic isoprene emissions and anthropogenic emissions affecting sulfate, acidity and particle water. Since the regional model captures the known relations of IEPOX-SOA with different chemistry and meteorological features, their simulated time series implicitly include their causal relations. We show that our causal model successfully infers the known major causal relations between total particle phase 2-methyl tetrols (the dominant component of IEPOX-SOA over the Amazon) and input features. We provide the first proof of concept that the application of our causal model better identifies causal relations compared to correlation and random forest analyses performed over the same dataset. Our work has tremendous implications, as our methodology of causal discovery could be used to identify unknown processes and features affecting fine particles and atmospheric chemistry in the Earth’s atmosphere.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Verification, Validation, and Calibration Through a Causal Lens

While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

Decoding crops one cell at a time: from cell atlases to single-cell genetics

Understanding the mechanisms underlying key agricultural traits remains a central challenge in crop research, but recent advances in technologies are providing powerful tools to address this issue. Among these, single-cell and spatial transcriptomics have revealed tissue heterogeneity and spatial organization, offering unique insights into cellular gene expression dynamics and the coordinated activity of multiple cell types. These approaches help uncover how specific cell types contribute to agricultural traits and refine candidate loci lists through integration with trait-associated loci. Additionally, single-cell and spatial transcriptomics have the potential to serve as cell-level readout platforms integrating cellular perturbations, enabling high-throughput discovery of causal relationships between genotype and gene expression at the cellular level in plants. Successful implementation will accelerate the identification of key genetic variants for crop improvement. Furthermore we review lessons learned from application of single-cell screening in mammalian cells, highlight major technical and biological barriers to its use in plants, and outline potential strategies to overcome these challenges. Together, the widespread application and integration of single-cell and spatial transcriptomics with other technologies enable not only the descriptive cataloging of cell states but also the causal interrogation of sequence functions and regulatory networks at cell type resolution, ultimately advancing gene function studies and accelerating crop improvement.

Cellular heterogeneity↗

Discovery of potent small-molecule inhibitors of lipoprotein(a) formation

Lipoprotein(a) (Lp(a)), an independent, causal cardiovascular risk factor, is a lipoprotein particle that is formed by the interaction of a low-density lipoprotein (LDL) particle and apolipoprotein(a) (apo(a)). Apo(a) first binds to lysine residues of apolipoprotein B-100 (apoB-100) on LDL through the Kringle IV (K IV ) 7 and 8 domains, before a disulfide bond forms between apo(a) and apoB-100 to create Lp(a). Here we show that the first step of Lp(a) formation can be inhibited through small-molecule interactions with apo(a) K IV 7–8. We identify compounds that bind to apo(a) K IV 7–8, and, through chemical optimization and further application of multivalency, we create compounds with subnanomolar potency that inhibit the formation of Lp(a). Oral doses of prototype compounds and a potent, multivalent disruptor, LY3473329 (muvalaplin), reduced the levels of Lp(a) in transgenic mice and in cynomolgus monkeys. Although multivalent molecules bind to the Kringle domains of rat plasminogen and reduce plasmin activity, species-selective differences in plasminogen sequences suggest that inhibitor molecules will reduce the levels of Lp(a), but not those of plasminogen, in humans. These data support the clinical development of LY3473329—which is already in phase 2 studies—as a potent and specific orally administered agent for reducing the levels of Lp(a).

59 BASIC BIOLOGICAL SCIENCES↗

Active causal learning for decoding chemical complexities with targeted interventions

Abstract Predicting and enhancing inherent properties based on molecular structures is paramount to design tasks in medicine, materials science, and environmental management. Most of the current machine learning and deep learning approaches have become standard for predictions, but they face challenges when applied across different datasets due to reliance on correlations between molecular representation and target properties. These approaches typically depend on large datasets to capture the diversity within the chemical space, facilitating a more accurate approximation, interpolation, or extrapolation of the chemical behavior of molecules. In our research, we introduce an active learning approach that discerns underlying cause-effect relationships through strategic sampling with the use of a graph loss function. This method identifies the smallest subset of the dataset capable of encoding the most information representative of a much larger chemical space. The identified causal relations are then leveraged to conduct systematic interventions, optimizing the design task within a chemical space that the models have not encountered previously. While our implementation focused on the QM9 quantum-chemical dataset for a specific design task—finding molecules with a large dipole moment—our active causal learning approach, driven by intelligent sampling and interventions, holds potential for broader applications in molecular, materials design and discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗