Search NASASearch

SEARCH · Search NASA

Results for “Abstract Interpretation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ

Integrating adaptive learning with post hoc model explanation and symbolic regression to build interpretable surrogate models

Abstract We develop a materials informatics workflow to build an interpretable surrogate model for micromagnetic simulations. Our goal is to predict the energy barrier of a moving isolated skyrmion in rare-earth-free $$\hbox {Mn}_4$$ Mn 4 N. Our approach integrates adaptive learning with post hoc model explanation and symbolic regression methods. We discuss an unexplored acquisition function (information condensing active learning) within the adaptive learning loop and compare it with the known standard deviation function for efficient navigation of the search space. Model-agnostic post hoc explanation techniques then uncover trends learned by the trained model, which we then leverage to constrain the expressions used for symbolic regression. Graphical abstract

Biswas, Ankita

Poster Abstract: Leveraging Large Language Models to Reveal Interpretable Cooling Behaviors from Smart Thermostat Data

Frequent heatwaves and hot summers increasingly challenge occupant comfort, health, and energy grid stability. Addressing these challenges requires a detailed understanding of household cooling behaviors, such as thermostat adjustments and adaptive responses to extreme conditions. Traditional analyses often rely on aggregated numerical metrics that overlook subtle but important household-specific variations. In this study, we introduce a generalizable methodology that integrates large language models (LLMs) with vision capabilities to enable scalable and detailed analysis of residential thermostat data. Using Ecobee's Donate Your Data (DYD) dataset—which provides five-minute records of indoor temperatures, thermostat setpoints, and HVAC runtimes—we focus on two U.S. cities with contrasting summer climates : Austin (TX) and Phoenix (AZ). Because raw time-series data are not well suited for direct LLM analysis, we transform them into visual representations, such as daily indoor temperature trajectories and weekly runtime histograms, to better capture behavioral variations. Leveraging LLMs' visual interpretation, we extract descriptive behavioral features, including temperature preferences, time-of-day cooling orientation, anticipatory versus reactive heatwave responses, and behavioral consistency. These semantic features support unsupervised clustering to identify distinct occupant archetypes at scale, revealing differences—such as morning-centric anticipatory coolers versus households that shift toward warmer setpoints during heatwaves—that can inform demand response, resilience planning, and health-aware interventions. By converting raw numerical data into interpretable behavioral patterns, this methodology enables scalable and practical analysis of occupant behavior, supporting actionable insights for comfort, resilience, and energy management.

Nihar, Kopal

A Unified Interpretation of Variability in Precipitation Isotope Ratios

Abstract Several mechanisms have been proposed to explain why the isotope ratios of precipitation vary in space and time and why they correlate with other climate variables like temperature and precipitation. Here, we argue that this behavior is best understood through the lens of radiative transfer, which treats the depletion of atmospheric vapor transport by precipitation as analogous to the attenuation of light by absorption or scattering. Building on earlier work by Siler et al., we introduce a simple model that uses the equations of radiative transfer to approximate the two-dimensional pattern of the oxygen isotope composition of precipitation ( δ p ) from monthly mean hydrologic variables. The model accurately simulates the spatial and seasonal variability in δ p within a state-of-the-art climate model and permits a simple decomposition of δ p variability into contributions from gradients in evaporation and the length scale of vapor transport. Outside the tropics, δ p is mostly controlled by gradients in evaporation, whose dependence on temperature explains the positive correlation between δ p and temperature (i.e., the temperature effect). At low latitudes, δ p is mostly controlled by gradients in the transport length scale, whose inverse relationship with precipitation explains the negative correlation between δ p and precipitation (i.e., the amount effect). This suggests that the temperature and amount effects are both mostly explained by the variability in upstream rainout, but they reflect distinct mechanisms governing rainout at different latitudes. Significance Statement The isotopic composition of precipitation has long been used to make inferences about past climates based on its observed relationship with precipitation in the tropics and with temperature at higher latitudes. These relationships—known as the “amount effect” and “temperature effect,” respectively—have been attributed to many different mechanisms, most of which are thought to operate at either high or low latitudes but not both. Here, we present a unified framework for interpreting the isotope variability that can explain the latitude dependence of the temperature and amount effects despite making no distinction between high and low latitudes. Although our results are generally consistent with certain interpretations of the amount effect, they suggest that the temperature effect is widely misunderstood.

54 ENVIRONMENTAL SCIENCES

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie

Interpreting Mass and Radius Measurements of Neutron Stars with Dark Matter Halos

Abstract The high densities of neutron stars (NSs) could provide astrophysical locations for dark matter (DM) to accumulate. Depending on the DM model, these DM admixed NSs (DANSs) could have significantly different properties than pure baryonic NSs, accessible through X-ray observations of rotation-powered pulsars. We adopt the two-fluid formalism in general relativity to numerically simulate stable configurations of DANSs, assuming a fermionic equation of state (EOS) for the DM with repulsive self-interaction. The distribution of DM in the DANS as a halo affects the path of X-rays emitted from hot spots on the visible baryonic surface, causing notable changes in the pulse profile observed by telescopes such as NICER, compared to pure baryonic NSs. We explore how various DM models affect the DM mass distribution, leading to different types of dark halos. We quantify the deviation in observed X-ray flux from stars with each of these halos. We identify the pitfalls in interpreting mass and radius measurements of NSs inferred from electromagnetic radiation and constraining the baryonic matter EOS if these dark halos exist.

Shawqi, Shafayat (ORCID:0000000210956183)

Commutative Algebra Modeling in Materials Science – A Case Study on Metal–Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry’s constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

Khaemba, Caleb S.

The Analysis Description Language Ecosystem: Latest developments and physics applications

We present latest developments in Analysis Description Language (ADL), a declarative domain-specific language describing the physics algorithm of a HEP data analysis decoupled from software frameworks. Analyses written in ADL can be integrated into any framework for various tasks. ADL is a multipurpose construct with uses ranging from analysis design to preservation, reinterpretation, queries, visualisation, combination, etc. The most advanced infrastructure to execute ADL on events is the CutLang runtime interpreter. Recent technical developments include an automated interface with different data types, generation of the abstract syntax tree, a visualization tool that that auto-converts analysis flows to graphs, incorporation of trained machine learning models and a Jupyter-based plotting tool. We also report physics implications including a large scale LHC analysis implementation and validation effort for beyond the standard model reinterpretation purposes and studies with ATLAS and CMS open data.

Sekmen, Sezen [Kyungpook National Univ., Daegu (Ko

Element Formation in Radiation-hydrodynamics Simulations of Kilonovae

Abstract Understanding the details of r -process nucleosynthesis in binary neutron star merger (BNSM) ejecta is key to interpreting kilonova observations and identifying the role of BNSMs in the origin of heavy elements. We present a self-consistent, two-dimensional, ray-by-ray radiation-hydrodynamic evolution of BNSM ejecta with an online nuclear network (NN) up to a timescale of days. For the first time, an initial numerical relativity ejecta profile composed of the dynamical component and spiral-wave and disk winds is evolved including detailed r -process reactions and nuclear heating effects. A simple model for the jet energy deposition is also included. Our simulation highlights that the common approach of relating in postprocessing the final nucleosynthesis yields to the initial thermodynamic profile of the ejecta can lead to inaccurate predictions. Moreover, we find that neglecting the details of the radiation-hydrodynamic evolution of the ejecta in nuclear calculations can introduce deviations of up to 1 order of magnitude in the final abundances of several elements, including very light and second r -process peak elements. The presence of a jet affects element production only in the innermost part of the polar ejecta, and it does not alter the global nucleosynthesis results. Overall, our analysis shows that employing an online NN improves the reliability of nucleosynthesis and kilonova light-curve predictions.

Magistrelli, Fabio (ORCID:0009000509767851)

The importance of electron scattering in the analysis of actinide X-ray spectroscopy

Abstract Manifestations of electron scattering in X-ray spectroscopy have been evident for decades. Here, it will be shown that the proper interpretation of variants of X-ray Absorption Spectroscopy (XAS) of actinide materials must include an accurate treatment of features caused by electron scattering, i.e., EXAFS or Extended X-ray Absorption Fine Structure. These EXAFS features can be of such low energy that they are within ten to twenty electron volts of the Unoccupied Density of States (UDOS), immediately above the Fermi Energy or Band Gap. The adaption of simple models using the FEFF simulation program will be presented, including the demonstration of the robust nature of the results from different models. Graphical abstract

Tobin, J. G. (ORCID:0000000322943301)

AI-powered topic modeling: comparing LDA and BERTopic in analyzing opioid-related cardiovascular risks in women

Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.

Research & Experimental Medicine

Replacing non-biomedical concepts improves embedding of biomedical concepts

Embeddings are semantically meaningful representations of words in a vector space, commonly used to enhance downstream machine learning applications. Traditional biomedical embedding techniques often replace all synonymous words representing biological or medical concepts with a unique token, ensuring consistent representation and improving embedding quality. However, the potential impact of replacing non-biomedical concept synonyms has received less attention. Embedding approaches often employ concept replacement to replace concepts that span multiple words, such as non-small-cell lung carcinoma, with a single concept identifier (e.g., D002289). Also, all synonyms of each concept are merged into the same identifier. Here, we additionally leveraged WordNet to identify and replace sets of non-biomedical synonyms with their most common representatives. This combined approach aimed to reduce embedding noise from non-biomedical terms while preserving the integrity of biomedical concept representations. We applied this method to 1,055 biomedical concept sets representing molecular signatures or medical categories and assessed the mean pairwise distance of embeddings with and without non-biomedical synonym replacement. A smaller mean pairwise distance was interpreted as greater intra-cluster coherence and higher embedding quality. Embeddings were generated using the Word2Vec algorithm applied to a corpus of 10 million PubMed abstracts. Our results demonstrate that the addition of non-biomedical synonym replacement reduced the mean intra-cluster distance by an average of 8%, suggesting that this complementary approach enhances embedding quality. Future work will assess its applicability to other embedding techniques and downstream tasks. Python code implementing this method is provided under an open-source license.

algorithms

Dynamic allostery in the peptide/MHC complex enables TCR neoantigen selectivity

Abstract The inherent antigen cross-reactivity of the T cell receptor (TCR) is balanced by high specificity. Surprisingly, TCR specificity often manifests in ways not easily interpreted from static structures. Here we show that TCR discrimination between an HLA-A*03:01 (HLA-A3)-restricted public neoantigen and its wild-type (WT) counterpart emerges from distinct motions within the HLA-A3 peptide binding groove that vary with the identity of the peptide’s first primary anchor. These motions create a dynamic gate that, in the presence of the WT peptide, impedes a large conformational change required for TCR binding. The neoantigen is insusceptible to this limiting dynamic, and, with the gate open, upon TCR binding the central tryptophan can transit underneath the peptide backbone to the opposing side of the HLA-A3 peptide binding groove. Our findings thus reveal a novel mechanism driving TCR specificity for a cancer neoantigen that is rooted in the dynamic and allosteric nature of peptide/MHC-I binding grooves, with implications for resolving long-standing and often confounding questions about T cell specificity.

Science & Technology - Other Topics

Chemical ionization mass spectrometry utilizing benzene cations for measurements of volatile organic compounds and nitric oxide

We evaluate the capability of chemical ionization mass spectrometry (CIMS) using benzene cations as reagent ions (benzene CIMS) for detecting atmospheric trace gases. We characterize the ionization pathways and product ion distributions for 27 analytes spanning diverse chemical classes. To interpret the complex ion chemistry involving two reagent ions (C 6 H$^{+}_{6}$ and (C 6 H 6 )$^{+}_{2}$) and multiple ionization pathways (charge transfer, proton transfer, adduct formation, and hydride abstraction), we introduce a thermodynamics-based framework that classifies analytes into three categories based on their ionization energy (IE), relative to those of benzene monomer (9.24 eV) and dimer (8.69 eV). Each class exhibits distinct ionization mechanisms and product ions. Analytes with IE smaller than 8.69 eV (low IE) undergo charge transfer with both reagent ions; analytes with IE between 8.69 and 9.24 eV (mid IE) undergo charge transfer with C 6 H$^{+}_{6}$ and potential adduct formation with (C 6 H 6 )$^{+}_{2}$; analytes with IE larger than 9.24 eV (high IE) could undergo adduct formation, proton transfer, or hydride abstraction. Analytes within each class also show similar sensitivity, enabling sensitivity estimation for compounds lacking calibration standards. In addition to volatile organic compounds (VOCs), benzene CIMS detects nitric oxide (NO) with a detection limit of 5 pptv for 1 min integration time, exceeding the performance of most commercial NOx analyzers. Field deployments in Chicago and St. Louis demonstrate good agreement with reference NO measurements. Isoprene measurements show good agreement with a co-located gas chromatography–photoionization detector (GC-PID) in St. Louis, but exhibit substantial positive bias in Chicago, likely due to interferences from anthropogenic VOCs in the polluted urban environment. These results highlight the potential of benzene CIMS for concurrent measurements of NO, VOCs, and their oxidation products using a single instrument, while also underscoring challenges in complex atmospheric conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Chiral-odd generalized parton distributions in the large-𝑁 𝑐 limit of QCD: Spin-flavor structure, polynomiality, and sum rules

We study the nonperturbative properties of the nucleon’s chiral-odd generalized parton distributions (transversity GPDs) in the large-𝑁 𝑐 limit of QCD. This includes the parametric ordering of the spin-flavor components, the polynomiality property of the moments, and the sum rules connecting the GPDs with the tensor form factors. A multipole expansion in the transverse momentum transfer is used to enumerate and interpret the structures in the nucleon matrix element of the chiral-odd partonic operator, including monopole, dipole and quadrupole terms. The 1/𝑁 𝑐 expansion of the GPDs is performed using the abstract mean-field picture of baryons in the large-𝑁 𝑐 limit and its symmetries. We derive a large-𝑁 𝑐 relation between the flavor-nonsinglet GPDs 𝐸$^{𝑢−𝑑}_𝑇$ and $\tilde{𝐻}^{𝑢−𝑑}_𝑇$ and test it with recent lattice QCD results. We show that the polynomiality property and sum rules of the GPDs are fulfilled with the restricted realization of translational and rotational invariance in the mean-field picture. The results provide a basis for the phenomenological analysis of chiral-odd GPDs and hard exclusive processes in the large-𝑁 𝑐 limit, and for calculations in specific dynamical models.

generalized parton distributions

Diagnosis of Alzheimer’s disease using plasma biomarkers adjusted to clinical probability

Abstract Recently approved anti-amyloid immunotherapies for Alzheimer’s disease (AD) require evidence of amyloid-β pathology from positron emission tomography (PET) or cerebrospinal fluid (CSF) before initiating treatment. Blood-based biomarkers promise to reduce the need for PET or CSF testing; however, their interpretation at the individual level and the circumstances requiring confirmatory testing are poorly understood. Individual-level interpretation of diagnostic test results requires knowledge of disease prevalence in relation to clinical presentation (clinical pretest probability). Here, in a study of 6,896 individuals evaluated from 11 cohort studies from six countries, we determined the positive and negative predictive value of five plasma biomarkers for amyloid-β pathology in cognitively impaired individuals in relation to clinical pretest probability. We observed that p-tau217 could rule in amyloid-β pathology in individuals with probable AD dementia (positive predictive value above 95%). In mild cognitive impairment, p-tau217 interpretation depended on patient age. Negative p-tau217 results could rule out amyloid-β pathology in individuals with non-AD dementia syndromes (negative predictive value between 90% and 99%). Our findings provide a framework for the individual-level interpretation of plasma biomarkers, suggesting that p-tau217 combined with clinical phenotyping can identify patients where amyloid-β pathology can be ruled in or out without the need for PET or CSF confirmatory testing.

Cell Biology

Thermal partition function of $$ {J}_3{\overline{J}}_3 $$ deformed AdS3

Abstract We derive a compact formula for the one-loop, bosonic string partition function of Euclideanized$$ {J}_3{\overline{J}}_3 $$ J 3 J ¯ 3 deformedAdS 3 with periodic Euclidean time as an integral transform of the partition function of the undeformed EuclideanizedAdS 3 . Such a deformation is interpretable as an irrelevant “single-trace$$ T\overline{T} $$ T T ¯ deformation” of the boundary. We will do this by first establishing a formal procedure to compute a worldsheet torus zero point function for an exactly marginal$$ J\overline{J} $$ J J ¯ deformation of a sigma model with U(1) L × U(1) R global symmetry. We then describe how this procedure is implemented on SL(2,R) sigma model and its Euclidean continuation. Finally, we describe the embedding of the deformed SL(2,R) torus amplitude into critical string theory and interpret the result as the leading perturbative contribution to the thermal partition function of the deformed theory.

Physics

Metastability and Ostwald step rule in the crystallisation of diamond and graphite from molten carbon

Abstract Experimental challenges in determining the phase diagram of carbon at temperatures and pressures near the graphite-diamond-liquid triple point are often related to the persistence of metastable crystalline or glassy phases, superheated crystals, or supercooled liquids. A deeper understanding of the crystallisation kinetics of diamond and graphite is crucial for effectively interpreting the outcomes of these experiments. Here, we reveal the microscopic mechanisms of diamond and graphite nucleation from liquid carbon through molecular simulations with first-principles machine learning potentials. Our simulations accurately reproduce the experimental phase diagram of carbon near the triple point and show that liquid carbon crystallises spontaneously upon cooling. Metastable graphite crystallises in the domain of diamond thermodynamic stability at pressures above the triple point. Furthermore, whereas diamond crystallises through a classical nucleation pathway, graphite follows a two-step process in which low-density fluctuations forego ordering. Calculations of the nucleation rates of the two competing phases confirm this result and reveal a manifestation of Ostwald’s step rule, where the strong metastability of graphite hinders the transformation to the stable diamond phase. Our results provide a key to interpreting melting and recrystallisation experiments and shed light on nucleation kinetics in polymorphic materials with deep metastable states.

Science & Technology - Other Topics