Search NASA⌕ Search

SEARCH · Search NASA

Results for “interpretable models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A methodology for decay heat characterization in molten salt reactors

Accurate decay heat prediction in molten salt reactors (MSRs) faces dual challenges: complex operational uncertainties and the need for interpretable models compatible with engineering workflows. This work presents a hybrid machine learning and segmented polynomial methodology that addresses both requirements through three key innovations. First, a modular data architecture encodes MSR-specific operational parameters (power density: 1-100 W cm -3 , humidity: 0-0.1 wt %, air ingress: 0-0.1 mol %) with uncertainty-aware temporal discretization spanning 15 orders of magnitude. Second, region-optimized machine learning models achieve 92.3 % root mean square error (RMSE) reduction over conventional polynomials while maintaining physical interpretability through automated piecewise equation generation. Third, dual front-end interfaces accelerate safety analyses — a Jupyter environment enables researchers to explore 10,000+ parameter combinations via interactive widgets, while a Streamlit web application reduces design iteration cycles through production-grade visualization tools. Operational deployment demonstrates prediction times of only a couple hundred milliseconds for 10 4 years decay profiles, enabling real-time optimization of spent fuel container designs.

42 - ENGINEERING↗

Including frameworks of public health ethics in computational modelling of infectious disease interventions

Decisions on public health interventions to control infectious diseases are often informed by computational models. Interpreting the predicted outcomes of a public health decision requires not only high-quality modelling but also an ethical framework for assessing the benefits and harms associated with different options. The design and specification of ethical frameworks matured independently of computational modelling, so many values recognized as important for ethical decision-making are missing from computational models. We demonstrate a proof-of-concept approach to incorporate multiple public health values into the evaluation of a simple computational model for vaccination against a pathogen such as SARS-CoV-2. By examining a bounded space of alternative prioritizations of three values relevant to public health ethics (aggregate clinical burden, equity in clinical burden, equity in adverse effects from vaccination), we identify value trade-offs, where the outcomes of optimal strategies differ depending on the ethical framework. This work demonstrates an approach to incorporating diverse values into decision criteria used to evaluate outcomes of models of infectious disease interventions.

"Mathematical Biology"↗

Basin & Range Investigation for Developing Geothermal Energy

Hidden geothermal systems represent a potentially prolific energy resource that could support critical U.S. public and government energy priorities. Basin and Range Investigations for Developing Geothermal Energy (BRIDGE) addressed some the challenges associated with hidden system exploration by prioritizing cost-effective exploration early on through strategic workflow and informed decision-making that mitigates early risk and shifts resources to later exploration stages (e.g., drilling). Sandia National Laboratories partnered with U.S. Navy Geothermal Office, Geologic Geothermal Group, and independent consultants, with additional collaboration with U.S. Geological Survey and private industry. The primary tool of the BRIDGE project was to deploy a regional-scale airborne electromagnetic method to investigate the shallow resistivity structure in areas with high prospectivity. This was followed up at several prospects by a multidisciplinary exploration approach, including additional geologic, geophysical and geochemical studies. A central tenet to the BRIDGE methodology is that zones of low resistivity frequently occur over geothermal systems in the Basin and Range, and when paired with other data constraints, imaging these zones can enable discovery of these systems. In addition to exploring greenfield areas (i.e., Grover Point), the BRIDGE project also flew HTEM resistivity surveys over known geothermal systems including those with established power plants (Don A. Campbell and Salt Wells) and prospects that are known to the literature but remain undeveloped, at least in part, due to a lack of understanding on the location of their producible reservoirs. BRIDGE produced a comprehensive set of data from prospects identified in the Nevada Play Fairway Analysis along with conceptual models for top ranking prospects, wherein all of the observations are used to inform an interpreted model of the system. These models present a range of possible system parameters such as temperature and size, and they are further informed by system analogues in the Basin and Range province and elsewhere. The results of this work leave space for further exploration that may now occur at prospects ‘down the list’ rather than distribution exploration resources evenly across all prospects.

15 GEOTHERMAL ENERGY↗

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling↗

Jet classification using high-level features from anatomy of top jets

Recent advancements in deep learning models have significantly enhanced jet classification performance by analyzing low-level features (LLFs). However, this approach often leads to less interpretable models, emphasizing the need to understand the decision-making process and to identify the high-level features (HLFs) crucial for explaining jet classification. To address this, we consider the top jet tagging problems and introduce an analysis model (AM) that analyzes selected HLFs designed to capture important features of top jets. Our AM mainly consists of the following three modules: a relation network analyzing two-point energy correlations, mathematical morphology and Minkowski functionals for generalizing jet constituent multiplicities, and a recursive neural network analyzing subjet constituent multiplicity to enhance sensitivity to subjet color charges. We demonstrate that our AM achieves performance comparable to the Particle Transformer (ParT) while requiring fewer computational resources in a comparison of top jet tagging using jets simulated at the hadronic calorimeter angular resolution scale. Furthermore, as a more constrained architecture than ParT, the AM exhibits smaller training uncertainties because of the bias-variance tradeoff. We also compare the information content of AM and ParT by decorrelating the features already learned by AM. Lastly, we briefly comment on the results of AM with finer angular resolution inputs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Fully generalized, turbulent trace impurity transport with Gkeyll and Flan in the DIII-D far-SOL

The Monte Carlo trace impurity turbulent transport code Flan is introduced for the first time. Flan follows impurities in a turbulent background plasma from Gkeyll using the Lorentz force to resolve the full particle gyro orbit. Collisions are handled using the Nanbu collision algorithm (Nanbu 1997 Phys. Rev. E 55 4642–52), and ionization/recombination is handled via ADAS coupling. The far-SOL of a generic DIII-D L-mode is simulated with and without the collision model to show how collisions affect radial tungsten transport. Anomalous diffusion coefficient (D r ) and pinch velocity (v p ) profiles are extracted from fits to the results. With collisions, D r and v p are between 0–1.0 m 2 s −1 and −100–100 m s −1 , respectively. Without collisions, D r and v p are between 0–0.3 m 2 s −1 and −50–50 m s −1 , respectively. Exponential fits to the radial W density profiles and experimental data from W deposition along a collector probe are in good agreement, demonstrating Flan as a useful interpretive modeling tool. Additional simulations show that impurity transport away from the wall increases with atomic number, though it is not clear why. Flan has the potential to better interpret existing data and improve reactor scale predictions of core contamination because the underlying physics model is very general and does not rely on arbitrary user-defined transport coefficients.

DIII-D↗

Interpretable machine learning-guided design of Fe-based soft magnetic alloys

Here, we present a machine learning (ML) guided approach to predict saturation magnetization (𝑀 S ) and coercivity (𝐻 C ) in Fe-rich soft magnetic alloys, particularly Fe-Si-B systems. ML models trained on experimental data reveal that increasing Si and B content reduces 𝑀 S from 1.81 T (DFT ≈ 2.04 T) to ≈1.54 T (DFT ≈ 1.56T) in Fe-Si-B, which is attributed to decreased magnetic density and structural modifications. Experimental validation of ML predicted magnetic saturation on Fe-1Si-1B (2.09 T), Fe-5Si-5B (2.01 T), and Fe-10Si-10B (1.54 T) alloy compositions further supports our findings. These trends are consistent with density functional theory predictions, which link increased electronic disorder and band broadening to lower 𝑀 S values. Experimental validation on selected alloys confirms the predictive accuracy of the ML model, with good agreement across compositions. Beyond predictive accuracy, detailed uncertainty quantification and model interpretability including through feature importance and partial dependence analysis reveal that 𝑀 S is governed by a nonlinear interplay between Fe content and early transition metal ratios, while 𝐻 C is more sensitive to processing conditions such as ribbon thickness and thermal treatment windows. The ML framework was further applied to Fe-Si-B/Cr/Cu/Zr/Nb alloys in a pseudoquaternary compositional space, which shows comparable magnetic properties to NANOMET (Fe 84.8 ⁢Si 0.5 ⁢B 9.4 ⁢Cu 0.8⁢ P 3.5 ⁢C 1 ), FINEMET (Fe 73.5 ⁢Si 13.5 ⁢B 9 Cu 1 ⁢Nb 3 ), NANOPERM (Fe 88 ⁢Zr 7⁢ B 4 ⁢Cu 1 ), and HITPERM (Fe 44 ⁢Co 44 ⁢Zr 7⁢ B 4 ⁢Cu 1 . Our findings demonstrate the potential of the ML framework for accelerated search of high-performance soft magnetic materials.

density functional theory↗

Prediction of Creep-Induced Strain Using a Symbolic Regression-Based Model

Material creep under high-temperature conditions limits the lifetime and safety of structural systems such as advanced nuclear reactors. Conventional creep testing is slow and often produces inconsistent results across nominally identical experiments, making lifetime prediction uncertain. Here, to address these challenges, this work develops a data-driven symbolic regression (SR) model that consolidates results from duplicate creep tests and predicts the remaining strain-time curve of an ongoing experiment. The method uses piece-wise multi-objective SR with physical constraints to generate analytic, interpretable functions describing transient creep strain. Applied to Inconel Alloy 617 data, the approach achieved relative mean absolute errors of 1.0–9.5%, providing closed-form predictions of strain evolution. These results demonstrate a first step toward reducing the duration and cost of long-term creep testing while retaining physically interpretable model forms.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

DECOVALEX-2023: Task A Final Report

Task A, also known as HGFrac, examines the fracturing processes that may occur in the Callovo-Oxfordian claystone (COx) in the context of the high-level (HLW) and intermediate-level long-lived (ILW-LL) radioactive waste repository in France. Understanding these processes and improving numerical models to reproduce them will aid in the design, optimization, and safety of the repository. Heat and gas fracturing are studied in two independent subtasks following a stepwise approach: laboratory tests/benchmark exercises, in-situ experiment, and finally, an application case. The in-situ heater experiment aimed to thermally induce a hydraulic fracture; temperature and pore pressure were monitored to detect any evidence of fracturing. The in-situ gas injection experiment aimed to study the effect of the stress orientation and gas injection kinetics on the gas fracturing process. The occurrence of fracturing was monitored by gas pressure measured in the injection interval. In both experiments, the excavation-induced fracture network around the heater/injection boreholes played an important role in the reduction of the compressive stress state, leading to both a tensile and shear failure response of the COx. During the first two years of the project, the research teams working on each task developed and/or proposed numerical approaches for reproducing the occurrence of fracturing in the in-situ experiments. The failure criteria were defined by reproducing the measurements from laboratory extension tests for the heat fracturing subtask. In the gas fracturing subtask, their approaches were used for simulating several benchmark exercises, and an inter-comparison between models was carried out. In both tasks, the developed approaches were compared with a simplified approach considering poro-elasticity for the mechanical behaviour of the COx. Most of the developed approaches are based on a continuous medium that takes into account variations in hydraulic properties due to mechanical degradation, such as plastic deformation or damage. Other approaches implicitly modelled weak planes or embedded discontinuities to reproduce fracture propagation. The potential for fracture initiation was also studied through of a discrete approach. In the second half of the project, the research teams mainly focused on interpretative modelling of two in-situ experiments and a blind prediction exercise to test their respective approaches. The models developed by the research teams were also applied at the repository scale to evaluate fracture initiation in a case study under vi unfavourable conditions, particularly in terms of spacing between High-Level Waste cells. The results showed that the poro-elasticity approach could be an efficient tool for understanding the main processes occurring in the COx. One example is the explicit representation of the excavation-induced fracture network around the boreholes, which yielded acceptable results compared to the measurement data. However, advanced approaches were needed to evaluate the potential increase of the excavation-induced fracture network extend and better understand fracture initiation. The stress analyses carried out by the teams revealed that hydraulic boundary conditions had a strong impact on fracture initiation in the heater experiment. Furthermore, in most cases, the results required higher pore pressure increments to reach fracturing than those measured in the experiment. This implies that the measurements may have been biased by the packer’s capacity to fully isolate the piezometric chambers, leading to lower pressures. On the contrary, there was no agreement on the fracturing mode, as some reported either shear or tensile fracturing, while others reported a combination of the two modes. In the case of gas fracturing, the research teams were limited to the comparison of a single point, which complicated their task. Nonetheless, the numerical results were able to reproduce the measurements and capture processes such as longitudinal gas flow through the excavation-induced fracture network, as suggested by some evidence in the observation piezometric chambers. The numerical models also agreed with the measurements in the sense of higher probability of developing along the injection borehole than radially towards the sound rock. The approaches developed by the research teams showed that they are capable of analysing and reproducing fracture initiation in the COx. However, areas of future work should focus on the fracture propagation and fracture aperture, which were out of the scope of this task. To this end, additional data must be gathered for the parameter characterisation and validation of the numerical models. Nonetheless, various approaches showed promising results as they were able to reproduce fracture development under certain conditions.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Balancing Trade-offs: Adaptive Differential Privacy in Interpretable Machine Learning Models

In the advancing field of machine learning, balancing accuracy, interpretability, and privacy represents a significant challenge. The problem is exacerbated by the widespread deployment of pre-trained models locally in diverse applications, which could lead to various amounts of privacy leakage. Conventional Differential Privacy strategies, in which uniform noises are applied to model gradients, guarantee data privacy at the expense of accuracy and interpretability. This paper introduces a Feature-Sensitive Adaptive Differential Privacy (FADP) framework with a unique noise-adding strategy. Noises are adaptively added based on feature importance clustering, where important features are considered for interpretability. By employing a unique masking technique, FADP selectively preserves crucial features with minimal noise interference, maintaining accuracy while enhancing interpretability. The FADP framework addresses the limitations of traditional DP methods by preserving critical channels and improving interpretability — a vital requirement in machine learning applications that demand transparency in model decisions. Through comprehensive testing, FADP is shown to balance the trade-offs among accuracy, privacy, and interpretability, marking a substantial advancement in the field of privacy-preserving machine learning.

Farhad Riya, Farhin [University of Tennessee, Knox↗

Particle Markov Chain Monte Carlo Approach to Inference in Transient Surface Kinetics

Here, in this work, we develop a novel Bayesian approach to study the adsorption and desorption of CO onto a Pd(111) surface, a process of great importance in natural sciences. The motivation for this work comes from the recent availability of time-resolved infrared spectroscopy data and the need for model interpretability and uncertainty quantification in chemical processes. The objective is to learn the relevant parameters that characterize the process: coverage with time, rate constants, activation energies, and pre-exponential factors. Our approach consists of three main schemes: (i) a problem design and probabilistic model for the whole system, (ii) a particle Markov chain Monte Carlo sampler to learn the hidden coverages and rate constant parameters, and (iii) two Bayesian formulations to infer the activation energies and pre-exponential factors. The flexibility of the Bayesian framework allows for uncertainty quantification where possible and integration of mathematical constraints in the model to reflect the system physically. We found that our results for the activation energies and pre-exponential factor are in agreement with those reported in the experimental literature, independently, and we provide discussions on the advantages and disadvantages as well as applicability to other systems.

36 MATERIALS SCIENCE↗

Street-level temperature estimation using graph neural networks: Performance, feature embedding and interpretability

Estimating street-level air temperature is a challenging task due to the highly heterogeneous urban surfaces, canyon-like street morphology, and the diverse physical processes in the built environment. Though pioneering studies have embarked on investigations via data-driven approaches, many questions remain to be answered. Here, in this study, we leveraged an innovative framework and redefined the street-level temperature estimation problem using Graph Neural Networks (GNN) with spatial embedding techniques. The results showed that GNN models are more capable and consistent of estimating street-level temperature among tested locations, benefiting from its unique strength in handling extensive data over unstructured graph topology. In addition, we conducted in-depth analysis of feature importance to enhance the model interpretability. Among the urban features analyzed in this study, the time-variant canopy density and meter-level land use data emerge as crucial factors. Our findings highlight GNN 's high potential in capturing the complex dynamics between urban elements and their impacts on microclimate, thus offering valuable insights for comprehensive urban data collection and urban climate modeling in general. Collectively, this study also contributes to urban planning and policy by providing avenues to enhance city resilience against climate change, thereby advancing the agenda for environmental stewardship and urban sustainability.

54 ENVIRONMENTAL SCIENCES↗

CMPLE: Correlation Modeling to Decode Photosynthesis Using the Minorize–Maximize Algorithm

In plant genomic experiments, correlations among various biological traits (phenotypes) give new insights into how genetic diversity may have tuned biological processes to enhance fitness under diverse conditions. Consequently, knowing how the correlations are affected by genetic (G) and environmental (E) factors helps develop climate-resilient plants. However, the current literature lacks any method for assessing the effect of predictors on pairwise correlations among multiple phenotypes together with easily interpretable model parameters. To address this need, we propose to model pairwise correlations directly in terms of G and E and develop a computationally efficient inference procedure. Two major novelties in our methodology are (1) the use of a composite pairwise likelihood method to avoid the positive definiteness restriction on the correlation matrix and (2) the use of a novel Minorize–Maximize (MM) algorithm for the efficient estimation of a large number of parameters. The proposed method shows excellent numerical performance on synthetic datasets. Here, the analysis of the motivating data on cowpea reveals that the rates of solar energy storage by photosynthesis (the aggregate trait) are differentially affected by different genetic loci through two distinct processes: “photoinhibition” which results from photodamage caused by excess light, and “photoprotection” which protects plants from photodamage but also results in energy loss.

Correlation modeling↗

DEMO-FTES: Development, Monitoring, and Control of Fracture Thermal Energy Storage in Crystalline Rock Formations (CRADA Final Report)

The DEMO-FTES project investigated the feasibility of Fracture Thermal Energy Storage (FTES) as a seasonal energy storage solution in crystalline rock formations. FTES leverages hydraulically induced fractures to exchange heat between circulating fluids and the surrounding rock mass, enabling long-term thermal energy retention due to the high specific heat and low thermal conductivity of rock. This approach has the potential to reduce heating and cooling energy demands and enhance building energy resilience. The project combined dimensional analysis, numerical modeling, laboratory experiments, and meso-scale field tests to evaluate FTES performance and advance its technology readiness level from 3 to 5. Scaling analysis identified key dimensionless parameters governing heat transfer and fluid flow, ensuring laboratory and field tests were representative of larger-scale systems. Numerical simulations using TOUGH and iTOUGH2 frameworks supported experiment design and interpretation, modeling fracture geometry, thermal-hydraulic behavior, and thermo-mechanical coupling. Laboratory tests at EPFL involved creating single and multiple fractures in 25 cm cubic samples of Gabbro and Granite under true triaxial stress.

25 ENERGY STORAGE↗

Statistical inference of anomalous thermal transport with uncertainty quantification for interpretive 2D SOL models

The critical task of inferring anomalous cross-field transport coefficients is addressed in simulations of boundary plasmas with fluid models. A workflow for parameter inference in the UEDGE fluid code is developed using Bayesian optimization with parallelized sampling and integrated uncertainty quantification. In this workflow, transport coefficients are inferred by maximizing their posterior probability distribution, which is generally multidimensional and non-Gaussian. Uncertainty quantification is integrated throughout the optimization within the Bayesian framework that combines diagnostic uncertainties and model limitations. As a concrete example, we infer the anomalous electron thermal diffusivity $\chi_\perp$ from an interpretive 2D model describing electron heat transport in the conduction-limited region with radiative power loss. The workflow is first benchmarked against synthetic data and then tested on H-, L-, and I-mode discharges to match their midplane temperature and divertor heat flux profiles. We demonstrate that the workflow efficiently infers diffusivity and its associated uncertainty, generating 2D profiles that match 1D measurements. Future efforts will focus on incorporating more complicated fluid models and analyzing transport coefficients inferred from a large database of experimental results.

Bayesian optimization↗

Promoting the regulatory acceptance of combined ion and neutron irradiation for material degradation in nuclear reactors

The Advanced Materials and Manufacturing Technologies (AMMT) program within the Department of Energy (DOE) Office of Nuclear Energy has developed its current recommendation for promoting the use of combined ion irradiation and neutron irradiation for the accelerated qualification of materials to be deployed in nuclear reactors. This plan is intended to provide a collaborative path forward that can be adopted by academia, national laboratories, and industry, and has been developed with input from the regulatory research arm of the U.S. Nuclear Regulatory Commission (NRC). To deploy new materials or materials manufactured with new technologies, such as additive manufacturing, materials must be evaluated for reactor-induced degradation from the combination of harsh temperatures, corrosive environments, and radiation fields. However, rapid deployment of materials necessitates accelerated testing methods rather than relying on years of neutron irradiation in a material test reactor. Ion irradiation has demonstrated success in reproducing material microstructure and select property evolution resulting from neutron irradiation with three to four orders of magnitude reduction in time and cost, making it an ideal candidate for accelerated irradiation testing. This presentation provides context governing both the scientific and regulatory aspects of the proposed goal. The discussion is aimed at a broad audience including researchers from industry, national laboratories, and academia. The recommended path forward is presented as a conceptual framework of specific steps. In brief, the strategy entails developing an integrated ion and neutron irradiation test plan for the material property of interest based on the fundamental tenet of the linkage of microstructure and properties in materials. Physics-based modeling interprets ion irradiation data and predicts neutron irradiation microstructure and properties with uncertainty bounds. The first round of testing is sufficient for an initial licensing application using a risk-informed approach, while a minimum required neutron irradiation test plan reduces cost and time requirements. A surveillance program with witness specimens in-reactor provides additional data over time to improve model predictions to higher damage levels and further reduce uncertainty bounds, which can be used for license extensions or longer lifetimes in new license applications.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

SODAs: sparse optimization for the discovery of differential and algebraic equations

Differential-algebraic equations (DAEs) integrate ordinary differential equations (ODEs) with algebraic constraints, providing a fundamental framework for developing models of dynamical systems characterized by time-scale separation, conservation laws and physical constraints. While sparse optimization has revolutionized model development by allowing data-driven discovery of parsimonious models from a library of possible equations, existing approaches for dynamical systems assume DAEs can be reduced to ODEs by eliminating variables before model discovery. This assumption limits the applicability of such methods for DAE systems with unknown constraints and time scales. We introduce sparse optimization for differential-algebraic systems (SODAs), a data-driven method for the identification of DAEs in their explicit form. By discovering the algebraic and dynamic components sequentially without prior identification of the algebraic variables, this approach leads to a sequence of convex optimization problems. It has the advantage of discovering interpretable models that preserve the structure of the underlying physical system. To this end, SODAs improves since SODAs is singular numerical stability when handling high correlations between library terms, caused by near-perfect algebraic relationships, by iteratively refining the conditioning of the candidate library. We demonstrate the performance of our method on biological, mechanical and electrical systems, showcasing its robustness to noise in both simulated time series and real-time experimental data.

DAE↗

Predictive models of the genetic bases underlying budding yeast fitness in multiple environments

Abstract The ability of organisms to adapt and survive depends on the effects of genes and the environment on fitness. However, the multigenic nature of fitness and genotype-by-environment interactions hinder our understanding of the genetic basis of fitness. Here, we established fitness prediction models for 35 environments using machine learning and existing fitness data and different genetic variant types for a Saccharomyces cerevisiae population. Models revealed that the predictive ability of genetic variants varied across environments, with copy number variants explaining the majority of fitness variation in most cases. Model interpretation showed that different variant types identified distinct gene sets associated with predictive variants. These gene sets were significantly enriched in experimentally validated genes affecting fitness in only a subset of environments, indicating that many genes influencing fitness remain unexplored. Notably, non-experimentally validated genes were more important than validated ones for fitness predictions. Gene contributions to predictions were both isolate- and environment-dependent, pointing to gene-by-gene and gene-by-environment interactions. Furthermore, models uncovered experimentally validated and novel candidate genetic interactions for a well-characterized stress, the fungicide benomyl. These findings highlight the feasibility of identifying the genetic basis of fitness by using different genetic variant types and offer novel targets for future functional analysis.

DNA copy number variations↗