Search NASASearch

SEARCH · Search NASA

Results for “Data Interpretation, Statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Historic climate, cosmogenic 10Be, denudation-rate, and geospatial datasets from the Pikes Peak region, Colorado, USA

This data package contains geographic information system (GIS) layers and tabular datasets associated with the study of elevation-dependent denudation rates on Pikes Peak in the Front Range of the Rocky Mountains, Colorado, USA. The package includes GIS layers used to produce the study-area map, including sample locations, sample watershed boundaries, the Pikes Peak batholith, Pleistocene glacier extent, weather station locations, and elevation and hillshade rasters, together with comma-separated value (CSV) tables and matching CSV data dictionaries. These mapped layers provide the geographic framework for interpreting denudation patterns across the Pikes Peak region and for relating sample locations to watershed geometry, bedrock setting, glacial history, and nearby climate stations. The first group of tables reports climate and geospatial context for the study area. These files include station-based temperature and precipitation data used to characterize elevational gradients in mean annual climate and monthly climate seasonality, sample locations, denudation-rate and topographic metrics, fixed frost-cracking model parameters, frost-cracking intensity and precipitation-frequency metrics, and stream-power inversion results. Together, these data provide the basis for evaluating how denudation varies with elevation, climate, and landscape form across sampled catchments on Pikes Peak. The second group of tables reports cosmogenic nuclide and erosion-model results used in the denudation analysis. Included files contain accelerator mass spectrometry (AMS) measurements for in situ-produced cosmogenic beryllium-10 (10Be), including sample identifiers, measured 10Be:9Be ratios, analytical uncertainties, carrier mass, quartz mass, blank corrections, blank-group statistics, and calculated 10Be concentrations and uncertainties. Additional tables summarize stream-power-law inversion results for sampled catchments, including optimized model parameters, predicted erosion rates, residual metrics, channel-pixel counts, and convergence status, as well as regression equations and summary statistics used to evaluate relationships among elevation, climate, frost cracking, precipitation forcing, and denudation rate. The package contains GIS files, comma-separated value files (.csv), Microsoft Excel files (.xlsx), CSV data dictionaries, a file-level metadata table, and a readme text file.

10Be cosmogenic nuclides

FORESTR: Finding, Organizing, Representing, Explaining, Summarizing, and Thinning Random forests

Random forests have become popular models used for data driven predictions. As a result, random forests are currently used or being considered for high-consequence mission applications in national security, such as the prediction of yield from optical signals and malware detection. While random forests may provide accurate predictions, the complexity of the algorithm causes a lack of interpretability. Random forests are an ensemble of regression or decision trees. Individual regression and decision trees are interpretable, but ensembles are inherently difficult to interpret due to the compilation of many models. We aim to increase the interpretability of random forests by finding patterns in the ensemble of trees that can be used to “thin” (or remove) trees. As a starting point, in this report, we develop a new distance metric for quantifying the similarity between trees based on their topologies (i.e., shapes). We base the metric on a novel distance metric for graphs that is a proper mathematical distance, is invariant to transformations, has registration between graphs, and computes topological evolutions between graphs. We use the tree distance metric to compute tree statistics such as a “mean tree” and to identify clusters of trees. We apply the developed methodology to a toy dataset and a mission relevant product inspection dataset to demonstrate how the metric can provide insight into random forests. Furthermore, we discuss the limitations of the approach and ideas for future research into how the metric could be used as a thinning tool to develop less complex models.

97 MATHEMATICS AND COMPUTING

The persistent shadow of the supermassive black hole of M87. II. Model comparisons and theoretical interpretations

The Event Horizon Telescope (EHT) observation of M87∗ in 2018 has revealed a ring with a diameter that is consistent with the 2017 observation. The brightest part of the ring is shifted to the southwest from the southeast. In this paper, we provide theoretical interpretations for the multi-epoch EHT observations for M87∗ by comparing a new general relativistic magnetohydrodynamics model image library with the EHT observations for M87∗ in both 2017 and 2018. The model images include aligned and tilted accretion with parameterized thermal and nonthermal synchrotron emission properties. The 2018 observation again shows that the spin vector of the M87∗ supermassive black hole is pointed away from Earth. A shift of the brightest part of the ring during the multi-epoch observations can naturally be explained by the turbulent nature of black hole accretion, which is supported by the fact that the more turbulent retrograde models can explain the multi-epoch observations better than the prograde models. The EHT data are inconsistent with the tilted models in our model image library. Assuming that the black hole spin axis and its large-scale jet direction are roughly aligned, we expect the brightest part of the ring to be most commonly observed 90 deg clockwise from the forward jet. This prediction can be statistically tested through future observations.

79 ASTRONOMY AND ASTROPHYSICS

Multimodality in the Search for New Physics in Pulsar Timing Data and the Case of Kination-amplified Gravitational-wave Background from Inflation

We investigate the kination-amplified inflationary gravitational-wave background (GWB) interpretation of the signal recently reported by various pulsar timing array (PTA) experiments. Kination is a post-inflationary phase in the expansion history dominated by the kinetic energy of some scalar field, characterized by a stiff equation of state w = 1. Within the inflationary GWB model, we identify two modes that can fit the current data sets (NANOGrav and EPTA) with equal likelihood: the kination-amplification (KA) mode and the ordinary, no-kination-amplification (no-KA) mode. The multimodality of the likelihood motivates a Bayesian analysis with nested sampling. We analyze the free spectra of current PTA data and mock free spectra constructed with higher signal-to-noise ratios using nested sampling. The analysis of the mock spectrum designed to be consistent with the best fit to the NANOGrav 15 yr (NG15) data successfully reveals the expected bimodal posterior for the first time while excluding the reheating mode that appears in the fit to the current NG15 data, making a case for our correct and comprehensive treatment of potential multimodal posteriors arising from future PTA data sets. The resultant Bayes factor is $\mathcal{B}$ $\equiv$ Z no–KA /Z KA = 2.9 ± 1.9, indicating comparable statistical significance between the two modes. Given the theoretical model-building challenges of producing highly blue-tilted primordial tensor spectra, the KA mode has the advantage of requiring less blue primordial spectra, compared with the no-KA mode. The synergy between future cosmic microwave background polarization, pulsar timing, and laser interferometer measurements of gravitational waves will help resolve the ambiguity implied by the multimodal posterior in PTA-only searches.

Cosmology

One-shot gas detection with transformer paired neural networks in Mako collected longwave infrared hyperspectral imagery

To date, careful data treatment workflows and statistical detectors are used to perform hyperspectral image (HSI) detection of any gas contained in a spectral library, which is often expanded with physics models to incorporate different spectral characteristics. In general, surrounding evidence or known gas-release parameters are used to provide confidence in or confirm detection capability, respectively. This makes quantifying detection performance difficult as it is nearly impossible to develop an absolute ground truth for gas target pixel presence in collected HSI. Consequently, developing and comparing new detection methods, especially machine learning (ML)-based methods, is susceptible to subjectivity in derived detection map quality. Here, in this work, we demonstrate the first use of transformer-based paired neural networks (PNNs) for one-shot gas target detection for multiple gases while providing quantitative classification and detection metrics for their use on labeled data. Terabytes of training data are generated from a database of long-wave infrared HSI obtained from historical Mako sensor campaigns over Los Angeles. By incorporating labels, singular signature representations, and a model development pipeline, we can tune and select PNNs to detect multiple gas targets that are not seen in training on a quantitative basis. We additionally assess our test set detections using interpretability techniques widely employed with ML-based predictors, but less common with detection methods relying on learned latent spaces.

Hyperspectral imaging

Shining a Light on the Nucleus: Photonuclear Measurements from Correlations to Charmonium

The atomic nucleus is comprised of a collection of nucleons (protons and neutrons), which are bound together by the nucleon-nucleon (NN) interaction that originates from Quantum Chromodynamics (QCD). While most nucleons experience the force from the rest of the nu- cleus as a single net “mean-field” interaction that binds them relatively weakly, a small but impactful fraction are in configurations called “Short-Range Correlations” (SRCs), in which they pair with another nucleon at very short distance to experience strong interactions, sig- nificant binding, and high momentum. Hard, high-energy scattering reactions in which an SRC pair is broken apart, knocking both nucleons out of the nucleus, provide the ability to probe the details of these SRC configurations in the nucleus. Previous measurements have had limited statistics and kinematic reach, and the theoretical tools available were in- sufficient to draw quantitative conclusions regarding the ground-state properties of SRCs. The studies described in this thesis represent the first global analysis of SRC breakup mea- surements in order to present a unified picture of SRCs within light- to medium-size nuclei. This includes the use of a novel theoretical framework, the Generalized Contact Formalism, which connects scattering cross-section measurements and the ground-state properties of the SRC pair, to quantitatively interpret a variety of electron-scattering measurements. This is brought to culmination by a report on the first measurement of SRC pairs via the use of hard meson photoproduction reactions, which, despite differing significantly from the me- chanics of electron-scattering, is well-described under a common framework, pointing to a consistent and universal picture of SRCs across reaction channels. I also report on the first measurement of J/¿ photoproduction in the near- and below-threshold kinematic region, giving the first insights to the gluonic structure of bound nucleons in the large-x “valence” region and providing constraints on a gluonic “EMC effect”. In addition to these studies, I provide details on the search for Primakoff production of axion-like particles using the pho- toproduction data taken for this experiment, and I conclude by describing studies of nucleon spin structure measurements that will be performed at the forthcoming U.S. Electron-Ion Collider.

Pybus, Jackson

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES

DESI DR2 results. I. Baryon acoustic oscillations from the Lyman alpha forest

We present the baryon acoustic oscillation (BAO) measurements with the Lyman-𝛼 (Ly⁢𝛼) forest from the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI) survey. Our BAO measurements include both the autocorrelation of the Ly⁢𝛼 forest absorption observed in the spectra of high-redshift quasars and the cross-correlation of the absorption with the quasar positions. The total sample size is approximately a factor of 2 larger than the DR1 dataset, with forest measurements in over 820,000 quasar spectra and the positions of over 1.2 million quasars. We describe several significant improvements to our analysis in this paper, and two supporting papers describe improvements to the synthetic datasets that we use for validation and how we identify damped Ly⁢𝛼 absorbers. Our main result is that we have measured the BAO scale with a statistical precision of 1.1% along and 1.3% transverse to the line of sight, for a combined precision of 0.65% on the isotropic BAO scale at 𝑧 eff =2.33. This excellent precision, combined with recent theoretical studies of the BAO shift due to nonlinear growth, motivated us to include a systematic error term in Ly⁢𝛼 BAO analysis for the first time. We measure the ratios 𝐷 𝐻 ⁡(𝑧 eff )/𝑟 𝑑 = 8.632 ± 0.098 ± 0.026 and 𝐷 𝑀 ⁡(𝑧 eff )/𝑟 𝑑 = 38.99 ± 0.52 ± 0.12, where 𝐷 𝐻 = 𝑐/𝐻⁡(𝑧) is the Hubble distance, 𝐷 𝑀 is the transverse comoving distance, 𝑟 𝑑 is the sound horizon at the drag epoch, and we quote both the statistical and the theoretical systematic uncertainty. The companion paper presents the BAO measurements at lower redshifts from the same dataset and the cosmological interpretation.

baryon acoustic oscillations

Using Separation-Enhanced Isotope Ratio Mass Spectrometry to Enable Increased Renewable Carbon Content in Transportation Fuels (CRADA 525)

Stable isotope ratio measurements of carbon atoms using isotope ratio mass spectrometry (IRMS) can be an effective tool for quantifying biogenic carbon in co-processed fuels, with results approaching the precision and accuracy of accelerator mass spectrometry (AMS). The lower cost of an IRMS may enable deployment to refineries, improving access and analysis turnaround times (≤2 hours), and, by extension, provide data that can allow process optimization to maximize renewable carbon in desired refinery products. This project explored the integration of chemical separation with IRMS analyses to enable highly detailed tracking of biogenic carbon into fuel product streams separated by boiling point range, chemical class, or specific compound. Forty-nine fuels and fuel components of fossil and biogenic origin, spanning gasoline and diesel boiling point ranges, were received from three refiners and were analyzed for their δ 13 C values via IRMS. Results spanned a 13 C range from ca. 10‰ to 44‰ and reflect materials derived from sustainable sources (e.g., C4 or C3 plants, animal-based pathways, syngas) or from fossil-derived fuels. Common ranges are approximately 18‰ to 9‰ and approximately 30‰ to 20‰ for C4 and C3 plants, respectively, and approximately 34‰ to 24‰ and approximately 70‰ to 33‰ for petroleum-derived fuels and methane, respectively. Fuel-like standards were developed and tested using direct-injection elemental analyzer (EA) IRMS for liquid fuels. This method was compared with the published methods, yielding statistically similar results. Four blend curve sets were produced ranging from 0% to 100% of a fuel containing biogenic carbon, focusing on 0% to 10% biogenic carbon. Linear fits were the most applicable for two of the four blend curve sets; however, two sets were found to exhibit slightly quadratic behavior, which was more pronounced in low biogenic blend samples, necessitating second-order fits. The origin of the slight quadratic behavior remains unclear; however, the discussion points to possible interpretations. CanmetENERGY thoroughly characterized a majority of the samples using one- and two-dimensional gas chromatography (GC and GC×GC, respectively) and other analyses. Selected samples were subjected to solid phase extraction (SPE) for saturate, olefin, aromatic, and polar (SOAP) analysis, and the resulting solvent-diluted fractions containing saturates and aromatics were returned to Pacific Northwest National Laboratory (PNNL), where the solvent was removed via evaporation or physical separation using GC techniques. Characterization and separations provided an understanding of saturate and aromatic content, as well as boiling point ranges for each sample and sample fraction. Samples resulting from SPE were examined using EA-IRMS and gas chromatography combustion IRMS (GC-C-IRMS) analyses. Both approaches suggest that the range in values between end-members can be increased by selecting the paraffinic or aromatic fraction of the end-member or by selecting among individual compounds resulting from GC separation of the paraffinic fractions. Considerable work remains to put these approaches into practice and statistically validate the benefit for using a fraction or individual compound over bulk analysis of a sample. However, initial results suggest that separations provide advantages for samples having blend ratios of less than 10% biogenic blendstocks. 13 C results showed statistically similar biofuel blend results to those obtained at PNNL, although additional work is needed to obtain better reproducibility. Select samples were sent to Los Alamos National Laboratory (LANL) for IRMS measurements and Beta Analytics for AMS measurements. This work suggests that IRMS and AMS yield closely comparable results and in some circumstances, IRMS could serve as a surrogate for AMS. While additional work is needed to better resolve statistical advantages for separations and better show the comparable nature of IRMS and AMS in both the biogenic carbon analysis of bulk chemical classes, initial results from this study suggest that these should be pursued in order to proliferate this approach for quantifying biogenic carbon in transportation fuels to the refinery level, thereby potentially enabling process optimization in co-processing scenarios.

09 BIOMASS FUELS

DESI DR2 Results I: Baryon Acoustic Oscillations from the Lyman Alpha Forest

We present the Baryon Acoustic Oscillation (BAO) measurements with the Lyman-alpha (LyA) forest from the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI) survey. Our BAO measurements include both the auto-correlation of the LyA forest absorption observed in the spectra of high-redshift quasars and the cross-correlation of the absorption with the quasar positions. The total sample size is approximately a factor of two larger than the DR1 dataset, with forest measurements in over 820,000 quasar spectra and the positions of over 1.2 million quasars. We describe several significant improvements to our analysis in this paper, and two supporting papers describe improvements to the synthetic datasets that we use for validation and how we identify damped LyA absorbers. Our main result is that we have measured the BAO scale with a statistical precision of 1.1% along and 1.3% transverse to the line of sight, for a combined precision of 0.65% on the isotropic BAO scale at $z_{eff} = 2.33$. This excellent precision, combined with recent theoretical studies of the BAO shift due to nonlinear growth, motivated us to include a systematic error term in LyA BAO analysis for the first time. We measure the ratios $D_H(z_{eff})/r_d = 8.632 \pm 0.098 \pm 0.026$ and $D_M(z_{eff})/r_d = 38.99 \pm 0.52 \pm 0.12$, where $D_H = c/H(z)$ is the Hubble distance, $D_M$ is the transverse comoving distance, $r_d$ is the sound horizon at the drag epoch, and we quote both the statistical and the theoretical systematic uncertainty. The companion paper presents the BAO measurements at lower redshifts from the same dataset and the cosmological interpretation.

79 ASTRONOMY AND ASTROPHYSICS

Interpreting the spatial distribution of soil properties with a physically-based distributed hydrological model

Digital soil maps are commonly data-driven as the development of physically-based models for soil mapping is difficult due to the complexity of soils. However, physically-based hydrologic models have been successful in simulating water dynamics. Since water movement is a major driver of pedogenesis, the physical rules that govern water movement might help explain and predict the spatial variation of soil properties. Here, we demonstrate the novel use of a physically-based, distributed hydrologic model to inform the spatial distribution of soil properties. The Distributed Hydrology Soil Vegetation Model (DHSVM) was utilized to simulate soil moisture content (SM) and water table depth (WTD) in two hillslope catchments under pasture and forest management wherein hydrologic model outputs were then compared with soil properties measured in situ. SM sensors and wells were installed in both catchments to validate simulations of soil water movement via Nash-Sutcliffe Efficiency (E). In-situ observations were made at 87 sites within both catchments to study the connection between simulated water movement (SM and WTD) and observed soil properties, namely the depth and thickness of the argillic (Bt), fragic (Btx), and C horizons, and the depth of redoximorphic features. The simulated time series of SM and WTD were also clustered per season using Dynamic Time Warping (DTW), which identified similarity among time series at varying timescales. Model validation suggested that simulations of surficial SM (0–20 cm) were reasonable (E = 0.45), however, simulated subsurface SM (45–60 cm) and WTD were not sufficiently accurate. The thickness of Btx horizons were spatially grouped into different populations by SM clusters from every season except spring. For the other properties, only SM dynamics of specific seasons grouped into significantly different populations, suggesting that the explanatory power of simulated water movement varies seasonally and was greater during winter. Here, we show clusters of simulated SM separated soil properties into statistically different populations, showing that hydrologic models could inform areas that followed different water dynamics related to pedogenic trajectories and related biogeochemical processes not necessarily simulated by the model. As such, physically-based modeling of water dynamics can, therefore, inform and advance digital soil mapping by linking water movement patterns stemming from hydrologic model outputs to spatial patterns of soil properties and pedogenesis.

54 ENVIRONMENTAL SCIENCES

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling

The role of source geometry and atmospheric propagation in global bolide infrasound detectability

Global infrasound monitoring provides a persistent means of detecting energetic bolide atmospheric entries, complementing optical observations and extending coverage over remote regions. We present a global assessment of the physical factors governing bolide infrasound detectability by correlating 623 bolide events reported by the Center for Near-Earth Object Studies between 2007 and 2025 with waveform data from the International Monitoring System. We identify 311 events with confirmed infrasound detections, corresponding to a detection rate of approximately 50%, substantially higher than inferred from earlier surveys, reflecting both the maturation of the global infrasound network and advances in automated, multi-frequency array processing. Analysis of flight parameters shows that infrasound detectability is selective rather than uniform across the bolide population. Detected events are preferentially associated with steeper entry angles and lower-altitude energy deposition, while shallow, high-altitude trajectories are less consistently observed. Very high-energy events remain detectable regardless of geometry, but for the more common lower-energy regime, observability depends on specific combinations of entry parameters and propagation conditions. This geometric dependence persists across comparable energy ranges and atmospheric conditions, indicating that entry angle exerts a primary control on detectability, with energy and propagation acting as secondary modulating factors. Furthermore, these results provide new physical constraints on bolide-atmosphere interactions and improve interpretation of global infrasound observations for planetary defense and atmospheric-entry studies.

Bolides

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ

Predictive models of the genetic bases underlying budding yeast fitness in multiple environments

Abstract The ability of organisms to adapt and survive depends on the effects of genes and the environment on fitness. However, the multigenic nature of fitness and genotype-by-environment interactions hinder our understanding of the genetic basis of fitness. Here, we established fitness prediction models for 35 environments using machine learning and existing fitness data and different genetic variant types for a Saccharomyces cerevisiae population. Models revealed that the predictive ability of genetic variants varied across environments, with copy number variants explaining the majority of fitness variation in most cases. Model interpretation showed that different variant types identified distinct gene sets associated with predictive variants. These gene sets were significantly enriched in experimentally validated genes affecting fitness in only a subset of environments, indicating that many genes influencing fitness remain unexplored. Notably, non-experimentally validated genes were more important than validated ones for fitness predictions. Gene contributions to predictions were both isolate- and environment-dependent, pointing to gene-by-gene and gene-by-environment interactions. Furthermore, models uncovered experimentally validated and novel candidate genetic interactions for a well-characterized stress, the fungicide benomyl. These findings highlight the feasibility of identifying the genetic basis of fitness by using different genetic variant types and offer novel targets for future functional analysis.

DNA copy number variations

Correct Interpretations of ENDF-102 Definitions for Resonance Effects

My Uncle Willie circa 1600 wrote “What’s in a name; a rose by any other name would smell as sweet.” I fear in this case we have a somewhat similar problem in that we may be using the same word but are not using the same definition; specifically, the word Unresolved. The simplest physics definition as it applies to neutron resonances, is the energy point where we can no longer see/measure ALL – let me repeat that – ALL - of the individual resonances. That seems simple and clear, but the question is: how to represent resonances beyond this point in order to accurately reproduce the effects we have seen in measurements and expect/need to reproduce in our applications. We know there are more, unseen resonances, otherwise we wouldn’t say Unresolved. The ENDF approach is well defined in ENDF-102 and simple: for ENDF data the only way to represent Unresolved data is by using a theoretical model to define the distribution of resonances, including those that are too narrow to measure (i.e., are unresolved). It is important to note that in ENDF this is the one and only Unresolved model, e.g., there is no provision in ENDF to accurately define individually ALL resonances above the Resolved energy range – by ALL here I mean both those that we can measure and those that we cannot individually measure, but that theory and integral measurements tells us are present. An alternative approach, which would appear to be equally valid, would be to include the latest measured data as tabulated energy expendent data extending upwards in energy above the Resolved energy range. In this approach the evaluation would not include an ENDF style Unresolved energy range; it would only include a Resolved resonance region, followed by tabulated higher energy points, representing the resonances that could be measured beyond the Resolved range. But an important point to note: By listing these resonances above the resolved energy one admits that at least some resonances in this energy range are missing as Unresolved; i.e., they are too narrow or overlapping to measure. The purpose of this paper is to illustrate that the later approach, while done with good intentions, and appearing to be valid/adequate in plots, does not meet the need of our engineering applications. Why? As we will see below, of these two possible approaches, only the ENDF use of a model to statistically include the missing, i.e., unresolved, resonances, can meet our engineering needs to reproduce the integral effects we have measured and understand. Only with this statistical model can we predict and include in our calculated results the important effects of temperature (Doppler broadening), and energy integrals (self-shielding). Below I will first present results using two ENDF/B-VIII.1 evaluations, U235 and U238, that use the correct ENDF-102 definition of an Unresolved resonance region, using a statistical model to include the effects of resonances that theory predicts are present, but are too narrow to measure. These two evaluations reproduce the expected temperature (Doppler) and energy integral (self-shielding) effects that we expect. Next I will present results using one ENDF/B-VIII.1 evaluation, 26-Fe-56, that does not use an ENDF-102 Unresolved resonance region; instead above its Resolved energy range it lists many tabulated energy points, that look like measured data, but by definition, since they are included above the ENDF Resolved energy range there are missing Unresolved resonances, i.e., there are missing the resonances that are too narrow to resolve, i.e., are unresolved. My conclusion, and I hope yours, is that the below figures illustrate that this approach does not reproduce the temperature and energy integrals that we expect and need to accurately calculate results for our fission reactor calculations. As such this approach should not be used in ENDF formatted evaluations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Analysis of streaked images of x-ray self-emission in laser-driven spherical implosions

Imaging of x-ray self-emission provides a powerful in situ measurement of the spatial and temporal evolution of high-energy-density plasmas. However, interpretation of these measurements requires detailed understanding of the data-generating process. This work presents a case study in the interpretation of x-ray self-emission data for the specific application of streaked one-dimensional slit imaging of spherical laser-driven implosions. A comprehensive generative model of the streaked slit-imaging diagnostic is developed including detailed treatments of the radiation transfer, photometrics, and photostatistics associated with the measurement. The model is used to generate realistic synthetic streaked images and to analyze experimental streaked images to extract important physical quantities of interest. An example analysis of streaked images from implosion experiments on the OMEGA laser is presented, where the model developed in this work is used to constrain the trajectory and peak velocity of the implosion using Bayesian inference.

Bayesian inference