Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

NASA Orbital Debris Large-Object Baseline Population in ORDEM 3.0

The NASA Orbital Debris Program Office (ODPO) has created and validated high fidelity populations of the debris environment for the latest Orbital Debris Engineering Model (ORDEM 3.0). Though the model includes fluxes of objects 10 um and larger, this paper considers particle fluxes for 1 cm and larger debris objects from low Earth orbit (LEO) through Geosynchronous Transfer Orbit (GTO). These are validated by several reliable radar observations through the Space Surveillance Network (SSN), Haystack, and HAX radars. ORDEM 3.0 populations were designed for the purpose of assisting, debris researchers and sensor developers in planning and testing. This environment includes a background derived from the LEO-to-GEO ENvironment Debris evolutionary model (LEGEND) with a Bayesian rescaling as well as specific events such as the FY-1C anti-satellite test, the Iridium 33/Cosmos 2251 accidental collision, and the Soviet/Russian Radar Ocean Reconnaissance Satellite (RORSAT) sodium-potassium droplet releases. The environment described in this paper is the most realistic orbital debris population larger than 1 cm, to date. We describe derivations of the background population and added specific populations. We present sample validation charts of our 1 cm and larger LEO population against Space Surveillance Network (SSN), Haystack, and HAX radar measurements.

Krisco, Paula H.↗

Measurement of the top-quark pole mass in dileptonic $t\overline{t}$ + 1-jet events at $\sqrt{s}=13$ TeV with the ATLAS experiment

A measurement of the top-quark pole mass $m$$^{pole}_{t}$ is presented in $t\bar{t}$ events with an additional jet, $t\bar{t}$+ 1-jet, produced in pp collisions at $\sqrt{s} = 13 TeV. The data sample, recorded with the ATLAS experiment during Run 2 of the LHC, corresponds to an integrated luminosity of 140 fb −1 . Events with one electron and one muon of opposite electric charge in the final state are selected to measure the $t\bar{t}$ + 1-jet differential cross-section as a function of the inverse of the invariant mass of the $t\bar{t}$ + 1-jet system. Iterative Bayesian Unfolding is used to correct the data to enable comparison with fixed-order calculations at next-to-leading-order accuracy in the strong coupling. The process pp → $t\bar{t}$j(2 → 3), where top quarks are taken as stable particles, and the process pp → $b\bar{b}$l + vl – $\overline{ν}$j (2 → 7), which includes top-quark decays to the dilepton final state and off-shell effects, are considered. The top-quark mass is extracted using a χ 2 fit of the unfolded normalized differential cross-section distribution. The results obtained with the 2 → 3 and 2 → 7 calculations are compatible within theoretical uncertainties, providing an important consistency check.

Hadron-Hadron Scattering↗

Targeted Adaptive Design

Modern advanced manufacturing and advanced materials design often require searches of relatively high-dimensional process control parameter spaces for settings that result in optimal structure, property, and performance parameters. The mapping from the former to the latter must be determined from noisy experiments or from expensive simulations. Here, we abstract this problem to a mathematical framework in which an unknown function from a control space to a design space must be ascertained by means of expensive noisy measurements, which locate control settings generating desired design features within specified tolerances, with quantified uncertainty. We describe targeted adaptive design (TAD), a new algorithm that performs this sampling task efficiently. TAD creates a Gaussian process surrogate model of the unknown mapping at each iterative stage, proposing a new batch of control settings to sample experimentally and optimizing the updated expected log-predictive probability density of the target design. TAD either stops upon locating a solution with uncertainties that fit inside the tolerance box or uses a measure of expected future information to determine that the search space has been exhausted with no solution. TAD thus embodies the exploration-exploitation tension in a manner that recalls, but is essentially different from, Bayesian optimization and optimal experimental design.

97 MATHEMATICS AND COMPUTING↗

Description of Pegethrix niliensis sp. nov., a Novel Cyanobacterium from the Nile River Basin, Egypt: A Polyphasic Analysis and Comparative Study of Related Genera in the Oculatellales Order

In this paper, we examine the filamentous cyanobacterial strain NILCB16 and describe it as a new species within the genus Pegethrix. The original population was sampled from a mat growing in an irrigation canal in the Nile River, Egypt. Initially classified under Plectonema or Planktolyngbya, the strain is a potential producer of the toxins microcystin and β-N-Methylamino-L-Alanine (BMAA). Additionally, we reviewed the taxonomic relationships between the Oculatellales genera. To describe the new species, we conducted a polyphasic study, encompassing 16S rRNA gene phylogenetic analyses performed using both Maximum Likelihood and Bayesian methods, sequence identity (p-distance) analysis, 16S-23S ITS secondary structures, and morphological and habitat comparisons. The phylogenetic analysis revealed that strain NILCB16 clustered within the Pegethrix clade with strong phylogenetic support, but in a distinct position from other species in the genus. The strain shared a maximum 16S rRNA gene identity of 97.3% with P. qiandaoensis and 96.1% with the type species, P. bostrychoides. Morphologically, NILCB16 can be differentiated from other species in the genus by its lack of false branching. Our phylogenetic analyses also show that Pegethrix, Cartusia, Elainella, and Maricoleus are clustered with strong phylogenetic support. They exhibit high 16S rRNA gene identity and are morphologically indistinguishable, suggesting they could potentially be merged into a single genus in the future.

Hentschke, Guilherme Scotta (ORCID:000000034396024↗

Systematic KMTNet Planetary Anomaly Search. I. OGLE-2019-BLG-1053Lb, a BuriedTerrestrial Planet

In order to exhume the buried signatures of “missing planetary caustics” in Korea Microlensing Telescope Network (KMTNet) data, we conducted a systematic anomaly search of the residuals from point-source point-lens fits, based on a modified version of the KMTNet Event Finder algorithm. This search revealed the lowest-mass-ratio planetary caustic to date in the microlensing event OGLE-2019-BLG-1053, for which the planetary signal had not been noticed before. The planetary system has a planet–host mass ratio ofq= (1.25±0.13) × 10−5. A Bayesian analysis yielded estimates of the mass of the host star, Mhost =-0.61+0.29 -0.24 Mo, the mass of its planet, Mplanet =-2.48 +1.19 -0.98 Mo, the projected planet – host separation, a^= 3.4 +0.5/-0.5 au, and the lens distance, DL =-6.8 +0.6 -0.90kpc.The discovery of this very-low-mass-ratio planet illustrates the utility of our method and opens a new window for a large and homogeneous sample to study the microlensing planet–host mass ratio function down to q∼ 10−5.

Exoplanet detection methods↗

DESI Spectroscopy of HETDEX Emission-line Candidates. I. Line Discrimination Validation

The Hobby–Eberly Dark Energy Experiment (HETDEX) is an untargeted spectroscopic galaxy survey that uses Lyα-emitting galaxies (LAEs) as tracers of 1.9 < z < 3.5 large-scale structure. Most detections consist of a single emission line, whose identity is inferred via a Bayesian analysis of ancillary data. To determine the accuracy of these line identifications, HETDEX detections were observed with the Dark Energy Spectroscopic Instrument (DESI). In two DESI pointings, high-confidence spectroscopic redshifts are obtained for 1157 sources, including 982 LAEs. The DESI spectra are used to evaluate the accuracy of the HETDEX object classifications and tune the methodology to achieve the HETDEX science requirement of ≲2% contamination of the LAE sample by low-redshift emission-line galaxies, while still assigning 96% of the true Lyα emission sample with the correct spectroscopic redshift. We compare emission-line measurements between the two experiments assuming a simple Gaussian line fitting model. Fitted values for the central wavelength of the emission line, the measured line flux, and line widths are consistent between the surveys within uncertainties. Derived spectroscopic redshifts, from the two classification pipelines, when both agree as an LAE classification, are consistent to within $\langle$Δz/(1 + z)$\rangle$ = 6.9 × 10 −5 with an rms scatter of 3.3 × 10 −4 . Data are available at https://data.desi.lbl.gov/desi/public/dr1/vac/dr1/hetdex.

79 ASTRONOMY AND ASTROPHYSICS↗

Probabilistic projections of the Amery Ice Shelf catchment, Antarctica, under conditions of high ice-shelf basal melt

Abstract. Antarctica's Lambert Glacier drains about one-sixth of the ice from the East Antarctic Ice Sheet and is considered stable due to the strong buttressing provided by the Amery Ice Shelf. While previous projections of the sea-level contribution from this sector of the ice sheet have predicted significant mass loss only with near-complete removal of the ice shelf, the ocean warming necessary for this was deemed unlikely. Recent climate projections through 2300 indicate that sufficient ocean warming is a distinct possibility after 2100. This work explores the impact of parametric uncertainty on projections of the response of the Lambert–Amery system (hereafter “the Amery sector”) to abrupt ocean warming through Bayesian calibration of a perturbed-parameter ice-sheet model ensemble. We address the computational cost of uncertainty quantification for ice-sheet model projections via statistical emulation, which employs surrogate models for fast and inexpensive parameter space exploration while retaining critical features of the high-fidelity simulations. To this end, we build Gaussian process (GP) emulators from simulations of the Amery sector at a medium resolution (4–20 km mesh) using the Model for Prediction Across Scales (MPAS)-Albany Land Ice (MALI) model. We consider six input parameters that control basal friction, ice stiffness, calving, and ice-shelf basal melting. From these, we generate 200 perturbed input parameter initializations using space filling Sobol sampling. For our end-to-end probabilistic modeling workflow, we first train emulators on the simulation ensemble and then calibrate the input parameters using observations of the mass balance, grounding line movement, and calving front movement with priors assigned via expert knowledge. Next, we use MALI to project a subset of simulations to 2300 using ocean and atmosphere forcings from a climate model for both low- and high-greenhouse-gas-emission scenarios. From these simulation outputs, we build multivariate emulators by combining GP regression with principal component dimension reduction to emulate multivariate sea-level contribution time series data from the MALI simulations. We then use these emulators to propagate uncertainty from model input parameters to predictions of glacier mass loss through 2300, demonstrating that the calibrated posterior distributions have both greater mass loss and reduced variance compared to the uncalibrated prior distributions. Parametric uncertainty is large enough through about 2130 that the two projections under different emission scenarios are indistinguishable from one another. However, after rapid ocean warming in the first half of the 22nd century, the projections become statistically distinct within decades. Overall, this study demonstrates an efficient Bayesian calibration and uncertainty propagation workflow for ice-sheet model projections and identifies the potential for large sea-level rise contributions from the Amery sector of the Antarctic Ice Sheet after 2100 under high-greenhouse-gas-emission scenarios.

54 ENVIRONMENTAL SCIENCES↗

Hazard Assessment from Storm Tides and Rainfall on a Tidal River Estuary

Here, we report on methods and results for a model-based flood hazard assessment we have conducted for the Hudson River from New York City to Troy/Albany at the head of tide. Our recent work showed that neglecting freshwater flows leads to underestimation of peak water levels at up-river sites and neglecting stratification (typical with two-dimensional modeling) leads to underestimation all along the Hudson. As a result, we use a three-dimensional hydrodynamic model and merge streamflows and storm tides from tropical and extratropical cyclones (TCs, ETCs), as well as wet extratropical cyclone (WETC) floods (e.g. freshets, rain-on-snow events). We validate the modeled flood levels and quantify error with comparisons to 76 historical events. A Bayesian statistical method is developed for tropical cyclone streamflows using historical data and consisting in the evaluation of (1) the peak discharge and its pdf as a function of TC characteristics, and (2) the temporal trend of the hydrograph as a function of temporal evolution of the cyclone track, its intensity and the response characteristics of the specific basin. A k-nearest-neighbors method is employed to determine the hydrograph shape. Out of sample validation tests demonstrate the effectiveness of the method. Thus, the combined effects of storm surge and runoff produced by tropical cyclones hitting the New York area can be included in flood hazard assessment. Results for the upper Hudson (Albany) suggest a dominance of WETCs, for the lower Hudson (at New York Harbor) a case where ETCs are dominant for shorter return periods and TCs are more important for longer return periods (over 150 years), and for the middle-Hudson (Poughkeepsie) a mix of all three flood events types is important. However, a possible low-bias for TC flood levels is inferred from a lower importance in the assessment results, versus historical event top-20 lists, and this will be further evaluated as these preliminary methods and results are finalized. Future funded work will quantify the influences of sea level rise and flood adaptation plans (e.g. surge barriers). It would also be valuable to examine how streamflows from tropical cyclones and wet cool-season storms will change, as this factor will dominate at upriver locations.

Hazard assessment↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

Development and Execution of the RUNSAFE Runway Safety Bayesian Belief Network Model

One focus area of the National Aeronautics and Space Administration (NASA) is to improve aviation safety. Runway safety is one such thrust of investigation and research. The two primary components of this runway safety research are in runway incursion (RI) and runway excursion (RE) events. These are adverse ground-based aviation incidents that endanger crew, passengers, aircraft and perhaps other nearby people or property. A runway incursion is the incorrect presence of an aircraft, vehicle or person on the protected area of a surface designated for the landing and take-off of aircraft; one class of RI events simultaneously involves two aircraft, such as one aircraft incorrectly landing on a runway while another aircraft is taking off from the same runway. A runway excursion is an incident involving only a single aircraft defined as a veer-off or overrun off the runway surface. Within the scope of this effort at NASA Langley Research Center (LaRC), generic RI, RE and combined (RI plus RE, or RUNSAFE) event models have each been developed and implemented as a Bayesian Belief Network (BBN). Descriptions of runway safety issues from the literature searches have been used to develop the BBN models. Numerous considerations surrounding the process of developing the event models have been documented in this report. The event models were then thoroughly reviewed by a Subject Matter Expert (SME) panel through multiple knowledge elicitation sessions. Numerous improvements to the model structure (definitions, node names, node states and the connecting link topology) were made by the SME panel. Sample executions of the final RUNSAFE model have been presented herein for baseline and worst-case scenarios. Finally, a parameter sensitivity analysis for a given scenario was performed to show the risk drivers. The NASA and LaRC research in runway safety event modeling through the use of BBN technology is important for several reasons. These include: 1) providing a means to clearly understand the cause and effect patterns leading to safety issues, incidents and accidents, 2) enabling the prioritization of specialty areas needing more attention to improve aviation safety, and 3) enabling the identification of gaps within NASA's Aviation Safety funding portfolio

Green, Lawrence L.↗

Galaxy cluster matter profiles - I. Self-similarity, mass calibration, and observable-mass relation validation employing cluster mass posteriors

We present a study of the weak lensing inferred matter profiles ΔΣ(R) of 698 South Pole Telescope (SPT) thermal Sunyaev-Zel’dovich effect (tSZE) selected and MCMF optically confirmed galaxy clusters in the redshift range 0.25 < z < 0.94 that have associated weak gravitational lensing shear profiles from the Dark Energy Survey (DES). Rescaling these profiles to account for the mass dependent size and the redshift dependent density produces average rescaled matter profiles ΔΣ(R/R200c)/(ρcritR200c) with a lower dispersion than the unscaled ΔΣ(R) versions, indicating a significant degree of self-similarity. Galaxy clusters from hydrodynamical simulations also exhibit matter profiles that suggest a high degree of self-similarity, with RMS variation among the average rescaled matter profiles with redshift and mass falling by a factor of approximately six and 23, respectively, compared to the unscaled average matter profiles. We employed this regularity in a new Bayesian method for weak lensing mass calibration that employs the so-called cluster mass posterior P(M200|ζ̂, λ̂, z), which describes the individual cluster masses given their tSZE (ζ̂) and optical (λ̂, z) observables. This method enables simultaneous constraints on richness λ-mass and tSZE detection significance ζ-mass relations using average rescaled cluster matter profiles. We validated the method using realistic mock datasets and present observable-mass relation constraints for the SPT×DES sample, where we constrained the amplitude, mass trend, redshift trend, and intrinsic scatter. Our observable-mass relation results are in agreement with the mass calibration derived from the recent cosmological analysis of the SPT×DES data based on a cluster-by-cluster lensing calibration. Our new mass calibration technique offers a higher efficiency when compared to the single cluster calibration technique. We present new validation tests of the observable-mass relation that indicate the underlying power-law form and scatter are adequate to describe the real cluster sample but that also suggest a redshift variation in the intrinsic scatter of the λ-mass relation may offer a better description. In addition, the average rescaled matter profiles offer high signal-to-noise ratio (S/N) constraints on the shape of real cluster matter profiles, which are in good agreement with available hydrodynamical ΛCDM simulations. This high S/N profile contains information about baryon feedback, the collisional nature of dark matter, and potential deviations from general relativity.Key words: gravitational lensing: weak / galaxies: clusters: general / large-scale structure of Universe

79 ASTRONOMY AND ASTROPHYSICS↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

Atacama Cosmology Telescope measurements of a large sample of candidates from the Massive and Distant Clusters of WISE Survey: Sunyaev-Zeldovich effect confirmation of MaDCoWS candidates using ACT

Context. Galaxy clusters are an important tool for cosmology, and their detection and characterization are key goals for current and future surveys. Using data from the Wide-field Infrared Survey Explorer (WISE), the Massive and Distant Clusters of WISE Survey (MaDCoWS) located 2839 significant galaxy overdensities at redshifts 0.7 . z . 1.5, which included extensive follow-up imaging from the Spitzer Space Telescope to determine cluster richnesses. Concurrently, the Atacama Cosmology Telescope (ACT) has produced large area millimeter-wave maps in three frequency bands along with a large catalog of Sunyaev-Zeldovich (SZ)-selected clusters as part of its Data Release 5 (DR5). Aims. We aim to verify and characterize MaDCoWS clusters using measurements of, or limits on, their thermal SZ effect signatures. We also use these detections to establish the scaling relation between SZ mass and the MaDCoWS-defined richness. Methods. Using the maps and cluster catalog from DR5, we explore the scaling between SZ mass and cluster richness. We do this by comparing cataloged detections and extracting individual and stacked SZ signals from the MaDCoWS cluster locations. We use complementary radio survey data from the Very Large Array, submillimeter data from Herschel, and ACT 224 GHz data to assess the impact of contaminating sources on the SZ signals from both ACT and MaDCoWS clusters. We use a hierarchical Bayesian model to fit the mass-richness scaling relation, allowing for clusters to be drawn from two populations: one, a Gaussian centered on the mass-richness relation, and the other, a Gaussian centered on zero SZ signal. Results. We find that MaDCoWS clusters have submillimeter contamination that is consistent with a gray-body spectrum, while the ACT clusters are consistent with no submillimeter emission on average. Additionally, the intrinsic radio intensities of ACT clusters are lower than those of MaDCoWS clusters, even when the ACT clusters are restricted to the same redshift range as the MaDCoWS clusters. We find the best-fit ACT SZ mass versus MaDCoWS richness scaling relation has a slope of p1 = 1.84+0.15 −0.14, where the slope is defined as M ∝ λ p1 15 and λ15 is the richness. We also find that the ACT SZ signals for a significant fraction (∼57%) of the MaDCoWS sample can statistically be described as being drawn from a noise-like distribution, indicating that the candidates are possibly dominated by low-mass and unvirialized systems that are below the mass limit of the ACT sample. Further, we note that a large portion of the optically confirmed ACT clusters located in the same volume of the sky as MaDCoWS are not selected by MaDCoWS, indicating that the MaDCoWS sample is not complete with respect to SZ selection. Finally, we find that the radio loud fraction of MaDCoWS clusters increases with richness, while we find no evidence that the submillimeter emission of the MaDCoWS clusters evolves with richness. Conclusions. We conclude that the original MaDCoWS selection function is not well defined and, as such, reiterate the MaDCoWS collaboration’s recommendation that the sample is suited for probing cluster and galaxy evolution, but not cosmological analyses. We find a best-fit mass-richness relation slope that agrees with the published MaDCoWS preliminary results. Additionally, we find that while the approximate level of infill of the ACT and MaDCoWS cluster SZ signals (1–2%) is subdominant to other sources of uncertainty for current generation experiments, characterizing and removing this bias will be critical for next-generation experiments hoping to constrain cluster masses at the sub-percent level.

large↗

Infusing Statistical Thinking into the NASA Quesst Community Test Campaign

Statistical thinking permeates many important decisions as NASA plans its Quesst mission, which will culminate in a series of community overflights using the X-59 aircraft to demonstrate low-noise supersonic flight. Month-long longitudinal surveys will be deployed to assess human perception and annoyance to this new acoustic phenomenon. NASA works with a large contractor team to develop systems and methodologies to estimate noise doses, to test and field socio-acoustic surveys, and to study the relationship between the two quantities, dose and response, through appropriate choices of statistical models. This latter dose-response relationship will serve as an important tool as national and international noise regulators debate whether overland supersonic flights could be permitted once again within permissible noise limits. In this presentation we highlight several areas where statistical thinking has come into play, including issues of sampling, classification and data fusion, and analysis of longitudinal survey data that are subject to rare events and the consequences of measurement error. We note several operational constraints that shape the appeal or feasibility of some decisions on statistical approaches, and we identify several important remaining questions to be addressed.

Bayesian model↗

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

$\mathrm{SageNet}$: Fast Neural Network Emulation of the Stiff-amplified Gravitational Waves from Inflation

Accurate modeling of the inflationary gravitational waves (GWs) requires time-consuming, iterative numerical integrations of differential equations to take into account their backreaction on the expansion history. To improve computational efficiency while preserving accuracy, we present the Stiff-amplified Gravitational-wave Emulator Network (SageNet), a deep learning framework designed to replace conventional numerical solvers (code available at https://github.com/YifangLuo/SageNet). SageNet employs a long short-term memory architecture to emulate the present-day energy density spectrum of the inflationary GWs with possible stiff amplification, Ω GW (f). Trained on a data set of 25,689 numerically generated solutions, SageNet allows accurate reconstructions of Ω GW (f) and generalizes well to a wide range of cosmological parameters; 90.9% of the test emulations with randomly distributed parameters exhibit errors of under 4%. In addition, SageNet demonstrates its ability to learn and reproduce the artificial, adaptive sampling patterns in numerical calculations, which implement denser sampling of frequencies around changes in spectral indices in Ω GW (f). The dual capability of learning both physical and artificial features of the numerical GW spectra establishes SageNet as a robust alternative to exact numerical methods. Finally, our benchmark tests show that SageNet reduces the computation time from tens of seconds to milliseconds, achieving a speedup of ∼10 4 times over standard CPU-based numerical solvers with the potential for further acceleration on GPU hardware. These capabilities make SageNet a powerful tool for accelerating Bayesian inference procedures for extended cosmological models. In a broad sense, the SageNet framework offers a fast, accurate, and generalizable solution to modeling cosmological observables whose theoretical predictions demand costly differential equation solvers.

Astronomy data modeling↗

Uncertainty quantification of a physics-informed model based on sparse identification of a Thermal Energy Distribution System

Integrated energy systems (IES)s are crucial for enhancing the economy and efficiency of power generation sources (e.g., nuclear energy) necessary to unleash American energy dominance. These systems can be integrated with thermal energy storage (TES) and intermittent renewable energies to optimize overall energy use, peak-load regulation, and demand-side responses. However, the stabilization of energy generation, transport, and utilization introduces operational complexities that exceed the challenges of managing each sub-component individually. Currently, though IESs rely on human operators for efficiency and stability, reducing human error risk and enhancing performance through automation is highly desirable. Recent advances at Idaho National Laboratory have demonstrated successful control of the Thermal Energy Distributed System (TEDS). However, the automatic control system depends on a deterministic Sparse Identification of Nonlinear Dynamics with Control (SINDyC) model, which are trained based on simulation data from physics-based simulations. Because of uncertainties in physics-based simulation, SINDyC model results in large discrepancies against experimental data and cannot be reliably used in automatic control. In this paper, we present an innovative approach to address these discrepancies by quantifying uncertainties and developing a more robust model. We first generated trajectories by using first-principles physics codes to encapsulate the experiment. Next, we trained thousands of models by randomly sampling these trajectories. We then collapsed all those models into one probabilistic SINDyC by fitting a multivariate Gaussian distribution onto the resulting coefficient’s distribution. Despite its simplicity, our approach successfully produced 95% confidence intervals that captured the experimental trajectories. It even did so with a higher probability and better U-pooling score across six of the seven relevant quantities of interest (QoIs), as compared to other classical approaches. In conclusion, ongoing research is focusing on generating new experimental trajectories to validate this approach, and on employing Bayesian calibration to refine parametric uncertainties and guide future model development efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗