Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data processing methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

Transforming Agricultural Productivity with AI-Driven Forecasting: Innovations in Food Security and Supply Chain Optimization

Global food security is under significant threat from climate change, population growth, and resource scarcity. This review examines how advanced AI-driven forecasting models, including machine learning (ML), deep learning (DL), and time-series forecasting models like SARIMA/ARIMA, are transforming regional agricultural practices and food supply chains. Through the integration of Internet of Things (IoT), remote sensing, and blockchain technologies, these models facilitate the real-time monitoring of crop growth, resource allocation, and market dynamics, enhancing decision making and sustainability. The study adopts a mixed-methods approach, including systematic literature analysis and regional case studies. Highlights include AI-driven yield forecasting in European hydroponic systems and resource optimization in southeast Asian aquaponics, showcasing localized efficiency gains. Furthermore, AI applications in food processing, such as plasma, ozone and Pulsed Electric Field (PEF) treatments, are shown to improve food preservation and reduce spoilage. Key challenges—such as data quality, model scalability, and prediction accuracy—are discussed, particularly in the context of data-poor environments, limiting broader model applicability. The paper concludes by outlining future directions, emphasizing context-specific AI implementations, the need for public–private collaboration, and policy interventions to enhance scalability and adoption in food security contexts.

99 GENERAL AND MISCELLANEOUS↗

Diagnostics of Magnetohydrodynamic Modes in the Interstellar Medium through Synchrotron Polarization Statistics

One of the biggest challenges in understanding magnetohydrodynamic (MHD) turbulence is identifying the plasma mode components from observational data. Previous studies on synchrotron polarization from the interstellar medium (ISM) suggest that the dominant MHD modes can be identified via statistics of Stokes parameters, which would be crucial for studying various ISM processes such as the scattering and acceleration of cosmic rays, star formation, and dynamo. In this paper, we present a numerical study of the synchrotron polarization analysis (SPA) method through systematic investigation of the statistical properties of the Stokes parameters. We derive the theoretical basis for our method from the fundamental statistics of MHD turbulence, recognizing that the projection of the MHD modes allows us to identify the modes dominating the energy fraction from synchrotron observations. Based on the discovery, we revise the SPA method using synthetic synchrotron polarization observations obtained from 3D ideal MHD simulations with a wide range of plasma parameters and driving mechanisms, and present a modified recipe for mode identification. We propose a classification criterion based on a new SPA+ fitting procedure, which allows us to distinguish between Alfvén mode and compressible/slow mode dominated turbulence. We further propose a new method to identify fast modes by analyzing the asymmetry of the SPA+ signature and establish a new asymmetry parameter to detect the presence of fast mode turbulence. Additionally, we confirm through numerical tests that the identification of the compressible and fast modes is not affected by Faraday rotation in both the emitting plasma and the foreground.

97 MATHEMATICS AND COMPUTING↗

Design Optimization of a Criticality Experiment for the Molten Chloride Reactor Experiment Facility

Neutronics simulations of Molten Chloride Fast Reactors have quantifiable biases that arise from nuclear data, modeling choices, or numerical methods. The multiphysics nature of molten salt reactors makes it challenging to disentangle neutronics modeling biases from biases originating from other physical phenomena. In comparison to a mock-up reactor, criticality experiments can specifically assess the neutronics modeling bias while limiting multiphysics effects. The criticality experiment must be neutronically representative of the full-scale reactor to be valuable. Here, in this paper, we describe the design of a criticality experiment to validate only the neutronics of TerraPower’s Molten Chloride Reactor Experiment (MCRE) and its criticality safety upset scenarios. The proposed experiment uses different chlorine-containing materials to maximize its similarity to the MCRE. The design process uses a constrained Bayesian optimization algorithm to investigate different objective functions that use covariance information for 35 Cl nuclear data. The experiments could reduce the nuclear data–induced uncertainty in k eff of the MCRE from 2161 to 886 pcm. They would also increase the upper subcritical limit of the MCRE criticality safety upset scenario from 0.94101 to 0.94476 when using the WHISPER analysis framework.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Data-driven Community-centered Resilient Assessment and Planning Toolkit for Nexus of Energy and Water (DCRAPT-NEW)

Urban areas, including Detroit and Pittsburgh, have suffered significant dual outages of the electrical and water infrastructure in the past decade due, in part, to the increasing number of extreme weather events. With increasing temperatures and rainfall intensity, these regions need to prepare for increasing extreme events through community-based energy and water resilience analysis, planning, and enhancement. This project developed a suite of open-source, open-access, community-centered, data-driven assessment and distributed energy resource (DER) and planning tools for energy and water resilience enhancement in urban areas. Through establishing a multi-level community awareness and engagement mechanism and a comprehensive collection of power outage and flooding data, an innovative group of community energy and water resilience assessment and planning tools have been developed for a wide range of users with differing and variable sets of data available to them. The developed tools include (1) DOE EAGLE-I data-driven, deep-learning assisted resilience assessment and DER planning tools at the county level with socioeconomic factors incorporated; (2) Utility annual power outage data-driven tools for long term resilience assessment and DER planning and 15-min power outage data-driven tools for short term resilience assessment and planning; (3) Detailed engineering tools for energy and water systems resilience assessment and planning when the system topology and component fragility curves are available; (4) Alternative Resiliency Metric Calculation that extracts and separates outage and restoration processes; and (5) Co-optimization tools that evaluate the resilience of the power and sewage system and allow users to conduct joint planning with energy and wastewater systems. The developed tools provide planners, decision-makers, and stakeholders with powerful capabilities to systematically evaluate system/community resilience and optimal and actionable guidance for enhancing resilience while prioritizing DER investments. The tools have been used and validated in Detroit and Pittsburgh and can be used in other areas of the nation. In addition, this project will (1) advance the knowledge and applications of machine-learning methods in analyzing and fusing different layers of information and generating meaningful data points such as generating rare weather events; (2) significantly improve the energy and water resilience of the identified communities in Detroit and Pittsburgh and prepare for more frequent and severe weather conditions; (3) help communities assess extreme weather event impacts and address short-term and long-term resilience-related issues The developed tools have been made public via GitHub and demonstrated to community stakeholders and utility companies via the two annual workshops and numerous community engagement meetings. The project outcomes are also disseminated through publications in various journals and conference proceedings, and presentations at top conferences.

13 HYDRO ENERGY↗

Initial Uncertainty Analysis of Carbon Tetrachloride Contamination and Remediation in the Ringold A and Lower Mud Units at the Central Plateau

The long-term effectiveness of groundwater cleanup at the Hanford Site Central Plateau depends on predictive models that can capture key uncertainties in contaminant fate and transport. Carbon tetrachloride (CCl 4 ), a persistent and toxic compound, presents particular challenges due to variability in degradation rates, uncertainty in initial plume distribution, and subsurface heterogeneity. These uncertainties directly influence plume persistence, migration pathways, and remedy performance, and thus must be systematically evaluated to support long-term remediation planning. To address these gaps, a large-scale Monte Carlo analysis was conducted using the Plateau to River (P2R) model framework. The modeling approach parameterized three primary uncertainty factors: (1) degradation rate, (2) initial plume distribution, and (3) hydraulic conductivity. Degradation was represented as a first-order process, with half-lives ranging from 70 to 700 years. Initial plume distributions were created using a geostatistical simulation method (sgsim), which generates many equally plausible versions of how contaminants might be distributed underground. From this, 100 different scenarios were mapped onto the P2R grid. Variability in hydraulic conductivity was represented in a similar way, with 100 scenarios each for the Ringold Lower Mud and Ringold A units (layers 6 and 7), based on fitted exponential variograms and conditioned to well data. In total, more than 1000 realizations were simulated to assess plume behavior under uncertainty. Results demonstrate that degradation kinetics exert the strongest control over plume persistence: Shorter half-lives produced rapid mass reduction, while longer half-lives yielded persistent plumes with limited attenuation. A nonlinear response was observed, with steep mass reductions at half-lives greater than 200 years and near-linear declines beyond this threshold, reflecting interactions between degradation and pumping. The initial plume distribution strongly influenced early transport patterns, with broader sources generating larger plume footprints, although pump-and-treat operations constrained plume migration to managed areas. By comparison, hydraulic conductivity variability in the Ringold units had only a secondary influence, modifying spreading behavior without altering the dominant migration pathways governed by source configuration and hydraulic controls. Overall, the analysis highlights that uncertainty in degradation rate and initial plume configuration are the primary drivers of variability in plume predictions, while conductivity heterogeneity plays a limited role. These findings underscore the need for improved site-specific data on degradation processes and source characterization to enhance the reliability of long-term performance assessments and to better inform remedial decision-making at the Central Plateau.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Rhenium Isotope Reconnaissance of Uranium Ore Concentrates

Exploration of natural isotopic variations of the element rhenium (Re) is in its infancy, with initial studies revealing isotopic fractionation in a variety of geological materials. Here, in this work, we investigate Re isotope variation as a new geochemical tool, given its redox-sensitive properties and affinity for organic matter and sulfides. In this work, Re abundance and isotope ratio data were collected from uranium ore concentrates (UOCs) across a variety of depositional ages, locations, geologic settings, and deposit types. Ore types from which the UOC were derived include sandstone, unconformity, and quartz-pebble (QP) conglomerate. To isolate Re from the U-rich matrix of UOCs, a new purification method utilizing DGA ion exchange resin was developed. We found that UOCs exhibit a wide range of Re isotope ratios, with sandstone ore-derived UOCs having the isotopically lightest values, QP conglomerate ore-derived UOCs having the heaviest, and unconformity ore-derived UOCs in between (with some overlap with sandstone UOCs). The Re isotope ratio range observed in UOCs extends previously reported values by more than a factor of two. Industrial processing (e.g., incomplete recovery of Re from ore, contamination, fractionation during processing) may play a role in the isotopic variability in the UOCs. However, systematic differences between ore types suggest that the depositional setting is a significant factor. For nuclear forensic investigations, Re isotopic compositions combined with data from other isotopic systems provide geochemical signatures that can aid in provenance assessment of UOCs. Regardless of the specific causes for the wide range of Re isotope ratios in UOCs, these initial data indicate Re is a promising tool for nuclear forensic investigations on samples from early in the nuclear fuel cycle.

58 GEOSCIENCES↗

Exploring Uncertainty in Moment Estimation for Small Earthquakes in Southern Nevada Using the Coda Envelope Method

Compiling source parameter estimates for small earthquakes is important both for our understanding of earthquake physics and for accurately assessing earthquake hazard. Reliable source parameter estimates are difficult to achieve for small earthquakes, in part due to our inability to accurately model the relevant physical processes at high frequencies. The coda envelope methodology developed by Mayeda and Walter (1996) and Mayeda et al. (2003) can mitigate this concern and estimate the moment of small earthquakes by determining the parameters that control the shape of the S-wave coda envelope while eliminating path effects by minimizing the scatter between seismic stations. Here, we use an open-source implementation of this technique called the Coda Calibration Tool (CCT; Barno, 2017) to calculate CCT-based moment magnitude estimates of small earthquakes (M L 0–3) in the Rock Valley, Nevada, region within the Nevada National Security Site. The Rock Valley data set is of particular interest because it allows us to explore the changes in uncertainties of the coda calibration method with earthquake size and depth. We found that a consistent linear relationship exists between the local magnitude M L and our coda-derived M w estimates for earthquakes as small as M L 0–3, but that current CCT workflows do not accurately characterize very shallow events. We also demonstrate that the epistemic uncertainty in the apparent stress value assumed by the CCT algorithm can influence magnitude estimates of small earthquakes. In conclusion, these results provide valuable insight into the seismicity of this region, and inform future analysis and modeling efforts for nuclear monitoring and seismic hazard.

58 GEOSCIENCES↗

Over three decades, and counting, of near-surface turbulent flux measurements from the Atmospheric Radiation Measurement (ARM) user facility

Processes mediating the coupling of terrestrial, aquatic, biospheric, and atmospheric systems influence weather, climate, and ecosystem dynamics via transfer of energy, momentum, water, and carbon (or other species). These exchange processes are quantified by measurements of near-surface turbulent fluxes. Understanding processes at these interfaces provides insight toward understanding and predicting current and future states within the Earth system. The Atmospheric Radiation Measurement (ARM) user facility has been conducting measurements of near-surface turbulent fluxes since the early 1990s at long-term fixed locations and shorter-term mobile deployments across the Earth. ARM has utilized two established methods for conducting these measurements: energy balance Bowen ratio (EBBR) and eddy covariance (EC). Primary measurements from the former include sensible and latent heat flux, while the latter also measures fluxes of momentum and carbon (primarily carbon dioxide, with methane fluxes measured at two locations to date). The EBBR systems have been deployed at 22 locations, and, to date, the EC systems have been deployed at over 50 sites, with plans for additional novel site locations in the future. Herein, the history, evolution, and key aspects of these instrument systems are documented, along with information on data quality assurance and post-processing, as well as best use practices. Additionally, three data validation experiments were recently conducted, and their key findings are summarized. Finally, ancillary datasets acquired by ARM, which can contextualize and aid interpretation of the near-surface turbulent flux measurements, are discussed. The datasets described herein include the eddy correlation flux measurement system: 30ECOR (https://doi.org/10.5439/1879993, Sullivan et al., 1997), 30QCECOR (https://doi.org/10.5439/1097546, Gaustad, 2003), ECORSF (https://doi.org/10.5439/1494128, Sullivan et al., 2019a), and associated AmeriFlux and Methane Value-Added Product, AMCMETHANE (https://doi.org/10.5439/1508268, Billesbach, 2011); the energy balance Bowen ratio system: 30EBBR (https://doi.org/10.5439/1023895, Sullivan et al., 1993) and 30BAEBBR (https://doi.org/10.5439/1027268, Gaustad and Xie, 1993); and the carbon dioxide flux measurement system: CO2FLX (https://doi.org/10.5439/1287574, https://doi.org/10.5439/1287575, https://doi.org/10.5439/1287576, Koontz et al., 2015a, b, c; https://doi.org/10.5439/1989774, https://doi.org/10.5439/1989776, https://doi.org/10.5439/1992202, Biraud and Chan, 2002a, b, c). These data can be found by searching the above data stream names at https://adc.arm.gov/discovery/#/results/ (last access: 8 September 2025).

Sullivan, Ryan C. [Argonne National Laboratory (AN↗

Optimizing bioenergy biofuel harvest: a comparative analysis of stepwise and integrated methods for economic and environmental sustainability

Switchgrass is a promising bioenergy feedstock due to its high biomass yield potential, adaptability to marginal lands, and low carbon intensity for feedstock production. However, accurate cost estimation and assessment of greenhouse gas (GHG) emissions for the energy-intensive harvesting process are essential for evaluating the sustainability of bioenergy. This study provides a comparative analysis of two harvesting methods: the Stepwise Method, which separates operations into multiple stages, and the Integrated Method, which combines mowing and raking into a single pass. The analysis was conducted under four scenarios based on field sizes and biomass yields. Using three years of field-scale switchgrass harvest data from 125 sites, GHG emissions, energy consumption, and harvesting costs were quantified using the GREET model and techno-economic analysis. Additionally, regression analysis identified key climate and operational factors affecting fuel consumption. The Stepwise method was the most cost-effective for large fields with high biomass yield, achieving the lowest harvesting costs ($37.70 per ton). In contrast, the Integrated Method performed better in small fields and low-yield conditions, reducing GHG emissions by 9 % and energy use by 5 %. Regression analysis confirmed that a larger field size reduced fuel consumption, while higher biomass yield and longer operational time increased fuel use. Maximum temperature also contributed to a slight increase in fuel consumption. Furthermore, these results provide actionable insights for optimizing harvesting strategies based on field-specific conditions and operational goals, contributing to the economic and environmental sustainability of bioenergy production.

60 APPLIED LIFE SCIENCES↗

Targeted materials discovery using Bayesian algorithm execution

Rapid discovery and synthesis of future materials requires intelligent data acquisition strategies to navigate large design spaces. A popular strategy is Bayesian optimization, which aims to find candidates that maximize material properties; however, materials design often requires finding specific subsets of the design space which meet more complex or specialized goals. We present a framework that captures experimental goals through straightforward user-defined filtering algorithms. These algorithms are automatically translated into one of three intelligent, parameter-free, sequential data collection strategies (SwitchBAX, InfoBAX, and MeanBAX), bypassing the time-consuming and difficult process of task-specific acquisition function design. Our framework is tailored for typical discrete search spaces involving multiple measured physical properties and short time-horizon decision making. We demonstrate this approach on datasets for TiO 2 nanoparticle synthesis and magnetic materials characterization, and show that our methods are significantly more efficient than state-of-the-art approaches. Overall, our framework provides a practical solution for navigating the complexities of materials design, and helps lay groundwork for the accelerated development of advanced materials.

42 ENGINEERING↗

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Sparse measurement medical CT reconstruction using multi-fused block matching denoising priors

A major challenge for medical X-ray CT imaging is reducing the number of X-ray projections to lower radiation dosage and reduce scan times without compromising image quality. However these under-determined inverse imaging problems rely on the formulation of an expressive prior model to constrain the solution space while remaining computationally tractable. Traditional analytical reconstruction methods like Filtered Back Projection (FBP) often fail with sparse measurements, producing artifacts due to their reliance on the Shannon-Nyquist Sampling Theorem. Consensus Equilibrium, which is a generalization of Plug and Play, is a recent advancement in Model-Based Iterative Reconstruction (MBIR), has facilitated the use of multiple denoisers are prior models in an optimization free framework to capture complex, non-linear prior information. However, 3D prior modelling in a Plug and Play approach for volumetric image reconstruction requires long processing time due to high computing requirement. Instead of directly using a 3D prior, this work proposes a BM3D Multi Slice Fusion (BM3D-MSF) prior that uses multiple 2D image denoisers fused to act as a fully 3D prior model in Plug and Play reconstruction approach. Our approach does not require training and are thus able to circumvent ethical issues related with patient training data and are readily deployable in varying noise and measurement sparsity levels. In addition, reconstruction with the BM3D-MSF prior achieves similar reconstruction image quality as fully 3D image priors, but with significantly reduced computational complexity. We test our method on clinical CT data and demonstrate that our approach improves reconstructed image quality.

Hossain, Maliha [ORNL]↗

Prediction of Redox Potentials for the Late Actinides Cm to Lr Using Electronic Structure Methods

Our previously developed computational method for calculating the aqueous redox potentials of the early actinides has been extended to the later elements in the actinide series: Cm, Bk, Cf, Es, Fm, Md, No, and Lr in multiple oxidation states. These calculations were performed using density functional theory with small-core pseudopotentials and their associated basis sets. Solvation effects were considered via a supermolecule-continuum approach, with 30 water molecules representing two solvation shells. Both the COSMO and SMD implicit solvation models were utilized. The structural parameters and hydration numbers for Cm(III), Bk(III), Bk(IV), and Cf(III) are in reasonable agreement with the available experimental data. For redox processes involving atomic cations in solution, the B3LYP/COSMO approach predicted redox potentials to within ±0.2 V of experiment for most redox couples, consistent with our prior work. Inclusion of spin-orbit corrections in specific redox pairs, especially those with the later actinides in high oxidation states, yields improved results relative to calculations including only scalar-relativistic corrections. The An +m /An(0) redox potentials were calculated using a Born-Haber cycle incorporating sublimation, ionization, and hydration energies. Due to a lack of experimental data, three sets of ionization energies were used for the Born-Haber cycle. The calculated An(III/0) potentials showed better agreement with experimental data when using the COSMO solvation model and the test set comprising the NIST recommended ionization energies. Furthermore, the Md(II/0) potential was better described with the SMD model, whereas No(II/0) was not well described by all methods. Finally, the computational approach was able to predict redox potentials that for most cases agreed with the current available experimental or estimated data.

Actinides↗

Beyond Optimization: Exploring Novelty Discovery in Autonomous Experiments

Autonomous experiments (AEs) are transforming how scientific research is conducted by integrating artificial intelligence with automated experimental platforms. Current AEs primarily focus on the optimization of a predefined target; while accelerating this goal, such an approach limits the discovery of unexpected or unknown physical phenomena. Here, we introduce a novel framework, INS 2 ANE (Integrated Novelty Score−Strategic Autonomous Non-Smooth Exploration), to enhance the discovery of novel phenomena in autonomous microscopy experimentation. Our method integrates two key components: (1) a novelty scoring system that evaluates the uniqueness of experimental results and (2) a strategic sampling mechanism that promotes exploration of under-sampled regions even if they appear less promising by conventional criteria. We validate this approach on a preacquired data set with a known ground truth comprising of image−spectral pairs. We further implement the process on autonomous scanning probe microscopy experiments. INS 2 ANE significantly increases the diversity of explored phenomena in comparison to conventional optimization routines, enhancing the likelihood of discovering previously unobserved phenomena. These results demonstrate the potential for autonomous microscopy experiments to enhance the scientific discovery by navigating complex experimental spaces to uncover novel phenomena.

Materials↗

Perspectives on Systematic Cloud Microphysics Scheme Development With Machine Learning

Cloud microphysics—the collection of processes that govern the small‐scale formation, evolution, and interactions of liquid droplets and ice crystals in clouds and precipitation—remains a major source of uncertainty in weather and climate models. Although too small in scale to be explicitly resolved in any large‐eddy simulation, weather, or climate model, the representation of cloud microphysical processes has significant impact at the climate scale. Current microphysical schemes are limited by both parametric uncertainty, linked to uncertainty in physical parameter values, and structural uncertainty, arising from incomplete physical understanding of the processes at play or approximations made for computational efficiency. Recent advances in the application of machine learning (ML) to the physical sciences show significant potential for minimizing these limitations by leveraging high‐fidelity simulations and observations. Here we outline the challenges that must be addressed to apply ML toward cloud microphysics scheme development. This perspectives paper synthesizes recent progress in using data‐driven methods, including ML, to improve cloud microphysics parameterizations and highlights opportunities to address key uncertainties. We discuss the roles of aleatoric (irreducible, or statistical) and epistemic (reducible, or systematic) errors in contributing to microphysics parameterization uncertainty. ML can leverage observations to improve microphysical schemes via bottom‐up and top‐down constraints. Methods such as differentiable programming and ML‐enhanced sampling strategies and the creation of large scale benchmark data sets promise to bridge the gap between observations and models and to improve the consistency of cloud microphysical representation across temporal and spatial scales.

Lamb, Kara D. [Columbia Univ., New York, NY (Unite↗

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology↗

Practical and Optimal Sequential Bayesian Experimental Design for Complex Systems Incorporating Human Experimenter Preferences (Final Scientific/Technical Report)

Experiments are indispensable for developing models of complex systems. Carefully designed experiments can provide substantial savings for these expensive data-acquisition opportunities. However, designs based on heuristics are often suboptimal for systems with multiphysics, nonlinear dynamics, and uncertain and noisy environments. Optimal experimental design, while leveraging predictive models, seeks to systematically quantify and maximize the value of experiments. In this project, we focused on the design of multiple experiments, where current approaches are largely suboptimal: batch-design does not adapt to new data acquired during the experiment campaign (no feedback), and greedy/myopic design ignores future dynamics and consequences (no lookahead). We developed the mathematical framework and computational methods for sequential optimal experimental design (sOED) for complex systems. We enabled tractable model-based sOED in a rigorous manner through novel algorithms based on reinforcement learning, and investigated the effects of human experimenters on the design process. Our methods are fully Bayesian, able to quantify and update uncertainty in a principled manner. The traits aimed by our approach—mathematical rigor and optimality, human effects and uncertainty quantification, computational practicality—are crucial for elevating the standards of artificial intelligence (AI) to support decision-making in scientific domains, and contribute toward trust and realistic adoption of AI in experimental design practice.

97 MATHEMATICS AND COMPUTING↗