Search NASA⌕ Search

SEARCH · Search NASA

Results for “model weights”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Bayesian Optimization of Catalysis with In-Context Learning

Large language models (LLMs) can perform accurate classification with zero or few examples through in-context learning (ICL), allowing the model to observe query-relevant examples at inference time and eliminating the need for additional weight updates to generalize beyond its original training data. We extend this capability to regression with uncertainty estimation using frozen LLMs (e.g., GPT-4o, Gemini), enabling Bayesian optimization (BO) in natural language without explicit model training or feature engineering. We apply this to materials discovery by representing materials as synthesis and testing procedures for use in natural language prompts. This Bayesian, design-first approach prioritizes optimization toward target material properties before detailed characterization, in contrast to conventional experimental workflows that often emphasize characterization of suboptimal materials. On benchmarks like aqueous solubility and oxidative coupling of methane (OCM), BO-ICL matches or outperforms Gaussian processes. In live experiments on the reverse water–gas shift (RWGS) reaction, BO-ICL identifies multimetallic catalysts that approach equilibrium CO yield within 6 and 10 iterations from a pool of 3,700 and 360,000 candidates, respectively. Our method redefines materials representation and accelerates discovery, with broad applications across catalysis, materials science, and AI.

Calibration↗

Synthetic method of analogues for emerging infectious disease forecasting

The Method of Analogues (MOA) has gained popularity in the past decade for infectious disease forecasting due to its non-parametric nature. In MOA, the local behavior observed in a time series is matched to the local behaviors of several historical time series. The known values that directly follow the historical time series that best match the observed time series are used to calculate a forecast. This non-parametric approach leverages historical trends to produce forecasts without extensive parameterization, making it highly adaptable. However, MOA is limited in scenarios where historical data is sparse. This limitation was particularly evident during the early stages of the COVID-19 pandemic, where the emerging global epidemic had little-to-no historical data. In this work, we propose a new method inspired by MOA, called the Synthetic Method of Analogues (sMOA). sMOA replaces historical disease data with a library of synthetic data that describe a broad range of possible disease trends. This model circumvents the need to estimate explicit parameter values by instead matching segments of ongoing time series data to a comprehensive library of synthetically generated segments of time series data. We demonstrate that sMOA has competitive performance with state-of-the-art infectious disease forecasting models, out-performing 78% of models from the COVID-19 Forecasting Hub in terms of averaged Mean Absolute Error and 76% of models from the COVID-19 Forecasting Hub in terms of averaged Weighted Interval Score. Additionally, we introduce a novel uncertainty quantification methodology designed for the onset of emerging epidemics. Developing versatile approaches that do not rely on historical data and can maintain high accuracy in the face of novel pandemics is critical for enhancing public health decision-making and strengthening preparedness for future outbreaks.

97 MATHEMATICS AND COMPUTING↗

ILAMBv2.7 benchmarking results comparing E3SMv2.1 land-atmosphere coupled (BGCv2LNDATM) and stand alone land (ELM) simulations with CMIP6 emission driven historical simulations

This dataset contains land model benchmarking results for the Energy Exascale Earth System Model version 2.1 (E3SMv2.1), including outputs from both coupled biogeochemistry simulations and stand-alone land model simulations. These results are compared against several emission-driven historical simulations from the Coupled Model Intercomparison Project Phase 6 (CMIP6). Benchmarking was conducted using the International Land Model Benchmarking (ILAMB) package, version 2.7 (ILAMBv2.7). CMIP6 model outputs were sourced from the Earth System Grid Federation (ESGF), while the E3SMv2.1 results were derived from raw model outputs. These outputs underwent processing steps such as time serialization, conservative regridding, and data standardization to ensure comparability. For spatial interpolation, the Earth System Modeling Framework (ESMF) tool, ESMF_RegridWeightGen, was employed to generate regridding weights, enabling the transformation of E3SM’s native cubed-sphere grid to a regular latitude-longitude grid.

Feng, Sha [PNNL]↗

Gaussian FLOWERS: Wind-rose-based analytical integration of Gaussian wake model for extremely fast AEP estimation

A major cost in the study of wind farm layout optimization is the repeated evaluation of the annual energy production (AEP). The current approach to estimating AEP requires a large set of flow simulations to be performed that cover each discrete wind speed and direction combination contained within the wind rose, followed by a probability-weighted sum of the power production resulting from each simulation. Even with inexpensive engineering wake models, this numerical integration scheme can lead to high computational costs. In this paper, we derive an analytical formulation for estimating farm AEP across every wind direction, based on a Gaussian wake velocity model, which reduces the number of wind farm simulations to a single function evaluation. As a result, we find that the Gaussian-FLOWERS approach reduces the time for AEP calculations by more than two orders of magnitude with a small trade-off in accuracy when compared to a conventional approach. This massive reduction in computation cost is useful to reduce overall costs in wind farm layout optimization studies.

17 WIND ENERGY↗

Weak-charge form-factor determination at the electron-ion collider

Determining the weak charge form factor, 𝐹 𝑊 ⁡(𝑄 2 ), of nuclei over a continuous range of momentum transfers, 0 ≲ 𝑄 2 ≲ 0.1 GeV 2 , is essential for mapping out the distribution of neutrons in nuclei. The neutron density distribution has significant implications for a broad range of areas, including studies of nuclear structure, neutron stars, and physics beyond the Standard Model. Currently, our knowledge of 𝐹 𝑊 ⁡(𝑄 2 ) comes primarily from fixed target experiments that measure the parity-violating asymmetry in coherent elastic electron-ion scattering. Fixed target experiments, such as CREX and PREX-1,2, have provided high-precision weak charge form factor extractions for the 48 Ca and 208 Pb nuclei, respectively. However, a major limitation of fixed target experiments is that they each provide data only at a single value of 𝑄 2 . With the proposed electron-ion collider (EIC) on the horizon, we explore its potential to impact the determination of the weak charge form factor. While it cannot compete with the precision of fixed target experiments, it can provide data over a wide and continuous range of 𝑄 2 values, and for a wide variety of nuclei. We show that with data corresponding to an integrated luminosity of ℒ ∼ 500/𝐴 fb −1 , where 𝐴 is the nucleus atomic weight, the EIC can significantly impact constraints by lifting degeneracies in theoretical models of the neutron density distribution. Ensuring EIC detector coverage at low 𝑄 2 and large negative pseudorapidities will be essential for such 𝐹 𝑊 ⁡(𝑄 2 ) measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Screening a knowledge‐based library of low molecular weight compounds against the proline biosynthetic enzyme 1‐pyrroline‐5‐carboxylate 1 ( PYCR1)

Abstract Δ 1 ‐pyrroline‐5‐carboxylate reductase isoform 1 (PYCR1) is the last enzyme of proline biosynthesis and catalyzes the NAD(P)H‐dependent reduction of Δ 1 ‐pyrroline‐5‐carboxylate toL‐proline. High PYCR1 gene expression is observed in many cancers and linked to poor patient outcomes and tumor aggressiveness. The knockdown of thePYCR1gene or the inhibition of PYCR1 enzyme has been shown to inhibit tumorigenesis in cancer cells and animal models of cancer, motivating inhibitor discovery. We screened a library of 71 low molecular weight compounds (average MW of 131 Da) against PYCR1 using an enzyme activity assay. Hit compounds were validated with X‐ray crystallography and kinetic assays to determine affinity parameters. The library was counter‐screened against human Δ 1 ‐pyrroline‐5‐carboxylate reductase isoform 3 and proline dehydrogenase (PRODH) to assess specificity/promiscuity. Twelve PYCR1 and one PRODH inhibitor crystal structures were determined. Three compounds inhibit PYCR1 with competitive inhibition parameter of 100 μM or lower. Among these, (S)‐tetrahydro‐2H‐pyran‐2‐carboxylic acid (70 μM) has higher affinity than the current best tool compoundN‐formyl‐l‐proline, is 30 times more specific for PYCR1 over human Δ 1 ‐pyrroline‐5‐carboxylate reductase isoform 3, and negligibly inhibits PRODH. Structure‐affinity relationships suggest that hydrogen bonding of the heteroatom of this compound is important for binding to PYCR1. The structures of PYCR1 and PRODH complexed with 1‐hydroxyethane‐1‐sulfonate demonstrate that the sulfonate group is a suitable replacement for the carboxylate anchor. This result suggests that the exploration of carboxylic acid isosteres may be a promising strategy for discovering new classes of PYCR1 and PRODH inhibitors. The structure of PYCR1 complexed withl‐pipecolate and NADH supports the hypothesis that PYCR1 has an alternative function in lysine metabolism.

Biochemistry & Molecular Biology↗

Tropical Cyclone Super Resolution using conditional diffusion denoising probabilistic model from mesoscale simulation to LES

Accurate modeling of tropical cyclone wind fields is essential for the design, risk assessment, and operational planning of offshore energy infrastructure. While mesoscale simulations are widely used thanks to their computational efficiency, they lack the necessary resolution to capture key features such as wind shear and veer profiles as well as the distribution turbulent kinetic energy (TKE). High-fidelity large-eddy simulation (LES) models on the other hand, can resolve turbulent structures and provide a more accurate representation of the complex wind field, albeit at a higher computational cost. To address this modeling gap, we introduce a two-part generative framework to enhance the resolution and physics-capturing ability of mesoscale simulations. First, a reduced-order model based on Karhunen–Loève (KL) decomposition is used to extract dominant spatial modes from one-dimensional mean wind profiles. A multilayer perceptron (MLP) is trained to map mesoscale mode weights to their LES counterparts, enabling accurate reconstruction of vertical velocity profiles. Second, a conditional Diffusion Denoising Probabilistic Model (DDPM) is developed to super-resolve coarse and low-fidelity mesoscale velocity fields, recovering fine-scale turbulence structures and stress distributions. The framework is evaluated across different tropical cyclone intensity categories defined by the Saffir–Simpson scale and demonstrates strong performance in both interpolation and extrapolation tasks. The generated fields accurately reproduce spatial coherence, stress distributions, and spectral energy characteristics observed in LES data. By bridging the fidelity gap between mesoscale and LES outputs, this approach offers a scalable, data-driven solution for enhancing the representation of tropical cyclone wind fields, enabling more robust offshore energy infrastructure systems design in tropical-cyclone-prone areas.

17 WIND ENERGY↗

Reweighting Monte Carlo predictions and automated fragmentation variations in Pythia 8

This work reports on a method for uncertainty estimation in simulated collider-event predictions. The method is based on a Monte Carlo-veto algorithm, and extends previous work on uncertainty estimates in parton showers by including uncertainty estimates for the Lund string-fragmentation model. This method is advantageous from the perspective of simulation costs: a single ensemble of generated events can be reinterpreted as though it was obtained using a different set of input parameters, where each event now is accompanied with a corresponding weight. This allows for a robust exploration of the uncertainties arising from the choice of input model parameters, without the need to rerun full simulation pipelines for each input parameter choice. Such explorations are important when determining the sensitivities of precision physics measurements. Accompanying code is available at https://gitlab.com/uchep/mlhad-weights-validation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Unified understanding of the impact of semiflexibility, concentration, and molecular weight on macromolecular-scale ring diffusion

Conformationally fluctuating, globally compact macromolecules such as polymeric rings, single-chain nanoparticles, microgels, and many-arm stars display complex dynamic behaviors due to their rich topological structure and intermolecular organization. Synthetic rings are hybrid objects with conformations that display both ideal random walk and compact globular features, which can serve as models of genomic DNA. To date, emphasis has been placed on the effect of ring molecular weight on their unusual behaviors. Here, we combine simulations and a microscopic force-level theory to build a unified understanding for how key aspects of ring dynamics depend on different tunable molecular properties including backbone rigidity, monomer concentration, degree of traditional entanglement, and molecular weight. Our large-scale molecular dynamics simulations of ring melts with very different backbone stiffnesses reveal unanticipated behaviors which agree well with our generalized theory. This includes a universal master curve for center-of-mass diffusion constants as a function of molecular weight scaled by a chemistry and thermodynamic state-dependent critical molecular weight that generalizes the concept of an entanglement cross-over for linear chains. The key physics is how backbone rigidity and monomer concentration induced changes of the entanglement length, interring packing, degree of interpenetration, and liquid compressibility slow down space-time dynamic-force correlations on macromolecular scales. A power law decay of the center-of-mass diffusion constant with inverse molecular weight squared is the first consequence, followed by an ultraslow activated hopping transport regime. Our results set the stage to address slow dynamics and kinetic arrest in different families of compact synthetic and biological polymeric systems.

Science & Technology - Other Topics↗

Marine Algae Industrialization Consortium (MAGIC): Combining biofuel and high-value bioproducts to meet the RFS

The Marine Algae Industrialization Consortium (MAGIC) was formed to address pressing challenges in the commercialization of microalgae as a source of biofuel. The “Marine Algae Industrialization Consortium (MAGIC): Combining biofuel and high-value bioproducts to meet the RFS” project formally addressed two US Department of Energy Bioenergy Technologies Office (BETO) goals: (1) Model the sustainable supply of 1 million metric tonnes ash free dry weight (AFDW) cultivated algal biomass and (2) Demonstrate valuable co-products produced along with biofuel intermediates to increase value of algal biomass by 30%. To achieve these goals, the project demonstrated and validated high-value co-products to drive down the cost of biofuel by increasing the value of algae “co-products” towards increasing the selling price of total algae biomass as one of the key drivers of economics and adoption. This was accomplished through five core, interdependent tasks including: (1) strain selection to identify and deliver strains for mass culture, (2) mass culture using a hybrid cultivation system and following key operating parameters for downstream applications to provide algae feedstock, (3) recovery and conversion to evaluate two alternative methods to separate dry algae biomass into oil and residuals for downstream testing, (4) product assessment to determine biofuel, aquafeed or poultry feed product efficacy using algae biomass fractions as well as to provide critical performance data for valuation and (5) commercialization to use technoeconomic and life cycle assessments (TEA/LCA) as iterative design and assessment tools including consideration of target markets, competitors, and distribution channels to guide product assessment, development and valuation. A total of 46 peer-review publications, many open-access, provide detail of much of the work carried out and the results of the tasks. Additional reports and presentations provide other technical and public engagement material. At a high level, using a variety of approaches, more than 1000 marine microalgae strains were evaluated to ultimately identify the seven winners that were down-selected to be grown in mass culture. Strain selection demonstrated that there were no ‘super strains’ and that each candidate had strengths and limitations for specific products, growth conditions or operational considerations. Mass culture growth of these seven strains at >5000 L / 29 m 2 scale found that four them were suitable for product assessment. More than 250 kg of biomass was produced across hundreds of pond runs along with thousands of cultivation entries on the growth and biomass characteristics as well as environmental parameters. In the process, dozens of standard operating procedures were generated as was custom software to process and analyze cultivation data. Recovery and conversion of algae biomass demonstrated that a hexane solvent based extraction protocol was most effective at recovering oil (biocrude) from algae and four strains were processed to produce oil and lipid extracted algae (residuals) for downstream testing. Membrane-based oil separation was less successful, but may still be applicable to other commercial applications in the future. Product testing demonstrated that algae biocrude is of high quality and hydrotreating generated numerous fractions of high quality composition for fuel and lubricate based applications. Aquafeed studies performed at a variety of scales showed that both whole and defatted (lipid extracted algae) microalgae were suitable as a feed ingredient, but that the specifics of the fed animal and biochemical composition of the algae are critical factors when determining formulation. Similarly, poultry studies on whole and defatted microalgae generally showed positive outcomes on animal growth and health, with some microalgae providing enhanced nutritional composition of the animal product. Economic and life cycle assessments covered a wide range of possible commercialization and sustainability scenarios. Replacement value, improved product value added, consumer values marketing added valuation and improved animal health were considered as alternatives for microalgae valuation. Using the open pond system, algae productivity was identified as the key driver of commercialization economics, but combination of co-products (e.g. animal feed) with biofuel production substantially increased the total selling price of algae. Modeled microalgae selling price exceeded $\$$1500/tonne and could generate competitive biofuel selling prices below $\$$5 gallon gas equivalents using realistic algal productivities. Short (process scale) and longer (decadal trends) sustainability assessments show that marine microalgae can enhance the sustainability of energy production and lead to other realized benefits in water, fertilizer and land use for other sectors (e.g. agriculture). This project successfully demonstrated all of the components of an end-to-end process from mass microalgae cultivation and dewatering, to recovery and conversion of algae biomass components, to final product demonstration and process valuation; the combined results provide a framework for future commercialization of algae based biofuels.

09 BIOMASS FUELS↗

More high-impact atmospheric river-induced extreme precipitation events under warming in a high-resolution model

Extreme precipitation events, as occurred in Europe 2021, or western North America 2023, with an intensity of 75 mm/day or 100 mm/day, respectively, exert a catastrophic impact. Lower-resolution (~100 km) climate models cannot simulate the intensity, nor their response to greenhouse warming. Using a high-resolution (~25 km) model capable of simulating such events, here we show that frequency-weighted area of such events over western Europe and the west coast of North America likely expands by more than 80% under an approximately 4°C of global warming from the historical levels. Along the west coasts of Europe and North America, area impacted by atmospheric rivers-induced extreme precipitation events is projected to double, driven by intensified landfalling atmospheric rivers. Thermodynamic processes drive the increase, whereas dynamic processes reduce the intensity over western Europe but enhance it along the west coast of North America. Here our findings provide policy-relevant information for climate adaptation strategies.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of FluSight influenza forecasting in the 2021–22 and 2022–23 seasons with a new target laboratory-confirmed influenza hospitalizations

Accurate forecasts can enable more effective public health responses during seasonal influenza epidemics. For the 2021–22 and 2022–23 influenza seasons, 26 forecasting teams provided national and jurisdiction-specific probabilistic predictions of weekly confirmed influenza hospital admissions for one-to-four weeks ahead. Forecast skill is evaluated using the Weighted Interval Score (WIS), relative WIS, and coverage. Six out of 23 models outperform the baseline model across forecast weeks and locations in 2021–22 and 12 out of 18 models in 2022–23. Averaging across all forecast targets, the FluSight ensemble is the 2nd most accurate model measured by WIS in 2021–22 and the 5th most accurate in the 2022–23 season. Forecast skill and 95% coverage for the FluSight ensemble and most component models degrade over longer forecast horizons. In this work we demonstrate that while the FluSight ensemble was a robust predictor, even ensembles face challenges during periods of rapid change.

59 BASIC BIOLOGICAL SCIENCES↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Tensor renormalization group for fermions

Abstract We review the basic ideas of the tensor renormalization group method and show how they can be applied for lattice field theory models involving relativistic fermions and Grassmann variables in arbitrary dimensions. We discuss recent progress for entanglement filtering, loop optimization, bond-weighting techniques and matrix product decompositions for Grassmann tensor networks. The new methods are tested with two-dimensional Wilson–Majorana fermions and multi-flavor Gross–Neveu models. We show that the methods can also be applied to the fermionic Hubbard model in 1+1 and 2+1 dimensions.

Physics↗

Mountain Basin Controls on the Snow-to-Streamflow Signal: An AIC-Weighted Multiple Linear Regression Framework

A regression-based analysis quantifies how basin characteristics modulate the snow-to-streamflow signal. First, we use the ERA5-Land reanalysis gridded product (European Centre for Medium Range Weather Forecasts reanalysis 5 -Land component) for 4,655 hydrologic unit code - 10 (HUC10) mountain basins across the western United States (US) for water years 1987–2024. Linear regressions are performed for peak snow water equivalent (SWE) and annual streamflow for each mountain basin. Models use ordinary least squares in Python’s statsmodels package. After which, an Akaike Information Criterion (AIC)–weighted ensemble multiple linear regression (MLR) framework with 47 watershed traits is used to predict the linear regression coefficient of determination (r-squared) defining the ability of peak SWE to predict annual streamflow across all mountain basin. Predictor sets are constrained to avoid multicollinearity by excluding models with variance inflation factors (VIF) greater than 5. Mountain basin traits included in the MLR include seasonal climate, topography, vegetation type and structure, and bedrock geology. Accepted models are considered if their AIC is within 2.0 of the model with the minimum AIC, or best model. To compare predictor influence across acceptable models, we computed standardized regression coefficients. To evaluate structural redundancy among models, we constructed binary inclusion vectors for each acceptable model, denoting whether a predictor was present (1) or absent (0). Core predictor variables are defined as occurring in at least 67% of the acceptable models. For this regional analysis, only one model was found acceptable, with higher snow-to-streamflow translation (higher r-squared) occurring in colder mountain basins with higher relative winter precipitation, more snow accumulation and a lower fraction of annual precipitation that falls in the spring and summer. The second component of the data package uses previously published, high-resolution output from an integrated hydrological model of the East River watershed using the U.S. Geological Survey Groundwater and Surface water Flow model (GSFLOW, doi:10.15485/1998576). East River MLR expands upon the approach described above to explore the response of five streamflow metrics—annual streamflow, runoff efficiency, 7-day minimum flow, low-flow duration, and non-perennial stream fraction to snow system indicators including peak SWE, snow-covered area, snow disappearance date, and the fraction of basin area characterized by low-to-no snow, as well as seasonal precipitation and temperature, and annual hydrologic variables representing soil moisture, evapotranspiration (ET), the partitioning of incoming precipitation to evapotranspiration (ET/P), groundwater storage, and groundwater inflow to streams. MLR was done on all water years (P0: 1987-2024) and for each period as determined in the split analysis using pooled regression techniques (P1: 1987-2011 and P2: 2012-2024) to evaluate shifting predictor variable emphasis on streamflow generation. Results indicate that since 2012, peak SWE has lost statistical strength in its prediction of annual streamflow and runoff efficiency, and the indirect influence of spring temperature has emerged as critically important. Low-flow metrics remain largely influenced by soil moisture, vegetation water use and groundwater inflows with summer precipitation becoming a direct influence on minimum summer flow. Together, these data and Python-based analysis tools provide a framework for identifying the key watershed characteristics that control how streamflow responds to snow from year to year. The package also helps quantify uncertainty in statistical models and assess how snow–streamflow relationships vary across regions and over time. This dataset contains comma-separated values files (.csv), text files (.txt), python code files (.py), figure files (.png), and shapefiles (.cpg, .dbf, .prj, .sbn, .sbx, .shp, .xml). Further details on file contents and MLR execution can be found in the readme file and the FLMD files. Work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

MLSPICE: Machine Learning based SPICE Modeling Platform for Power Magnetics

Electrical power converters are critical to a wide range of applications ranging from renewable integration to transportation electrification, and can be a key factor determining the size, weight, and efficiency of energy conversion systems. Magnetic components are typically the largest and least efficient components in power electronics. While there have been major strides in the modeling and analysis of power semiconductor devices and circuit simulations, the necessary advances in the design of power magnetics have lagged. In this project, we have transformed the modeling and design of power magnetics with machine learning enabled methods and catalyze simultaneous disruptive improvements for ML-based power electronics design tools. A fully automated open-source machine learning based magnetics modeling platform – the MagNet project - with innovations in full stack have been developed to greatly accelerate the design process and provide new insights to magnetic material and geometry design. The ARPA-E funded MagNet platform contains three major building blocks: 1) a ML-Integrated Data Acquisition System (MIDAS): a highly automated data acquisition testbed which is capable of measuring a large number of magnetic cores with a wide range of electrical circuit excitations; 2) a ML-integrated Core Loss Model (MICLM): a machine-learning trained modeling method for modeling the core loss and saturation effects of magnetic materials for arbitrary excitation waveforms; 3) ML-guided Magnetics SPICE Simulation Tool (PMSPICE): a fully integrated CAD tool which can simulate the magnetics in SPICE. It can help the designers to quickly model the linear and non-linear characteristics of magnetic components and evaluate their behavior in SPICE simulations. The developed MagNet system has fully demonstrated the proposed performance target and has been open sourced to the entire power electronics community to advance the modeling and design of power magnetics from many different angles.

36 MATERIALS SCIENCE↗

Modeling injection-induced fault slip using long short-term memory networks

Stress changes due to changes in fluid pressure and temperature in a faulted formation may lead to the opening/shearing of the fault. This can be due to subsurface (geo)engineering activities such as fluid injections and geologic disposal of nuclear waste. Such activities are expected to rise in the future making it necessary to assess their short- and long-term safety. Here, a new machine learning (ML) approach to model pore pressure and fault displacements in response to high-pressure fluid injection cycles is developed. The focus is on fault behavior near the injection borehole. To capture the temporal dependencies in the data, long short-term memory (LSTM) networks are utilized. To prevent error accumulation within the forecast window, four critical measures to train a robust LSTM model for predicting fault response are highlighted: (i) setting an appropriate value of LSTM lag, (ii) calibrating the LSTM cell dimension, (iii) learning rate reduction during weight optimization, and (iv) not adopting an independent injection cycle as a validation set. Several numerical experiments were conducted, which demonstrated that the ML model can capture peaks in pressure and associated fault displacement that accompany an increase in fluid injection. The model also captured the decay in pressure and displacement during the injection shut-in period. Further, the ability of an ML model to highlight key changes in fault hydromechanical activation processes was investigated, which shows that ML can be used to monitor risk of fault activation and leakage during high pressure fluid injections.

58 GEOSCIENCES↗

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗