Search NASA⌕ Search

SEARCH · Search NASA

Results for “data harmonization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Quadrature Based Neural Network Learning of Stochastic Hamiltonian Systems

Hamiltonian Neural Networks (HNNs) provide structure-preserving learning of Hamiltonian systems. In this paper, we extend HNNs to structure-preserving inversion of stochastic Hamiltonian systems (SHSs) from observational data. We propose the quadrature-based models according to the integral form of the SHSs’ solutions, where we denoise the loss-by-moment calculations of the solutions. The integral pattern of the models transforms the source of the essential learning error from the discrepancy between the modified Hamiltonian and the true Hamiltonian in the classical HNN models into that between the integrals and their quadrature approximations. This transforms the challenging task of deriving the relation between the modified and the true Hamiltonians from the (stochastic) Hamilton–Jacobi PDEs, into the one that only requires invoking results from the numerical quadrature theory. Meanwhile, denoising via moments calculations gives a simpler data fitting method than, e.g., via probability density fitting, which may imply better generalization ability in certain circumstances. Numerical experiments validate the proposed learning strategy on several concrete Hamiltonian systems. The experimental results show that both the learned Hamiltonian function and the predicted solution of our quadrature-based model are more accurate than that of the corrected symplectic HNN method on a harmonic oscillator, and the three-point Gaussian quadrature-based model produces higher accuracy in long-time prediction than the Kramers–Moyal method and the numerics-informed likelihood method on the stochastic Kubo oscillator as well as other two stochastic systems with non-polynomial Hamiltonian functions. Moreover, the Hamiltonian learning error εH arising from the Gaussian quadrature-based model is lower than that from Simpson’s quadrature-based model. These demonstrate the superiority of our approach in learning accuracy and long-time prediction ability compared to certain existing methods and exhibit its potential to improve learning accuracy via applying precise quadrature formulae.

Mathematics↗

The role of the droplet interface in controlling the multiphase oxidation of thiosulfate by ozone

Predicting reaction kinetics in aqueous microdroplets, including aerosols and cloud droplets, is challenging due to the probability that the underlying reaction mechanism can occur both at the surface and in the interior of the droplet. Additionally, few studies directly measure the surface activities of doubly charged anions, despite their prevalence in the atmosphere. Here, deep-UV second harmonic generation spectroscopy is used to probe surface affinities of the doubly charged anions thiosulfate, sulfate, and sulfite, key species in the thiosulfate ozonation reaction mechanism. Thiosulfate has an appreciable surface affinity with a measured Gibbs free energy of adsorption of -7.3 ± 2.5 kJ mol -1 in neutral solution, while sulfate and sulfite exhibit negligible surface propensity. The Gibbs free energy is combined with data from liquid flat jet ambient pressure X-ray photoelectron spectroscopy to constrain the concentration of thiosulfate at the surface in our model. Stochastic kinetic simulations leveraging these novel measurements show that the primary reaction between thiosulfate and ozone occurs at the interface and in the bulk, with the contribution of the interface decreasing from ~65% at pH 5 to ~45% at pH 13. Additionally, sulfate, the major product of thiosulfate ozonation and an important species in atmospheric processes, can be produced by two different pathways at pH 5, one with a contribution from the interface of >70% and the other occurring predominantly in the bulk (>98%). The observations in this work have implications for mining wastewater remediation, atmospheric chemistry, and understanding other complex reaction mechanisms in multiphase environments. Future interfacial or microdroplet/aerosol chemistry studies should carefully consider the role of both surface and bulk chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrated Distribution Planning

The contemporary distribution planning landscape is comprised of an increasing number of factors that require integration into the engineering of the modern electric grid. Expectations for electric utilities to accommodate heightened awareness of stakeholders' interest in things like decarbonization, resilience and equity are growing. As these interests are formed into objectives, many jurisdictions will experience increasing levels of load modifying technologies like DER, building and industrial electrification and electric vehicles which prove not only to challenge the capabilities of the grid; but the processes by which planning for it is traditionally done. Other related factors that strain the conventional distribution planning mold are the swelling amount and sources of data associated with these technologies and the need it creates for improved capabilities in the processes and tools that manage it. As the complexity of the distribution system expands, so will the distribution system's effects on the transmission and generation systems that it is a part of. Forecasting distribution system load and DER are examples of areas where this complexity will manifest, and harmonizing distribution forecasting with transmission and generation forecasting requires higher amounts of intentionality as these typically separate processes become a solitary one. Of course, core activities do not cease as a utility begins to integrate these other factors, and in this webinar we explore specifics of how distribution planning can be expected to evolve as progress towards Integrated Distribution System Planning is made.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reducing the Cost of CCSD Basis Set Extrapolation in Ab Initio Computational Thermochemistry

Here, a series of approximations to CCSD contributions in computational model chemistries is presented in the context of kcal mol –1 , kJ mol –1 , and 20 cm –1 theoretical predictions of total atomization energies, benchmarked within the HEAT+CH 4 test suite. A specific set of circumstances where MP2, without empirical scaling, may be used as an effective intermediate in the first two of these accuracy ranges was determined. However, SDQ-MP4, a method long used in pursuit of kcal mol –1 accuracy but relatively unstudied in the subchemical accuracy community, offers significant improvement over the quality of MP2 as a basis-set intermediate at significantly reduced cost compared to CCSD. Given this, we argue for SDQ-MP4 as the de facto CCSD basis-set intermediate in sub-chemical accuracy calculations when CCSD in a desired basis set becomes unaffordable. We additionally report on a “CBS-like” scheme, where MP2 and SDQ-MP4 are used in conjunction to create a “cheap” three-part approximation of large CCSD basis set limits. The data for the CCSD approximation schemes are organized in such a way that model chemistry developers can locate an analog of their current approach for the CCSD basis set limit and explore alternative intermediates that either decrease computational cost or increase computational accuracy. We also show, for a handful of molecules, that SDQ-MP4 shows promise as an effective basis-set intermediate for harmonic and fundamental frequency computations, allowing for zero-point corrections of nearly CCSD(T)/ANO1 quality using simple composite methods that only require CCSD(T)/ANO0.

Thorpe, James H. [Argonne National Laboratory (ANL↗

RectifHydPlus: Forty Year Hydropower Generation Reanalysis for Conterminous United States, Version 1.1.

This dataset contains monthly hydropower net-generation totals for 590 plants (each >10 MW) across the conterminous United States (CONUS) from 1980 to 2019. RectifHydPlus v1.1 includes one harmonized table of historical monthly generation—backfilled with observed monthly values where available—and two companion tables: (i) an estimates-only version with no backfill and (ii) a hydrological-control version that removes the effects of capacity and operational change. Each table comprises 23,600 records (590 plants × 40 years). The dataset was developed to address temporal gaps and inconsistencies in publicly available hydropower generation data as available through EIA-923 survey reports. Each record includes a quality label denoting the underlying proxy—from best (direct reservoir releases) to weakest (pattern copied from similar years). By combining the agency-reported survey records with observed and simulated hydrologic releases, RectifHydPlus offers complete, quality-labeled monthly estimates suitable for trend analysis and generation of hydropower generation inputs for energy-water modeling.

Turner, Sean [Oak Ridge National Laboratory (ORNL)↗

Reinforcement Learning for In-Spill Optimization of the Mu2e Resonant Extraction: Compensating Non-Stationarity

We present design considerations and challenges for the fast machine learning component of a third-order resonant beam extraction regulation system being commissioned to deliver steady beam rates to the mu2e experiment at Fermilab. Dedicated quadrupoles drive the tune toward the 29/3 resonance each spill, extracting beam at kV multiwire septa. The overall Spill Regulation System consists of (1) a “slow” process using ~100-spill averages to adjust the base quad ramp infrequently, (2) a feedforward harmonic content compensator, and (3) the “fast” ML agent reacting during each ongoing spill with on-the-fly additive corrections to the sum of (1) and (2). We have demonstrated improved beam-rate steadying for a fast ML agent compared to a PID controller using a quasi-physical spill simulation, and demonstrated distillation of that simulation into a predictive surrogate model. Current work includes a data-and-training pipeline to generate data-aware surrogates with real-world dynamics, even as the dynamics shift unpredictably. The surrogates are to act as RL environments against which to train our fast ML control agents before deploying them on FPGA in the live system. Further current efforts focus on modeling and controlling beam loss around the storage ring, understanding additional available hardware inputs to the model, and the interplay of these with beam-steadying performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Life-Cycle Assessment Integration into Scalable Open-Source Numerical Models (LiAISON) for Prospective Impact Analysis of Novel Technologies

Decarbonizing the industrial sector is a significant challenge in achieving a net-zero greenhouse gas (GHG) emissions economy by 2050 and the Paris Agreement, i.e., a global climate change mitigation target of achieving a maximum average temperature change potential of 1.5 Degrees Celsius or less by 2100 with respect to pre-industrial levels. In the United States (US), the industrial sector accounts for 23% of total GHG emissions and is home to a number of hard-to-electrify activities. The chemicals subsector has the single largest subsector emissions profile after direct emissions from fossil fuel combustion and leakage from fossil fuel distribution systems. Within the chemicals subsector, many processes depend on hydrogen or ammonia precursors. Decarbonizing these two commodities would contribute significantly to decarbonizing the industrial sector as hydrogen could also be used for low carbon steel production (e.g., hydrogen-based direct reduction of iron) and other industrial applications. Emerging technologies require the application of prospective life cycle assessment (LCA), which can account for technology (foreground) scaling and process improvements via learning-by-doing, among others. In many cases, the future system context (background) in which the technologies are assumed to operate in is equally relevant. Background scenarios generated by integrated assessment models (IAM) can coherently incorporate potential future dynamics of the energy-climate-human-land system. Further, IAM scenarios are harmonized across socioeconomic and climate change mitigation pathways, which facilitates the comparability of prospective LCAs using different IAMs. We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics. The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE). We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Additionally we compare our results by linking two other prospective models with LiAISON - GCAM (Global Change Assessment Model) and ReEDS (Regional Energy Deployment System) to analyze the effect of changing background scenarios using varying predictions in life cycle analysis.

decarbonizing↗

Life-Cycle Assessment Integration into Scalable Open-Source Numerical Models (LiAISON) for Analyzing Emerging Low-Carbon Technologies

Decarbonizing the industrial sector is a significant challenge in achieving a net-zero greenhouse gas (GHG) emissions economy by 2050 and the Paris Agreement, i.e., a global climate change mitigation target of achieving a maximum average temperature change potential of 1.5 degrees Celsius or less by 2100 with respect to pre-industrial levels. In the United States (US), the industrial sector accounts for 23% of total GHG emissions and is home to a number of hard-to-electrify activities. The chemicals subsector has the single largest subsector emissions profile after direct emissions from fossil fuel combustion and leakage from fossil fuel distribution systems. Within the chemicals subsector, many processes depend on hydrogen or ammonia precursors. Decarbonizing these two commodities would contribute significantly to decarbonizing the industrial sector as hydrogen could also be used for low carbon steel production (e.g., hydrogen-based direct reduction of iron) and other industrial applications. Emerging technologies require the application of prospective life cycle assessment (LCA), which can account for technology (foreground) scaling and process improvements via learning-by-doing, among others. In many cases, the future system context (background) in which the technologies are assumed to operate in is equally relevant. Background scenarios generated by integrated assessment models (IAM) can coherently incorporate potential future dynamics of the energy-climate-human-land system. Further, IAM scenarios are harmonized across socioeconomic and climate change mitigation pathways, which facilitates the comparability of prospective LCAs using different IAMs. We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics. The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE). We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Under baseline projections (i.e., no decarbonization goals), neither process reaches parity with the incumbent technology across several environmental metrics. Under the decarbonization scenarios, the underlying sectoral shifts result in declining impacts over time, compared to 2020 levels, except for metal depletion levels, which increase. The background shifts postulate a heavily decarbonized economy and energy system, which help technologies reach parity with SMR between 2040-2050 (RCP2.6) and 2030-2040 (RCP1.9) for global warming. Despite declines across several other metrics over time, neither PtH2 technology break even with SMR by 2100 besides for global warming.

decarbonizing↗

Life-Cycle Assessment Integration into Scalable Open-Source Numerical Models (LiAISON) for Analyzing Emerging Low-Carbon Technologies

Decarbonizing the industrial sector is a significant challenge in achieving a net-zero greenhouse gas (GHG) emissions economy by 2050 and the Paris Agreement, i.e., a global climate change mitigation target of achieving a maximum average temperature change potential of 1.5 degrees C or less by 2100 with respect to pre-industrial levels. In the United States (US), the industrial sector accounts for 23% of total GHG emissions and is home to a number of hard-to-electrify activities. The chemicals subsector has the single largest subsector emissions profile after direct emissions from fossil fuel combustion and leakage from fossil fuel distribution systems. Within the chemicals subsector, many processes depend on hydrogen or ammonia precursors. Decarbonizing these two commodities would contribute significantly to decarbonizing the industrial sector as hydrogen could also be used for low carbon steel production (e.g., hydrogen-based direct reduction of iron) and other industrial applications. Emerging technologies require the application of prospective life cycle assessment (LCA), which can account for technology (foreground) scaling and process improvements via learning-by-doing, among others. In many cases, the future system context (background) in which the technologies are assumed to operate in is equally relevant. Background scenarios generated by integrated assessment models (IAM) can coherently incorporate potential future dynamics of the energy-climate-human-land system. Further, IAM scenarios are harmonized across socioeconomic and climate change mitigation pathways, which facilitates the comparability of prospective LCAs using different IAMs. We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE). We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Under baseline projections (i.e., no decarbonization goals), neither process reaches parity with the incumbent technology across several environmental metrics. Under the decarbonization scenarios, the underlying sectoral shifts result in declining impacts over time, compared to 2020 levels, except for metal depletion levels, which increase. The background shifts postulate a heavily decarbonized economy and energy system, which help technologies reach parity with SMR between 2040-2050 (RCP2.6) and 2030-2040 (RCP1.9) for global warming. Despite declines across several other metrics over time, neither PtH2 technology break even with SMR by 2100 besides for global warming.

decarbonizing↗

Towards Prospective LCA Using Life-Cycle Assessment Integration into Scalable Open-Source Numerical Models (LiAISON) Framework for Analyzing Emerging Low-Carbon Technologies

Decarbonizing the industrial sector is a significant challenge in achieving a net-zero greenhouse gas (GHG) emissions economy by 2050 and the Paris Agreement, i.e., a global climate change mitigation target of achieving a maximum average temperature change potential of 1.5 Degrees Celsius or less by 2100 with respect to pre-industrial levels. In the United States (US), the industrial sector accounts for 23% of total GHG emissions and is home to a number of hard-to-electrify activities. The chemicals subsector has the single largest subsector emissions profile after direct emissions from fossil fuel combustion and leakage from fossil fuel distribution systems. Within the chemicals subsector, many processes depend on hydrogen or ammonia precursors. Decarbonizing these two commodities would contribute significantly to decarbonizing the industrial sector as hydrogen could also be used for low carbon steel production (e.g., hydrogen-based direct reduction of iron) and other industrial applications. Emerging technologies require the application of prospective life cycle assessment (LCA), which can account for technology (foreground) scaling and process improvements via learning-by-doing, among others. In many cases, the future system context (background) in which the technologies are assumed to operate in is equally relevant. Background scenarios generated by integrated assessment models (IAM) can coherently incorporate potential future dynamics of the energy-climate-human-land system. Further, IAM scenarios are harmonized across socioeconomic and climate change mitigation pathways, which facilitates the comparability of prospective LCAs using different IAMs. We introduce an open source prospective LCA framework, the Life-cycle Assessment Integration into Scalable Open-source Numerical models (LiAISON), to analyze the non-linear relationships between technology foreground and the future energy system background across a series of midpoint and resource use metrics The integration of LCA and IAM data is achieved using prospective environmental Impact assessment (PREMISE). We showcase it by assessing two Power-to-Hydrogen (PtH2) processes, namely Solid Oxide Electrolysis (SOE) and Polymer Electrolyte Membrane Electrolysis (PEME). We compare the technologies to a baseline of hydrogen production via natural gas-based Steam Methane Reforming (SMR) in a US context of multiple energy system and climate change mitigation futures. Besides providing an analysis that specifies the LCA results ranges with temporal and geospatial explicitness across the two technologies, metrics, and impact assessment methods, this research also aims to establish a base framework that can be expanded to use other IAM generated scenarios and US open-source life cycle inventory (LCI) databases. We find that the temporal environmental performance of either technology or their difference to SMR is directly influenced by the underlying background dynamics. Additionally we compare our results by linking two other prospective models with LiAISON - GCAM(Global Change Assessment Model) and ReEDS (Regional Energy Deployment System) to analyze the effect of changing background scenarios using varying predictions in life cycle analysis.

emissions↗

Persistent global greening over the last four decades using novel long-term vegetation index data with enhanced temporal consistency

Advanced Very High-Resolution Radiometer (AVHRR) satellite observations have provided the longest global daily records from 1980s, but the remaining temporal inconsistency in vegetation index datasets has hindered reliable assessment of vegetation greenness trends. To tackle this, we generated novel global long-term Normalized Difference Vegetation Index (NDVI) and Near-Infrared Reflectance of vegetation (NIRv) datasets derived from AVHRR and Moderate Resolution Imaging Spectroradiometer (MODIS). We addressed residual temporal inconsistency through three-step post processing including cross-sensor calibration among AVHRR sensors, orbital drifting correction for AVHRR sensors, and machine learning-based harmonization between AVHRR and MODIS. After applying each processing step, we confirmed the enhanced temporal consistency in terms of detrended anomaly, trend and interannual variability of NDVI and NIRv at calibration sites. Our refined NDVI and NIRv datasets showed a persistent global greening trend over the last four decades (NDVI: 0.0008 yr -1 ; NIRv: 0.0003 yr -1 ), contrasting with those without the three processing steps that showed rapid greening trends before 2000 (NDVI: 0.0017 yr -1 ; NIRv: 0.0008 yr -1 ) and weakened greening trends after 2000 (NDVI: 0.0004 yr -1 ; NIRv: 0.0001 yr -1 ). These findings highlight the importance of minimizing temporal inconsistency in long-term vegetation index datasets, which can support more reliable trend analysis in global vegetation response to climate changes.

54 ENVIRONMENTAL SCIENCES↗

Verbal Learning and Memory Deficits across Neurological and Neuropsychiatric Disorders: Insights from an ENIGMA Mega Analysis

Deficits in memory performance have been linked to a wide range of neurological and neuropsychiatric conditions. While many studies have assessed the memory impacts of individual conditions, this study considers a broader perspective by evaluating how memory recall is differentially associated with nine common neuropsychiatric conditions using data drawn from 55 international studies, aggregating 15,883 unique participants aged 15–90. The effects of dementia, mild cognitive impairment, Parkinson’s disease, traumatic brain injury, stroke, depression, attention-deficit/hyperactivity disorder (ADHD), schizophrenia, and bipolar disorder on immediate, short-, and long-delay verbal learning and memory (VLM) scores were estimated relative to matched healthy individuals. Random forest models identified age, years of education, and site as important VLM covariates. A Bayesian harmonization approach was used to isolate and remove site effects. Regression estimated the adjusted association of each clinical group with VLM scores. Memory deficits were strongly associated with dementia and schizophrenia (p < 0.001), while neither depression nor ADHD showed consistent associations with VLM scores (p > 0.05). Differences associated with clinical conditions were larger for longer delayed recall duration items. By comparing VLM across clinical conditions, this study provides a foundation for enhanced diagnostic precision and offers new insights into disease management of comorbid disorders.

Neurosciences & Neurology↗

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Structure Sensitive Reaction Kinetics of Chiral Molecules on Intrinsically Chiral Surfaces

Enantiospecific heterogeneous catalysis utilizes chiral surfaces to resolve enantiomers via structure sensitive surface chemistry. The catalyst design challenge is the identification of chiral surface structures that maximize enantiospecificity. Herein, we develop data driven models for the enantiospecificity of tartaric acid reactions on chiral Cu(hkl) R&S surfaces. Measurements of enantiospecific rate constants were obtained by using curved Cu(hkl) R&S surfaces that enable kinetic measurements on hundreds of chiral surface orientations. One model uses feature vectors derived from generalized coordination numbers to capture the local structure around Cu atoms exposed by the Cu(hkl) R&S surfaces. The second model introduces the use of chiral cubic harmonic functions to capture the symmetry constraints of the face-centered cubic Cu structure. The model using 58 generalized coordination numbers has a fitting error similar to that of the model using only 5 cubic harmonic functions. The two models predict maxima in the enantiospecificity on surfaces with very similar surface orientations. The models developed in this work are applicable for any enantiospecific reaction happening on any chiral material with a cubic lattice structure, opening the way to understanding the surface structure sensitivity of the enantiospecific reaction kinetics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Window convolution of the galaxy clustering bispectrum

In galaxy survey analysis, the observed clustering statistics do not directly match theoretical predictions but rather have been processed by a window function that arises from the survey geometry including the sky footprint, redshift-dependent background number density and systematic weights. While window convolution of the power spectrum is well studied, for the bispectrum with a larger number of degrees of freedom, it poses a significant numerical and computational challenge. In this work, we consider the effect of the survey window in the tripolar spherical harmonic decomposition of the bispectrum and lay down a formal procedure for their convolution via a series expansion of configuration-space three-point correlation functions, which was first proposed by Sugiyama et al. (2019). We then provide a linear algebra formulation of the full window convolution, where an unwindowed bispectrum model vector can be directly premultiplied by a window matrix specific to each survey geometry. To validate the pipeline, we focus on the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) luminous red galaxy (LRG) sample in the South Galactic Cap (SGC) in the redshift bin 0.4 ≤ z ≤ 0.6. We first perform convergence checks on the measurement of the window function from discrete random catalogues, and then investigate the convergence of the window convolution series expansion truncated at a finite of number of terms as well as the performance of the window matrix. This work highlights the differences in window convolution between the power spectrum and bispectrum, and provides a streamlined pipeline for the latter for current surveys such as DESI and the Euclid mission.

79 ASTRONOMY AND ASTROPHYSICS↗

Myriad World Baseline: Global Geodemographic Estimates

The LandScan Myriad World Baseline (MWB) method produces global, residential (nighttime/home-location) gridded geodemographic estimates based on 5-year age/gender cohorts—at 30-arcsecond (≈1 km) resolution. MWB is designed to fill gaps where detailed, georeferenced survey data (e.g., Demographic and Health Surveys (DHS)) are missing or outdated, and to provide a baseline that can support human security analysis, including consequence assessment, “patterns of life” modeling, and scenario-based population futures. MWB’s workflow spatializes household-level age/gender characteristics from the GLOPOP-S dataset by conflating household and gridded expected relative wealth adapted from Global Gridded Relative Deprivation Index (GRDI), then adjusts them to a target year of interest. Age/gender estimates are then applied to harmonize lowest-administrative-level statistics with LandScan residential counts, yielding final geodemographic estimates. Two validation case studies are presented: Ghana (2021) and Tokyo/Kanagawa, Japan (2020), illustrating spatial variability in demographic cohorts and comparing MWB outputs to official gridded statistics. Results show close overall alignment relative to validation criteria including population pyramids and age-dependency ratios.

Tuccillo, Joe [ORNL] (ORCID:0000000259300943)↗

Machine learned potential for high-throughput phonon calculations of metal—organic frameworks

Metal–organic frameworks (MOFs) are highly porous and versatile materials studied extensively for applications such as carbon capture and water harvesting. However, computing phonon-mediated properties in MOFs, like thermal expansion and mechanical stability, remains challenging due to the large number of atoms per unit cell, making traditional Density Functional Theory (DFT) methods impractical for high-throughput screening. Recent advances in machine learning potentials have led to foundation atomistic models, such as MACE-MP-0, that accurately predict equilibrium structures but struggle with phonon properties of MOFs. In this work, we developed a workflow for computing phonons in MOFs within the quasi-harmonic approximation with a fine-tuned MACE model, MACE-MP-MOF0. The model was trained on a curated dataset of 127 representative and diverse MOFs. The fine-tuned MACE-MP-MOF0 improves the accuracy of phonon density of states and corrects the imaginary phonon modes of MACE-MP-0, enabling high-throughput phonon calculations with state-of-the-art precision. The model successfully predicts thermal expansion and bulk moduli in agreement with DFT and experimental data for several well-known MOFs. These results highlight the potential of MACE-MP-MOF0 in guiding MOF design for applications in energy storage and thermoelectrics.

Elena, Alin Marin↗

The Unified Phenotype Ontology : a framework for cross-species integrative phenomics

Phenotypic data are critical for understanding biological mechanisms and consequences of genomic variation, and are pivotal for clinical use cases such as disease diagnostics and treatment development. For over a century, vast quantities of phenotype data have been collected in many different contexts covering a variety of organisms. The emerging field of phenomics focuses on integrating and interpreting these data to inform biological hypotheses. A major impediment in phenomics is the wide range of distinct and disconnected approaches to recording the observable characteristics of an organism. Phenotype data are collected and curated using free text, single terms or combinations of terms, using multiple vocabularies, terminologies, or ontologies. Integrating these heterogeneous and often siloed data enables the application of biological knowledge both within and across species. Existing integration efforts are typically limited to mappings between pairs of terminologies; a generic knowledge representation that captures the full range of cross-species phenomics data is much needed. We have developed the Unified Phenotype Ontology (uPheno) framework, a community effort to provide an integration layer over domain-specific phenotype ontologies, as a single, unified, logical representation. uPheno comprises (1) a system for consistent computational definition of phenotype terms using ontology design patterns, maintained as a community library; (2) a hierarchical vocabulary of species-neutral phenotype terms under which their species-specific counterparts are grouped; and (3) mapping tables between species-specific ontologies. This harmonized representation supports use cases such as cross-species integration of genotype-phenotype associations from different organisms and cross-species informed variant prioritization.

59 BASIC BIOLOGICAL SCIENCES↗