Search NASA⌕ Search

SEARCH · Search NASA

Results for “data gap”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Mesoscale Cellular Convection Detection and Classification Using Convolutional Neural Networks: Insights From Long-Term Observations at ARM Eastern North Atlantic Site

Marine boundary layer clouds are crucial in Earth's climate system. They frequently manifest as closed or open cell mesoscale cellular convection (MCC). MCC clouds are challenging to represent accurately in current climate models, highlighting the need for detailed observational data sets and in-depth analyses. This study utilizes over 8 years of observations from the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) User Facility Eastern North Atlantic (ENA) site at Graciosa Island, Azores, to investigate these clouds. We first apply a convolutional neural network with a U-Net architecture to classify open and closed cells, marking the first application of such an approach for automatically detecting MCC patterns from ground-based radar measurements. This method addresses some observational gaps in satellite data related to low temporal resolution, nighttime challenges, and limited vertical structure capture. The analysis of the MCC cases shows clear differences between closed and open MCCs: Closed MCC clouds are characterized by lower cloud tops and bases, shallower cloud geometrical depth, weaker horizontal wind speeds, stronger atmospheric stability, and a more homogeneous liquid water path than open MCCs. Finally, we demonstrate two potential applications of our radar-based MCC classifications: (a) facilitating the investigation of aerosol-cloud interactions and (b) exploring meteorological factors along with MCC's evolution by integrating satellite imagery and back-trajectory analysis. The identified MCC cases offer a valuable resource for the scientific community to study MCC processes further and improve climate model accuracy.

54 ENVIRONMENTAL SCIENCES↗

Microbiome data management in action workshop: Atlanta, GA, USA, June 12–13, 2024

Microbiome research is revolutionizing human and environmental health, but the value and reuse of microbiome data are significantly hampered by the limited development and adoption of data standards. While several ongoing efforts are aimed at improving microbiome data management, significant gaps still remain in terms of defining and promoting adoption of consensus standards for these datasets. The Strengthening the Organization and Reporting of Microbiome Studies (STORMS) guidelines for human microbiome research have been endorsed and successfully utilized by many research organizations, publishers, and funding agencies, and have been recognized as a consensus community standard. No equivalent effort has occurred for environmental, synthetic, and non-human host-associated microbiomes. To address this growing need within the microbiome research community, we convened the Microbiome Data Management in Action Workshop (June 12–13, 2024, in Atlanta, GA, USA), to bring together key decision makers in microbiome science including researchers, publishers, funders, and data repositories. The 50 attendees, representing the diverse and interdisciplinary nature of microbiome research, discussed recent progress and challenges, and brainstormed actionable recommendations and paths forward for coordinated environmental microbiome data management and the modifications necessary for the STORMS guidelines to be applied to environmental, non-human host, and synthetic microbiomes. The outcomes of this workshop will form the basis of a formalized data management roadmap to be implemented across the field. These best practices will drive scientific innovation now and in years to come as these data continue to be used not only in targeted reanalyses but in large-scale models and machine learning efforts.

54 ENVIRONMENTAL SCIENCES↗

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

Demonstrate new plasticity models for doped UO 2 that capture dislocation mechanisms

In light water reactors, fuel vendors are investigating the use of dopants to modify the properties of UO 2 pellets, with the goal of improving pellet-cladding mechanical interactions during operation. Dopants are expected to ‘soften’ the pellets; that is, the doped pellets have higher plastic deformation than conventional UO 2 . This leads to a reduction in the severity of mechanical pellet-cladding interactions, helping to reduce the hoop strain on the cladding. By minimizing the strain exerted by the pellet on the cladding, it is anticipated that cladding performance under accident conditions can be enhanced (i.e., lowering the risk of burst during a LOCA). Dopants such as chromium (Cr) promote grain growth during pellet fabrication, leading to larger grains; therefore, understanding the link between chemistry, microstructure and mechanical deformation (enhanced creep rates) behavior of UO 2 is critical to helping operators further substantiate the benefits of doping UO 2 . Historically, the nuclear energy industry has relied on empirical models to make assessments of performance. Compared to empirical models, mechanistic physics-based models provide benefits, such as, fewer data points for validation and better extrapolation where experimental data is scarce or non-existent. In this report, Bayesian inference techniques have been applied to a previously developed lower length-scale-informed diffusional creep model. The objective is to i) infer lower-length-scale parameter distributions from available experiment and then ii) determine the uncertainties in the measurable quantity (in this case creep rates) after propagating the inferred lower length scale parameter uncertainties. The approach requires many evaluations of the model, which becomes computationally insurmountable; therefore, a neural-network model is trained to data obtained by sampling the full model over the most important parameters. This neural-network is then used in the Bayesian inference approach to determine probability distributions in the parameter values that represent the uncertainty in the model given what is known from the experiments (posterior). A significant reduction compared to conservative initial (prior) uncertainties is achieved through inference against the experimental data, demonstrating the efficacy of this approach. Furthermore, by accounting for uncertainties in the experimental conditions and sample non-stoichiometry, it is possible to resolve apparent discrepancies in experimental measurements within a self-consistent grain boundary (Coble) creep model that is sensitive to chemistry. This work has been written up and submitted to Nuclear Technology for a special issue on accelerated fuel qualification (AFQ). This uncertainty quantification (UQ) work not only improves the diffusional model, while accounting for uncertainty, but also establishes a framework which can readily be applied to the mechanistic models of dislocation deformation developed in this study. The most likely values from the Bayesian analysis are incorporated into our UO 2 diffusional creep model and a lower length scale-informed irradiation UO 2 creep mechanistic model to generate a dataset. This dataset has been provided to our INL collaborators for training an artificial neural network surrogate model, which will be implemented in the BISON fuel performance code to assess how the results differ from those currently obtained using a fully empirical model and that of using the nominal (uncalibrated) atomic scale parameters in our mechanistic model. Plastic deformation (creep and glide) in UO 2 is a complex phenomenon, governed by multiple underlying processes such as local defect concentrations, applied stresses, and microstructural characteristics. Consequently, there is a need for a meso-scale model with polycrystalline resolution capable of extrapolating to large grain sizes applicable to doped UO 2 , where data is limited and the model can help bridge the knowledge gap. By integrating atomistic data into the polycrystal LApx code, it becomes possible to predict dislocation climb and glide plasticity that simple analytical models cannot accurately represent. The application of atomic-scale data within LApx demonstrated the importance of climb and glide mechanisms in reproducing high-stress UO 2 behavior. Behaviors such as this are crucial to capture and implement in BISON, as parts of the fuel pellet can reach temperatures where glide can occur before pellet cracking. This model which captures dislocation based mechanisms for UO 2 is then used to stand up the doped model accounting for larger grain sizes. It was found that larger grain sizes can lead to enhanced deformation rates in the glide regime, and therefore can help with the pellet cladding mechanical interaction. Therefore if the fuel pellet reaches conditions (stress/temperature) where glide is active, the enhanced creep rates for larger grains in the glide regime (doped UO 2 ) can help with pellet cladding mechanical interactions. Plastic deformation in UO 2 involves multiple mechanisms, including diffusional creep, dislocation climb, and glide. This milestone contains two parts: (1) UQ of a pre-existing lower length scale informed mechanistic diffusional creep model, and (2) development of a new LApx based model for dislocation-mediated creep mechanisms in UO 2 , with application to large-grain doped UO 2 .

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Nuclear excitation functions for medical isotope production: Targeted radionuclide therapy via nat IR$(d, x)$ 193m Pt

193m Pt is an Auger emitting radionuclide which may have therapeutic potential, particularly when labeled to the chemotherapeutic drug cisplatin. One challenge to broader explorations of its clinical potential is the need for production routes with high specific activity. As part of a larger campaign to address gaps in reaction data for emerging medical radionuclides, this work seeks to characterize the nat Ir(d,x) reactions as a potential production pathway for 193m Pt. A stacked target irradiation, consisting of natural iridium, iron, nickel, and copper foils, was performed using a 33 MeV deuteron beam at the Lawrence Berkeley National Laboratory 88-Inch Cyclotron. This measurement, along with previous experimental data, suggests an energy window between 11 to 18 MeV to maximize the production and radiopurity of 193m Pt. This experiment has yielded cross sections for 43 channels of deuteron-induced reactions from threshold to 30 MeV, including the first experimental results of nat Ir(d,x) 188m1+g,190m1+g Ir (cumulative), nat Ni(d,x) 56,57,58 m,58g Co (independent), nat Cu(d,x) 61 Co (cumulative) and nat Fe(d,x) 53 Fe, 48 V (cumulative). The results were compared with literature data, the TENDL-2023 database, and default theoretical calculations from the TALYS-2.04, CoH-3.6.0, EMPIRE 3.2.3, and ALICE-2020 reaction modeling codes. Here, this work presents another example of the lack of predictive capabilities for this set of modern nuclear-reaction modeling codes, and highlights the unsatisfactory modeling of experimental cross sections. Experimental data are important to improve the codes in general, and new experimental results can be used to improve the models. Finally, this measurement has revealed the need for an updated evaluation of the nat Cu(d,x) 63 Zn deuteron monitor reaction.

193mPt↗

Investigating the opioid epidemic across the United States: Associations between county-level characteristics and overdose mortality

The opioid crisis remains a critical public health challenge in the United States. Despite national efforts that reduced opioid prescribing by nearly 44% between 2011 and 2021, opioid overdose deaths more than tripled during the same period. This alarming trend reflects a major shift in the crisis, with illegal opioids now driving the majority of overdose deaths instead of prescription opioids. Although supply-side factors fueling this transition have been widely studied, the structural and community-level conditions that shape overdose mortality are less well understood. To help address this gap, this study has three primary objectives: (1) overcome structural gaps in national data to construct a complete nationwide county-level dataset from 2010 to 2022; (2) using data analysis, identify and investigate spatiotemporal anomalies in overdose mortality; and (3) using two machine-learning models, quantify the importance of thirteen social vulnerability variables in predicting overdose mortality. Our results identify unemployment and limited vehicle access as key county-level predictors of overdose mortality. Higher levels of these vulnerabilities are associated with elevated mortality, whereas lower levels are associated with reduced mortality. These findings highlight factors that may be relevant for public health planning and policy prioritization within the context of the opioid crisis.

Anomaly analysis↗

Deimos: HALEU TRISO Heated Critical Experiment Data

Deimos was the first critical experiment using high-assay low-enriched uranium (HALEU) TRistructural ISOtropic (TRISO) fuel in over 40 years. HALEU TRISO is the desired fuel form for many of the advanced reactor designs in development; however, very little experimental data are available for this fuel type. Deimos was designed to utilize existing HALEU TRISO fuel in a large graphite moderator to obtain nuclear and reactor physics data to fill the gaps surrounding this fuel type and enrichment. In addition to cold critical data, three separate heated experiments were conducted to measure the temperature reactivity coefficient for this type of system. These measured coefficients were then compared to simulated coefficients to a first level order of fidelity. This comparison showed very good agreement for the experiment where only the inner core was heated and good agreement for the other two configurations, which included heating portions of the outer core. Less agreement when the outer core was heated is attributed to potential heating in the beryllium reflector, which has a positive temperature reactivity coefficient and was unaccounted for in the first-order models. Future heated experiments with Deimos will include temperature monitoring of the beryllium reflector to account for beryllium heating in the simulations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Novel, active, and uncultured hydrocarbon-degrading microbes in the ocean

ABSTRACT Given the vast quantity of oil and gas input to the marine environment annually, hydrocarbon degradation by marine microorganisms is an essential ecosystem service. Linkages between taxonomy and hydrocarbon degradation capabilities are largely based on cultivation studies, leaving a knowledge gap regarding the intrinsic ability of uncultured marine microbes to degrade hydrocarbons. To address this knowledge gap, metagenomic sequence data from the Deepwater Horizon (DWH) oil spill deep-sea plume was assembled to which metagenomic and metatranscriptomic reads were mapped. Assembly and binning produced new DWH metagenome-assembled genomes that were evaluated along with their close relatives, all of which are from the marine environment (38 total). These analyses revealed globally distributed hydrocarbon-degrading microbes with clade-specific substrate degradation potentials that have not been reported previously. For example, methane oxidation capabilities were identified in all Cycloclasticus . Furthermore, all Bermanella encoded and expressed genes for non-gaseous n -alkane degradation; however, DWH Bermanella encoded alkane hydroxylase, not alkane 1-monooxygenase. All but one previously unrecognized DWH plume member in the SAR324 and UBA11654 have the capacity for aromatic hydrocarbon degradation. In contrast, Colwellia were diverse in the hydrocarbon substrates they could degrade. All clades encoded nutrient acquisition strategies and response to cold temperatures, while sensory and acquisition capabilities were clade specific. These novel insights regarding hydrocarbon degradation by uncultured planktonic microbes provides missing data, allowing for better prediction of the fate of oil and gas when hydrocarbons are input to the ocean, leading to a greater understanding of the ecological consequences to the marine environment. IMPORTANCE Microbial degradation of hydrocarbons is a critically important process promoting ecosystem health, yet much of what is known about this process is based on physiological experiments with a few hydrocarbon substrates and cultured microbes. Thus, the ability to degrade the diversity of hydrocarbons that comprise oil and gas by microbes in the environment, particularly in the ocean, is not well characterized. Therefore, this study aimed to utilize non-cultivation-based ‘omics data to explore novel genomes of uncultured marine microbes involved in degradation of oil and gas. Analyses of newly assembled metagenomic data and previously existing genomes from other marine data sets, with metagenomic and metatranscriptomic read recruitment, revealed globally distributed hydrocarbon-degrading marine microbes with clade-specific substrate degradation potentials that have not been previously reported. This new understanding of oil and gas degradation by uncultured marine microbes suggested that the global ocean harbors a diversity of hydrocarbon-degrading bacteria, which can act as primary agents regulating ecosystem health.

Howe, Kathryn L.↗

Investigations of Relative Stability and Distribution of Mo Species in Pore Space of Mo/ZSM-5 Systems

Reduced forms of molybdenum (Mo) carbides anchoring in pore space of zeolites are considered as activated catalysts for the methane dehydroaromatization reaction. However, the atomic structure of the active sites and their stability under reactive conditions are still not clearly understood. Herein, we employ a synergistic theoretical and experimental approach to investigate this research gap. The Raman data shows the existence of the distributed binuclear and mononuclear Mo oxides in the pores that are corroborated by the hydrogen temperature-programmed reduction experiments. The binding free energies of mono- and bi-nuclear Mo carbides in straight and sinusoidal channels and their intersections as a function of temperature and anchoring sites have been determined. The catalysts with Mo coordinated at two anchoring sites are more stable compared to those located at the rings with one Al substitution. The activation energies of the catalysts migration to adjacent rings were estimated as a function of temperature and kinetics parameters of such migrations were determined. The migration is facilitated if there is one anchoring site shared between two rings. This process can contribute to an agglomeration mechanism leading to loss of catalytic activity.

Myshakin, Evgeniy↗

Decision Tree for Variable Selection vs. Impact on Durability for Biomass and Biochar Burial Pathways [Slides]

Quantifying durability for lower-TRL BiCRS pathways has been challenging as limited data are available from real-world projects and long-term experiments, resulting in an overall lack of scientific consensus. We develop a decision tree that aims to summarize the current scientific understanding and state-of-the-art project experience. The decision tree can be used to (1) guide the selection of key variables and evaluate their relative impact on durability, (2) identify data and knowledge gaps for future research.

09 BIOMASS FUELS↗

Queued Up: 2025 Edition – Characteristics of Power Plants Seeking Transmission Interconnection As of the End of 2024 [Slides]

Electric transmission system operators (ISOs, RTOs, or utilities) require proposed power plants seeking to connect to the transmission grid to undergo a series of impact studies before they can be built. This process establishes what new transmission equipment or upgrades may be needed before a project can connect to the system and assigns the costs of that equipment. The lists of projects in this process are known as “interconnection queues”. In collaboration with interconnection.fyi, Berkeley Lab compiled, aggregated, and cleaned interconnection queue data from >50 transmission grid operators (7 ISO/RTOs and 49 non-ISO balancing areas), which collectively represent ~97% of currently installed U.S. electric generating capacity. The dataset includes requests submitted to queues through the end of 2024, and only includes requests seeking to connect to the transmission grid (not distribution-connected or behind-the-meter projects). The files below include both a PDF report and an Excel data file. The PDF report analyzes interconnection data and metrics through the end of 2024. The Excel data file includes (a) the full project-level interconnection queue dataset through 2024, (b) a codebook (data dictionary) describing each data field, and (c) 35 additional tabs featuring tables summarizing a range of interconnection metrics. Key highlights from the Queued Up: 2025 Edition (featuring data through 2024) include: • As of the end of 2024, there were ~10,300 projects actively seeking grid interconnection in the U.S., representing 1,400 GW of generation and approximately 890 GW of storage. • Historic withdrawal rates alongside relatively fewer new requests resulted in a 12% decrease in total active queue volume compared to the prior year. • Active natural gas capacity (136 GW, +72% year-over-year) increased in 2024, while solar (956 GW, -12%), storage (890 GW, -13%), and wind (271 GW, -26%) capacity decreased. • 408 GW of capacity already has a draft or executed interconnection agreement (IA) but has not yet reached commercial operations. • The time projects spend in queues before reaching COD is increasing. For the regions with available data, the median duration from IR to COD has doubled from <2 years for projects built in 2000-2007 to over 4 years for those built in 2018-2024. • Ultimately, most of this proposed capacity will not be built. Only 13% of capacity that submitted interconnection requests from 2000-2019 had reached commercial operations by the end of 2024; 77% of that capacity had been withdrawn and 10% was still active. • FERC Order 2023 and various other reforms are being implemented. These are important measures to reduce interconnection bottlenecks and enhance grid system reliability, but it is too early to measure and assess their full impact. • New additions for the 2025 edition include: (a) additional detail on data processing and gaps; (b) updates on interconnection reforms; (c) new analysis on interconnection agreements, and more.

24 POWER TRANSMISSION AND DISTRIBUTION↗

U.S. Pacific Coast Workshop Report on Preconstruction Research Recommendations (U.S. Offshore Wind Synthesis of Environmental Effects Research (SEER) Project)

In May 2022, the U.S. Offshore Wind Synthesis of Environmental Effects Research (SEER) project team hosted a stakeholder workshop focused on preconstruction (baseline) research needs for potential floating offshore wind (OSW) energy development on the U.S. Pacific Coast, including California, Oregon, and Washington. Prior to the workshop, the SEER team developed a set of initial synthesized research recommendations that were identified based on a review of relevant, publicly available resources and with advisory group input. The workshop covered three marine life breakout groups on subsequent days to discuss research recommendations related to 1) marine mammals and sea turtles, 2) fish and invertebrates, and 3) birds and bats. As part of the workshop, over a hundred participants from the public and private sectors provided feedback on various aspects of the initial research recommendations, including associated data and knowledge gaps, benefits/limitations of available methods and technologies, and technological advancements or infrastructure needed to address the recommendation. Approximately 1,000 total comments were received on the workshop MURAL boards and were synthesized in this report. Based on workshop feedback, SEER developed a final database of over 500 specific research recommendations based on more than 40 resources. In Fall 2022, the full database and a tool with updated synthesized research recommendations were disseminated on Tethys (https://tethys.pnnl.gov) to assist with informing future funding opportunities and research programming. There is a continued need to improve awareness of the potential environmental effects, monitoring technologies, and management strategies for floating OSW energy development on the U.S. Pacific Coast. Coordination of these activities will require the sustained involvement of multiple stakeholders from across sectors. Beyond the baseline considerations discussed in this workshop, future state-of-the-science activities should be planned to consider research needs across wind energy life cycle phases for all relevant wildlife taxa and associated habitat and ecosystem processes.

17 WIND ENERGY↗

High Fidelity Simulations of Air-Cooled Reactor Cavity Cooling System

High Temperature Gas Reactor (HTGR) designs incorporate passive safety systems (e.g., the Reactor Cavity Cooling System [RCCS]) that utilize natural principles to manage heat dissipation from the reactor pressure vessel (RPV) during accidents or routine shutdowns. The industry community is experiencing a pressing need for advanced simulation tools that can accurately assess the performance of these types of systems. In the literature, a knowledge gap exists concerning high-fidelity data for the RCCS, and this gap is one of the areas of focus of the present study. This research focuses on a specific RCCS designed for General Atomics' Modular High-Temperature Gas Reactor (GA-MHTGR). Experimental studies on a scaled version (can be seen in Figure 1) of the air-cooled RCCS used in GA-MHTGR were conducted by the University of Wisconsin-Madison (UW-Madison). This work contributes to a broader initiative aimed at establishing a numerical benchmark based on the UW-Madison experiments.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Evaluation of LLVM Flang for Production HPC Applications and Modern Fortran Features

In 2025, LLVM released its first Flang Fortran compiler version considered ready for widespread evaluation. We know of no published assessment of Flang compiling a workload- derived portfolio of high-performance computing (HPC) applications. We address this gap using workload data from the National Energy Research Scientific Computing Center (NERSC), which supports more than 10,000 scientists on approximately 1,000 projects. The NERSC workload analyses identify many Fortran components in heavily used applications. We selected 10 such packages with available source code. We compiled them with Flang 22.1.3 on NERSC’s Perlmutter system. Six compiled without code modifications, though some required build-system changes. Three compiled after minor source edits, mostly to address Fortran standard violations. One built only without OpenMP enabled. We evaluated seven additional packages selected for their use of, or enablement of, standard Fortran parallel features: multi-image execution and do concurrent. Six such codes compiled with most or all unit tests passing.

Rasmussen, Katherine↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

Streaming Data in HPC Workflows Using ADIOS

The “IO Wall” problem, in which the gap between computation rate and data access rate grows continuously, poses significant problems to scientific workflows which have traditionally relied upon using the filesystem for intermediate storage between workflow stages. One way to avoid this problem in scientific workflows is to stream data directly from producers to consumers and avoiding storage entirely. However, the manner in which this is accomplished is key to both performance and usability. This paper presents the Sustainable Staging Transport, an approach which allows direct streaming between traditional file writers and readers with few application changes. SST is an ADIOS “engine”, accessible via standard ADIOS APIs, and because ADIOS allows engines to be chosen at run-time, many existing file-oriented ADIOS workflows can utilize SST for direct application-to-application communication without any source code changes. This paper describes the design of SST and presents performance results from various applications that use SST, for feeding model training with simulation data with substantially higher bandwidth than the theoretical limits of Frontier’s file system, for strong coupling of separately developed applications for multiphysics multiscale simulation, or for in situ analysis and visualization of data to complete all data processing shortly after the simulation finishes.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X↗

Data From: "Warming and snow loss increase reliance on old groundwater in a Colorado River headwater"

This repository contains the data and code associated with the paper titled "Warming and snow loss increase reliance on old groundwater in a Colorado River headwater," published in Nature Geoscience, 2026. This study seeks to answer how various ages of groundwater interact with mountainous streamflow in mountainous headwaters such as the East River. It includes various model-data processing scripts, primarily for ParFlow-CLM analysis of simulated water years 2015-2021, and two numerical warming experiments (+2.5 and +4.0 degrees C), including run scripts, forcing scripts, and post-processing, as well as comparison to observation datasets, detailed below. This data requires the use of R (.r, .rmd), Python (.py), Jupyter Notebook or Jupyter Lab (.ipynb), ParFLOW-CLM, EcoSLIM. Further information on the use of all file formats mentioned below (e.g. .tff. .nc) are provided within the associated scripts and directory where the files are located. Contents & Usage ASO/: ​​Contains the bash and python scripts used to convert airborne snow observatory (ASO) data (ASO, 2023) in various data formats (georeferenced tiff file, NetCDF, UTM, and to latitude/longitude) then regrided to the ParFlow equivalent grid. Output data are in regrid_regll_data.zip and subsequently visualized and analyzed in plot_and_compare.py for Supplementary Figures A14 and A15. The wksht_ASO_comparison.xlsx spreadsheet is used to calculate the data for Supplementary Figure A16. EcoSLIM/: Contains the scripts and input files to run the EcoSLIM particle tracking simulations (/run_scripts) and the post-processing python script (/plot_scripts/eco_agedist_plots.ipynb). Jasechko et al./: Contains the jupyter notebook (Extract_Elevation.ipynb) to determine the outlet elevations of the 260 watersheds used in Jasechko et al. (2016), and the corresponding table, Table_S1_Watersheds_alt.csv. Used to create Supplementary Information Figure A2. PLM_Wells/: Contains the QA/QC-ed groundwater level time series of the PLM-1 and PLM-6 Monitoring Wells from Faybishenko et al. (2023), reformatted to water years used for Supplementary Figures A19 and and A20. ParFlow/: Contains the input files and run scripts to run ParFlow-CLM (/run_scripts), the python and tool command language (Tcl) scripts to create and distribute the ParFlow forcing simulation files (/forcing), and various scripts and intermediary files to analyze the model outputs (/post_process). SQUIRE/: Contains the processing scripts and intermediary files for the Surface QUantitatIve pRecipitation Estimation (SQUIRE) data (Grover, 2023) used to generate Supplementary Figure A18. USGS_Streamflow/: Contains the raw and gap-filled United States Geological Survey streamflow data (U.S. Geological Survey, 2026) used at the Almont station (site number 09112500). Gap-filling is performed in the R script with data from the Taylor station (site number 09110000). (/USGS_09112500_EAST_RIVER_AT_ALMONT_GAP_FILLED/code_almont_streamflow_gap_fill.Rmd). discharge/: Contains the gap-filled discharge data at the Watershed Function SFA East River pumphouse site (Newcomer et al., 2022) used to generate Supplementary Figure A13 and to compute hourly Nash-Sutcliffe model efficiency coefficients (NSE) in Table A4. snotel_and_flux_tower/: Contains the snow telemetry data (U.S. Department of Agriculture, 2024) from the Butte (site ID 380) and Schofield (site ID 737) stations, reformatted by water year, accessed with the snotelr R package. Used to create Supplementary Figure A17. Also contains the flux tower observational data (FluxTower_Pumphouse_ESS-DIVE.ET_only.h.txt) from Ryken et al. (2022) and sap flux transpiration data (MaxB_Transpiration_5Sites.daily_sums.h.txt) from Ryken (2021), used to create Supplementary Figures A22 and A23, respectively. Raw EcoSLIM model outputs are in excess of 24TB, and are stored on National Energy Research Scientific Computing Center (NERSC) and publicly available via the external link provided in the paper.

atmospheric warming↗