Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Waveform resampling with LMN method

In this article, resampling is a common technique applied in digital signal processing. Based on the Fast Fourier Transformation (FFT), we apply an optimization called here the LMN method to achieve fast and robust re-sampling. In addition to performance comparisons with some other popular methods, we illustrate the effectiveness of this LMN method in a particle physics experiment: re-sampling of waveforms from Liquid Argon Time Projection Chambers.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Improvements to the characterization of Agfa x-ray film for use on opacity spectroscopy diagnostics

The National Ignition Facility uses a soft x-ray opacity spectrometer for x-ray spectral imaging in high-energy-density experiments. The increased demand for a better spectral resolution prompted the investigation into the Agfa D4 film. Characterization is already under way for the film. A Manson x-ray source using six different anodes was used to expose film to the linear optical density (OD) region. This is a continuation of the previous work, and the updated analysis process is communicated here. The identified uncertainties have been reduced with the updated steps that improve the results of the characterization process. In conclusion, when the Stanford Synchrotron Radiation Lightsource Beamline 16-2 was operational, the film was characterized at that source. Its beam offered a higher fluency with a lower exposure time needed to reach saturation. Results for both sources are compared in this paper.

47 OTHER INSTRUMENTATION↗

Taylor approximation variance reduction for approximation errors in PDE-constrained Bayesian inverse problems

In numerous applications, surrogate models are used as a replacement for accurate parameter-to-observable mappings when solving large-scale inverse problems governed by partial differential equations (PDEs). The surrogate model may be a computationally cheaper alternative to the accurate parameter-to-observable mappings and/or may ignore additional unknowns or sources of uncertainty. The Bayesian approximation error (BAE) approach provides a means to account for the induced uncertainties and approximation errors, i.e. the errors between the accurate parameter-to-observable mapping and the surrogate. The statistics of these errors are, however, in general unknown a priori, and are thus calculated using Monte Carlo sampling. Although the sampling is typically carried out offline, i.e. before considering the data, the process can still represent a computational bottleneck. In this work, we develop a scalable computational approach for reducing the costs associated with the sampling stage of the BAE approach. Specifically, we consider the Taylor expansion of the accurate and surrogate forward models with respect to the uncertain parameter fields either as a control variate for variance reduction or as a means to directly and efficiently approximate the mean and covariance of the approximation errors. We propose efficient methods for evaluating the expressions for the mean and covariance of the Taylor approximations based on linear(-ized) PDE solves. Furthermore, the proposed approach is independent of the dimension of the uncertain parameter, depending instead on the intrinsic dimension of the data, ensuring scalability to high-dimensional problems. The potential benefits of the proposed approach are demonstrated for two high-dimensional inverse problems governed by PDE examples, namely for the estimation of a distributed Robin boundary coefficient in a linear diffusion problem, and for a coefficient estimation problem governed by a nonlinear diffusion problem.

Bayesian approximation error↗

HERO WEC V1.0 2024 - WEC-Sim Detailed Simulation Runs and Summary Data

This dataset includes results from simulations of NREL's hydraulic and electric reverse osmosis wave energy converter (HEREO WEC). Simulation runs include 135 wave cases that were based on the updated WEC-Sim model, which is linked below. The data represented in this repository is based on an updated WEC-Sim model using laboratory data to tune and refine the original WEC-Sim model for the V1.0 HERO WEC. The 135 wave cases represent waves with the following wave height and wave period ranges: - Significant Wave Height: 0.25 - 3.75m in 0.25m increments - Wave Period: 5 - 13 sec in 1 sec increments Each run was simulated using a Pierson-Moskowitz irregular wave spectrum with a 100 second ramp time, a total simulation time of 3,100 seconds, and a simulation time-step of 0.005s. A reference table has been included to map each multi condition run (MCR) case with each wave condition. Summary data set includes a spreadsheet and image files with matrices that are associated with data from simulation runs. All matrices cover the same significant wave height and wave periods from the simulation runs, in the same increments. The following matrices are included: - Power Abs: The average absorbed power from the WEC (calculated from anchor reaction force and heave velocity) - Power Hyd: The average hydraulic power output at pump (calculated from pump output flow and pressure) - Power - Hyd ROi: The average hydraulic power measured at the RO system inlet (calculated from RO system pressure and flow (pre-accumulator)) - Flow - Pump out: The average flowrate measured at the pump outlet - Flow - Perm: The average permeate (clean water) production - Flow - RO (pre): The average flowrate measured at the inlet of the RO system before the accumulators - Flow - RO (post): The average flowrate measured after the accumulator bank in the RO system - Pressure - RO: The average pressure measured at the inlet of the RO system This data set has been developed by the National Renewable Energy Laboratory, operated by the Alliance for Sustainable Energy, LLC, for the U.S. Department of Energy (DOE) under Contract No. DE-AC36-08GO28308. Funding provided by the U.S. Department of Energy Office of Energy Efficiency and Renewable Energy Water Power Technologies Office.

16 TIDAL AND WAVE POWER↗

Renewable Energy Potential Model: Priority Geothermal Leasing Areas ReEDs Results

This dataset contains the results of a study conducted by the National Renewable Energy Laboratory (NREL) to identify potential future priority geothermal leasing areas on Bureau of Land Management (BLM) and United States Forest Service (USFS) lands. The analysis uses the Regional Energy Deployment System (ReEDS) model to evaluate geothermal resource potential under different scenarios of resource depth and technology combinations through the year 2050. The study considers geothermal resource potential, natural resource conflicts, and transmission access to categorize areas into near, mid, and far deployment priorities. The dataset includes outputs from the ReEDS model, such as geothermal capacity, generation, system costs, and emissions under various economic and technical scenarios. Favorability site data with geographic coordinates and site-specific attributes (e.g., resource favorability, land type) are also provided. Supporting resources include a technical report detailing methodologies and assumptions, along with a link to the ReEDS model GitHub repository, which requires GAMS and Python software for execution.

15 GEOTHERMAL ENERGY↗

Massive all-atom analysis of 2D materials with quantum properties (Final report)

Improvements in microscopy have enabled the acquisition of data at a scale that is difficult to process manually, making automated machine learning approaches to analyzing experimental images essential. In this project, we developed and applied machine learning (ML) workflows for atomic resolution scanning transmission electron microscopy (STEM) images. This development included improving both methodology as well as generating user-friendly codes. We developed machine learning architectures which, after training, automatically identify the location and types of defects throughout a material. We used these data to produce class-averaged images of 2D atomic coordinates with up to 0.3 pm precision, uncovering the structure and oscillations of long-range strain fields around point defects in WSe 2-2x Te 2x . We also resolved a long-standing problem in this field in the training of ML models, a lack of labeled experimental data, by developing a cycle-GAN that transformed simulated-generated labeled data into labeled data indistinguishable from experiment and therefore suitable for training. This removed the remaining parts of the ML data processing workflow where human intervention was still critical and therefore a bottleneck to working at scale. Codes have been developed and released for this full machine learning workflow. ML approaches to partially automate STEM acquisition were also developed. Finally we applied ML and other advanced data processing methods to several materials science problems in two-dimensional materials, including studying the evolution of hyperuniformity with defect concentration in WSe2, understanding phase transformations in transition metal dichalcogenides during in-situ heating in the STEM, and exploring how 2D interfaces transform from twisted into aligned structures.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Extraction of Vibration Data with Imaging

To date, the primary sensing technology used to measure the vibration response has been accelerometers and strain gages mounted directly to the structure and using either wired or, more recently, wireless telemetry. Cost issues with these sensors and the associated data acquisition systems typically limit the numbers that are deployed on in situ structures. Although there are a few structures with larger sensing counts that in some cases exceed over 1000 sensors, more typical numbers range from ten to one hundred sensors resulting in low spatial resolution when they are applied to physically large systems. When one considers that nuclear power plant structures usually have complex geometries, material properties, connectivity and boundary conditions, it is clear these current approaches to vibration measurements can only provide limited information about a system’s dynamics response characteristics. As an alternative, many non-contact measurement technologies have emerged, including point wise measurement methods such as Global Positioning System (GPS), microwave interferometry, and laser Doppler vibrometry (LDV), as well as simultaneous full-field measurement methods such as electronic speckle pattern interferometry, holography interferometry, and muon tomography, some of which can provide high spatial resolution measurements. Among these methods, digital video imaging techniques have emerged as a feasible solution for full-field vibration measurements that provide significantly more detailed dynamic response information because every pixel becomes a measurement point. Furthermore, recent advances in image processing and computer vision algorithms have been successfully used to process video data for experimental and operational modal analysis. Such full-field measurements have the potential to significantly improve many current structural assessment procedures including system identification (modal parameter estimation), structural health monitoring, load reconstruction, model validation, and model updating. Furthermore, more recent full-field imaging techniques can be accomplished with relatively low-cost, commercially-available off-the-shelf cameras. However, these measurement procedures have other limitations that must be considered such as the ability to only measure visibly accessible points on a structure and a more limited dynamic range and bandwidth than can be achieved with accelerometers or strain gages.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

CHESS 2025: Discrete-return LiDAR point clouds from NEON AOP surveys

This dataset provides Level 1 (L1) discrete-return light detection and ranging (LiDAR) point cloud data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary unclassified discrete-return LiDAR data delivered by NEON and are provided per flightline as LASzip (LAZ) 1.4 Format 6 files. Data were processed following the workflow described in the NEON L0-to-L1 Discrete Return LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022). Each record in the unclassified point clouds represents a geolocated laser target/return recorded by the LiDAR system, with values for X, Y, Z position and return intensity. All point coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Flight metadata describing flightline boundaries and positional uncertainty by point are also included. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Utah FORGE: 2024 Discrete Fracture Network Model Data

The Utah FORGE 2024 Discrete Fracture Network (DFN) Model dataset provides a set of files representing discrete fracture network modeling for the FORGE site near Milford, Utah. The dataset includes four distinct DFN model file sets, each corresponding to different time frames and modeling approaches in 2024. These models characterize both natural and induced fractures in the geothermal reservoir, which consists of crystalline granitic and metamorphic rock approximately 8,000 feet below the ground surface. The dataset includes a reference DFN model from February 2024 that incorporates planar fractures and well trajectories, as well as upscaled permeability, porosity, compressibility, and storage values on specified grids. Additionally, there are models based on new microseismic (MEQ) data from May and July 2024, including fracture planes fitted to the latest MEQ catalog datasets, tensile fractures from hydraulic stimulation, and an alternative connected DFN for modeling purposes. Coordinate data is provided in both global and local frames, with detailed instructions on the transformations used to align with principal stress orientations. The dataset also includes notes and calculation files for estimating fracture sizes and differences between various fracture sets. There are subfolders for Global Coordinates and Local Coordinates. To move from the global to the local coordinate frame, fractures and wells were a) rotated 20 degrees counterclockwise looking down about the global point (335376.400482041, 4263189.99998761, 250.093546450195) to better align with the principal stresses; and b) translated by (-335408.68, -4263010.9, 1150). Upscaled permeability values using the _XYZ suffix show directions with respect to the global XYZ coordinate frame, while those using the _IJK suffix are aligned with local coordinate frame.

15 GEOTHERMAL ENERGY↗

Two-dimensional heteronuclear single quantum coherence (HSQC) NMR spectra of lignin isolated from Populus trichocarpa residues after CELF pretreatment and CBP fermentation

Here we present a curated dataset of a series of two-dimensional heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra of lignin isolated from a woody energy crop (Populus trichocarpa) residues after co-solvent enhanced lignocellulosic fractionation (CELF) pretreatment and consolidated bioprocessing (CBP) process. The natural poplar variant GW-9947 from the Center for Bioenergy Innovation (CBI) was used. The poplar was knife milled and passed through a 1 mm sieve. The CELF pretreatment was performed in a Parr autoclave reactor with 7.5 wt % solids loading, 0.5 wt% H2SO4 as catalyst at 150°C with 15, 25 and 30 minutes, respectively. Tetrahydrofuran was added in a 1:1 mass ratio with water as the pretreatment solvent. The residues from CELF pretreatment were then subjected to CBP using the bacterium C. thermocellum DSM 1313. CBP fermentations were performed at 60 °C in a shaker at 50 grams/L solids loadings. Lignin was isolated from the pretreated samples after ball-milling in a porcelain jar with ceramic balls via Retsch PM 200 at 580 rpm for 2.5 h followed by enzymatic hydrolysis in acetate buffer (pH 4.8, 50 °C) for 48 h. The lignin samples were characterized using 13C–1H HSQC experiments which were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus. A standard Bruker pulse sequence was used on a Prodigy platform cryoprobe. The dry lignin samples were dissolved in deuterated dimethylsulfoxide for HSQC experiments. The spectra were acquired under the following acquisition conditions: 210 ppm spectral width in F1 (13C) dimension with 256 data points and 11 ppm spectral width in F2 (1H) dimension with 1024 data points, a 90° pulse, a one bond C–H coupling constant of 145 Hz, a 1.0 s pulse delay, and 64 scans. All the data was processed using the TopSpin 3.6 software (Bruker BioSpin). The NMR spectra provides structural characteristics information about lignin remaining in solids after CELF (150 °C with 15, 25 and 30 minutes) process and C. thermocellum CBP.

Lignin structure, HSQC, poplar, CELF, CBP, CBI↗

Utah FORGE: Geochemical Data for Cold Groundwaters and Produced Geothermal Fluids

Geochemical data for cold groundwaters and produced geothermal fluids around the Utah FORGE site. The data is compiled into four tables in the attached Excel File. Table 1 is a compilation of compositions (anions, cations, weak acids, oxygen, hydrogen, and carbon isotopes) for cold groundwaters and produced geothermal waters in the Milford valley, Utah. Table 2 is a compilation of noble gas (He, Ne, Ar) and He and Ne isotopic compositions for cold groundwaters and produced geothermal waters in the Milford valley, Utah. Table 3 provides values for calculated advective and diffusive fluxes of helium. Table 4 provides values of calculated subsurface stored heat between the Opal Mound fault and the Utah FORGE site, which are related to volumes of recently solidified magmatic heat sources.

15 GEOTHERMAL ENERGY↗

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Future Wind Energy Resources and Cost Uncertainties Across the United States

This dataset contains results estimating projections of change of annual capacity factors and levelized cost of energy for several turbine technologies in the 2024 Annual Technology Baseline (ATB). Projections of change are based on downscaled earth system model (ESM) data from Sup3rCC. There has been evidence of reductions in average wind speeds over land in North America since the 1980s, and several models project that average wind speeds will continue to decrease. Concurrently, the cost of wind energy systems in the United States has been decreasing since around 2010, a trend also projected to continue. There is considerable uncertainty in these future projections, with quantitative estimates of future wind resource and system costs varying widely. To study this, we run land-based wind energy models with a range of possible future system costs, turbine designs, and meteorological inputs from multiple downscaled earth system models over the contiguous United States to estimate critical system performance metrics such as annual energy production (AEP) and levelized cost of energy. Where multiple earth system models agree, changes in mean AEP from the time period 2000-2019 to 2040-2059 can be as high as +10% in South Texas or as low as -20% in Iowa. Several additional states in the Midwest that currently have considerable wind generation capacity show the possibility of substantial decreases in AEP by mid-century. Larger turbines and moderate reductions in system costs can offset even the largest projected decreases in wind resource, but much uncertainty remains in the extent to which wind resources will actually change into the future and to what extent wind energy systems can drive down future costs. An analysis of variance shows, in several states in the Midwest, the uncertainty in future wind resource can be almost as important for future changes in the cost of wind energy as the uncertainty in future system costs.

17 WIND ENERGY↗

Data and scripts associated with “Non-random processes impacting organic matter chemistry are maximized in mid-order streams”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Non-random processes impacting organic matter chemistry are maximized in mid-order streams” submitted to Limnology and Oceanography (L&O) by Danczak et al. (in review). This package contains data and scripts used to investigate dissolved organic matter (DOM) molecular chemistry and diversification processes across 47 surface-water sampling sites in the Yakima River Basin, Washington, USA, during an August 2021 sampling campaign. The package contains analyses of ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS), geochemical measurements, geospatial attributes, molecular diversity, and meta-metabolome ecological null models needed to reproduce the main manuscript results. The underlying field data were pulled from exising data packages at https://doi.org/10.15485/1892052 (Fulton et al., 2022) and https://doi.org/10.15485/1898914 (Grieger et al., 2022). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. We thank the following organizations for providing access to field locations for sample collection: the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, the Confederated Tribes and Bands of the Yakama Nation, and the Cowiche Canyon Conservatory. Research was conducted under Washington State Parks and Recreation Commission Scientific Research Permit #210901. We are grateful to the Yakama Nation Tribal Council and Yakama Nation Fisheries for their collaboration in facilitating sample collection and ensuring data usage aligns with their values and worldview. This data package contains an R-Markdown file for analyses and five folders: (1) Data, (2) Geospatial Data, (3) Supplemental_Files, (5) Figures_pdf, (4) and src. The Data folder contains tabular inputs and derived files used in the manuscript analysis. The Geospatial Data folder contains climate and water-balance, hydrologic, land-cover, population/regional water-use, stream, topographic, and stream-order attribute CSV files. The src folder contains scripts used to process data, run analyses, and generate figures. The Figures_pdf folder contains manuscript figure outputs. The Supplemental_Files folder contains supplemental analysis products. All files are .csv, .pdf, .html, .png, .R, .Rmd, .svg, or .tre. This data package is associated with the rcfsa-RC2-SPS_Null_Modeling repository found at https://github.com/river-corridors-sfa/rcfsa-RC2-SPS_Null_Modeling.

54 ENVIRONMENTAL SCIENCES↗

Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNN

We present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier.

97 MATHEMATICS AND COMPUTING↗

Survey-wide asteroid discovery with a high-performance computing enabled non-linear digital tracking framework

Modern astronomical surveys detect asteroids by linking together their appearances across multiple images taken over time. This approach faces limitations in detecting faint asteroids and handling the computational complexity of trajectory linking. Here, we present a novel method that adapts “digital tracking” – traditionally used for short-term linear asteroid motion across images – to work with large-scale synoptic surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time (Rubin/LSST). Our approach combines hundreds of sparse observations of individual asteroids across their non-linear orbital paths to enhance detection sensitivity by several magnitudes. To address the computational challenges of processing massive data sets and dense orbital phase spaces, we developed a specialized high-performance computing architecture. We demonstrate the effectiveness of our method through experiments that take advantage of the extensive computational resources at Lawrence Livermore National Laboratory. This work enables the detection of significantly fainter asteroids in existing and future survey data, potentially increasing the observable asteroid population by orders of magnitude across different orbital families, from near-Earth objects (NEOs) to Kuiper belt objects (KBOs).

Asteroid discovery↗

Deep Learning Advances Arctic River Water Temperature Predictions

The accelerated warming in the Arctic poses serious risks to freshwater ecosystems by altering streamflow and river thermal regimes. However, limited research on Arctic River water temperatures exists due to data scarcity and the absence of robust methodologies, which often focus on large, major river basins. To address this, we leveraged the newly released, extensive AKTEMP data set and advanced machine learning techniques to develop a Long Short-Term Memory (LSTM) model. By incorporating ERA5-Land reanalysis data and integrating physical understanding into data-driven processes, our model advanced river water temperature predictions in ungauged, snow- and permafrost-affected basins in Alaska. Our model outperformed existing approaches in high-latitude regions, achieving a median Nash-Sutcliffe Efficiency of 0.95 and root mean squared error of 1.0°C. The LSTM model learned air temperature, soil temperature, solar radiation, and thermal radiation—factors associated with energy balance—were the most important drivers of river temperature dynamics. Soil moisture and snow water equivalent were highlighted as critical factors representing key processes such as thawing, melting, and groundwater contributions. Glaciers and permafrost were also identified as important covariates, particularly in seasonal river water temperature predictions. Our LSTM model successfully captured the complex relationships between hydrometeorological factors and river water temperatures across varying timescales and hydrological conditions. This scalable and transferable approach can be potentially applied across the Arctic, offering valuable insights for future conservation and management efforts.

54 ENVIRONMENTAL SCIENCES↗

Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures

The advent of single-particle cryo-electron microscopy (cryo-EM) has brought forth a new era of structural biology, enabling the routine determination of large biological molecules and their complexes at atomic resolution. The high-resolution structures of biological macromolecules and their complexes significantly expedite biomedical research and drug discovery. However, automatically and accurately building atomic models from high-resolution cryo-EM density maps is still time-consuming and challenging when template-based models are unavailable. Artificial intelligence (AI) methods such as deep learning trained on limited amount of labeled cryo-EM density maps generate inaccurate atomic models. To address this issue, we created a dataset called Cryo2StructData consisting of 7,600 preprocessed cryo-EM density maps whose voxels are labelled according to their corresponding known atomic structures for training and testing AI methods to build atomic models from cryo-EM density maps. Cryo2StructData is larger than existing, publicly available datasets for training AI methods to build atomic protein structures from cryo-EM density maps. We trained and tested deep learning models on Cryo2StructData to validate its quality showing that it is ready for being used to train and test AI methods for building atomic models.

59 BASIC BIOLOGICAL SCIENCES↗