Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Future Wind Energy Resources and Cost Uncertainties Across the United States

This dataset contains results estimating projections of change of annual capacity factors and levelized cost of energy for several turbine technologies in the 2024 Annual Technology Baseline (ATB). Projections of change are based on downscaled earth system model (ESM) data from Sup3rCC. There has been evidence of reductions in average wind speeds over land in North America since the 1980s, and several models project that average wind speeds will continue to decrease. Concurrently, the cost of wind energy systems in the United States has been decreasing since around 2010, a trend also projected to continue. There is considerable uncertainty in these future projections, with quantitative estimates of future wind resource and system costs varying widely. To study this, we run land-based wind energy models with a range of possible future system costs, turbine designs, and meteorological inputs from multiple downscaled earth system models over the contiguous United States to estimate critical system performance metrics such as annual energy production (AEP) and levelized cost of energy. Where multiple earth system models agree, changes in mean AEP from the time period 2000-2019 to 2040-2059 can be as high as +10% in South Texas or as low as -20% in Iowa. Several additional states in the Midwest that currently have considerable wind generation capacity show the possibility of substantial decreases in AEP by mid-century. Larger turbines and moderate reductions in system costs can offset even the largest projected decreases in wind resource, but much uncertainty remains in the extent to which wind resources will actually change into the future and to what extent wind energy systems can drive down future costs. An analysis of variance shows, in several states in the Midwest, the uncertainty in future wind resource can be almost as important for future changes in the cost of wind energy as the uncertainty in future system costs.

17 WIND ENERGY↗

ALICE luminosity determination for Pb–Pb collisions at $\sqrt{s_{NN}}$ = 5.02 TeV

Luminosity determination within the ALICE experiment is based on the measurement, in van der Meer scans, of the cross sections for visible processes involving one or more detectors (visible cross sections). In 2015 and 2018, the Large Hadron Collider provided Pb–Pb collisions at a centre-of-mass energy per nucleon pair of $\sqrt{s_{NN}}$ = 5.02 TeV. Two visible cross sections, associated with particle detection in the Zero Degree Calorimeter (ZDC) and in the V0 detector, were measured in a van der Meer scan. This article describes the experimental set-up and the analysis procedure, and presents the measurement results. The analysis involves a comprehensive study of beam-related effects and an improved fitting procedure, compared to previous ALICE studies, for the extraction of the visible cross section. The resulting uncertainty of both the ZDC-based and the V0-based luminosity measurement for the full sample is 2.5%. The inelastic cross section for hadronic interactions in Pb–Pb collisions at $\sqrt{s_{NN}}$ = 5.02 TeV, obtained by efficiency correction of the V0-based visible cross section, was measured to be 7.67 ± 0.25 b, in agreement with predictions using the Glauber model.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data and scripts associated with “Non-random processes impacting organic matter chemistry are maximized in mid-order streams”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Non-random processes impacting organic matter chemistry are maximized in mid-order streams” submitted to Limnology and Oceanography (L&O) by Danczak et al. (in review). This package contains data and scripts used to investigate dissolved organic matter (DOM) molecular chemistry and diversification processes across 47 surface-water sampling sites in the Yakima River Basin, Washington, USA, during an August 2021 sampling campaign. The package contains analyses of ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS), geochemical measurements, geospatial attributes, molecular diversity, and meta-metabolome ecological null models needed to reproduce the main manuscript results. The underlying field data were pulled from exising data packages at https://doi.org/10.15485/1892052 (Fulton et al., 2022) and https://doi.org/10.15485/1898914 (Grieger et al., 2022). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. We thank the following organizations for providing access to field locations for sample collection: the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, the Confederated Tribes and Bands of the Yakama Nation, and the Cowiche Canyon Conservatory. Research was conducted under Washington State Parks and Recreation Commission Scientific Research Permit #210901. We are grateful to the Yakama Nation Tribal Council and Yakama Nation Fisheries for their collaboration in facilitating sample collection and ensuring data usage aligns with their values and worldview. This data package contains an R-Markdown file for analyses and five folders: (1) Data, (2) Geospatial Data, (3) Supplemental_Files, (5) Figures_pdf, (4) and src. The Data folder contains tabular inputs and derived files used in the manuscript analysis. The Geospatial Data folder contains climate and water-balance, hydrologic, land-cover, population/regional water-use, stream, topographic, and stream-order attribute CSV files. The src folder contains scripts used to process data, run analyses, and generate figures. The Figures_pdf folder contains manuscript figure outputs. The Supplemental_Files folder contains supplemental analysis products. All files are .csv, .pdf, .html, .png, .R, .Rmd, .svg, or .tre. This data package is associated with the rcfsa-RC2-SPS_Null_Modeling repository found at https://github.com/river-corridors-sfa/rcfsa-RC2-SPS_Null_Modeling.

54 ENVIRONMENTAL SCIENCES↗

Computational tools and algorithms for ion mobility spectrometry-mass spectrometry

Ion mobility spectrometry-mass spectrometry (IMS-MS or IM-MS) is a powerful analytical technique that combines the gas-phase separation capabilities of IM with the identification and quantification capabilities of MS. IM-MS can differentiate molecules with indistinguishable masses but different structures (e.g., isomers, isobars, molecular classes, and contaminant ions). The importance of this analytical technique is reflected by a staged increase in the number of applications for molecular characterization across a variety of fields, from different MS-based omics (proteomics, metabolomics, lipidomics, etc.) to the structural characterization of glycans, organic matter, proteins, and macromolecular complexes. With the increasing application of IM-MS there is a pressing need for effective and accessible computational tools. This article presents an overview of the most recent free and open-source software tools specifically tailored for the analysis and interpretation of data derived from IM-MS instrumentation. This review enumerates these tools and outlines their main algorithmic approaches, while highlighting representative applications across different fields. Finally, a discussion of current limitations and expectable improvements is presented.

59 BASIC BIOLOGICAL SCIENCES↗

Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNN

We present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier.

97 MATHEMATICS AND COMPUTING↗

Survey-wide asteroid discovery with a high-performance computing enabled non-linear digital tracking framework

Modern astronomical surveys detect asteroids by linking together their appearances across multiple images taken over time. This approach faces limitations in detecting faint asteroids and handling the computational complexity of trajectory linking. Here, we present a novel method that adapts “digital tracking” – traditionally used for short-term linear asteroid motion across images – to work with large-scale synoptic surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time (Rubin/LSST). Our approach combines hundreds of sparse observations of individual asteroids across their non-linear orbital paths to enhance detection sensitivity by several magnitudes. To address the computational challenges of processing massive data sets and dense orbital phase spaces, we developed a specialized high-performance computing architecture. We demonstrate the effectiveness of our method through experiments that take advantage of the extensive computational resources at Lawrence Livermore National Laboratory. This work enables the detection of significantly fainter asteroids in existing and future survey data, potentially increasing the observable asteroid population by orders of magnitude across different orbital families, from near-Earth objects (NEOs) to Kuiper belt objects (KBOs).

Asteroid discovery↗

Deep Learning Advances Arctic River Water Temperature Predictions

The accelerated warming in the Arctic poses serious risks to freshwater ecosystems by altering streamflow and river thermal regimes. However, limited research on Arctic River water temperatures exists due to data scarcity and the absence of robust methodologies, which often focus on large, major river basins. To address this, we leveraged the newly released, extensive AKTEMP data set and advanced machine learning techniques to develop a Long Short-Term Memory (LSTM) model. By incorporating ERA5-Land reanalysis data and integrating physical understanding into data-driven processes, our model advanced river water temperature predictions in ungauged, snow- and permafrost-affected basins in Alaska. Our model outperformed existing approaches in high-latitude regions, achieving a median Nash-Sutcliffe Efficiency of 0.95 and root mean squared error of 1.0°C. The LSTM model learned air temperature, soil temperature, solar radiation, and thermal radiation—factors associated with energy balance—were the most important drivers of river temperature dynamics. Soil moisture and snow water equivalent were highlighted as critical factors representing key processes such as thawing, melting, and groundwater contributions. Glaciers and permafrost were also identified as important covariates, particularly in seasonal river water temperature predictions. Our LSTM model successfully captured the complex relationships between hydrometeorological factors and river water temperatures across varying timescales and hydrological conditions. This scalable and transferable approach can be potentially applied across the Arctic, offering valuable insights for future conservation and management efforts.

54 ENVIRONMENTAL SCIENCES↗

Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures

The advent of single-particle cryo-electron microscopy (cryo-EM) has brought forth a new era of structural biology, enabling the routine determination of large biological molecules and their complexes at atomic resolution. The high-resolution structures of biological macromolecules and their complexes significantly expedite biomedical research and drug discovery. However, automatically and accurately building atomic models from high-resolution cryo-EM density maps is still time-consuming and challenging when template-based models are unavailable. Artificial intelligence (AI) methods such as deep learning trained on limited amount of labeled cryo-EM density maps generate inaccurate atomic models. To address this issue, we created a dataset called Cryo2StructData consisting of 7,600 preprocessed cryo-EM density maps whose voxels are labelled according to their corresponding known atomic structures for training and testing AI methods to build atomic models from cryo-EM density maps. Cryo2StructData is larger than existing, publicly available datasets for training AI methods to build atomic protein structures from cryo-EM density maps. We trained and tested deep learning models on Cryo2StructData to validate its quality showing that it is ready for being used to train and test AI methods for building atomic models.

59 BASIC BIOLOGICAL SCIENCES↗

Coassembly and binning of a twenty-year metagenomic time-series from Lake Mendota

Abstract The North Temperate Lakes Long-Term Ecological Research (NTL-LTER) program has been extensively used to improve understanding of how aquatic ecosystems respond to environmental stressors, climate fluctuations, and human activities. Here, we report on the metagenomes of samples collected between 2000 and 2019 from Lake Mendota, a freshwater eutrophic lake within the NTL-LTER site. We utilized the distributed metagenome assembler MetaHipMer to coassemble over 10 terabases (Tbp) of data from 471 individual Illumina-sequenced metagenomes. A total of 95,523,664 contigs were assembled and binned to generate 1,894 non-redundant metagenome-assembled genomes (MAGs) with ≥50% completeness and ≤10% contamination. Phylogenomic analysis revealed that the MAGs were nearly exclusively bacterial, dominated by Pseudomonadota (Proteobacteria, N = 623) and Bacteroidota (N = 321). Nine eukaryotic MAGs were identified by eukCC with six assigned to the phylum Chlorophyta. Additionally, 6,350 high-quality viral sequences were identified by geNomad with the majority classified in the phylum Uroviricota. This expansive coassembled metagenomic dataset provides an unprecedented foundation to advance understanding of microbial communities in freshwater ecosystems and explore temporal ecosystem dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-Attribute Subset Selection enables prediction of representative phenotypes across microbial populations

The interpretation of complex biological datasets requires the identification of representative variables that describe the data without critical information loss. This is particularly important in the analysis of large phenotypic datasets (phenomics). Here we introduce Multi-Attribute Subset Selection (MASS), an algorithm which separates a matrix of phenotypes (e.g., yield across microbial species and environmental conditions) into predictor and response sets of conditions. Using mixed integer linear programming, MASS expresses the response conditions as a linear combination of the predictor conditions, while simultaneously searching for the optimally descriptive set of predictors. We apply the algorithm to three microbial datasets and identify environmental conditions that predict phenotypes under other conditions, providing biologically interpretable axes for strain discrimination. MASS could be used to reduce the number of experiments needed to identify species or to map their metabolic capabilities. The generality of the algorithm allows addressing subset selection problems in areas beyond biology.

59 BASIC BIOLOGICAL SCIENCES↗

Thoroughly testing and integrating hundreds of Pull Requests per month: ROOT’s new Cost-efficient and Feature Rich GitHub-based CI

ROOT is an open source framework, freely available on GitHub, at the heart of data acquisition, processing and analysis of HE(N)P experiments, and beyond. It is developed collaboratively: contributions are not authored only by ROOT team members, but also by the user community at large: developers and scientists from universities, labs as well as the private sector. More than 1500 GitHub Pull Requests are merged on average per year. It is in this context that code integration acquires a primary role. The review of code contributions isn’t enough: not only they need to be thoroughly reviewed, they also need to be thoroughly tested through a powerful CI infrastructure on several different platforms to comply with the high code quality standards of the project. Since the end of 2023, ROOT moved its continuous integration system from Jenkins to GitHub Actions. In this contribution, we characterise the transition to the GitHub CI, focussing on our strategy, its implementation and the lessons learned, as well as the advantages the new system offers with respect to the previous one. Particular emphasis will be given to the evaluation of the cost-benefit ratio for Jenkins and GitHub Actions for the ROOT project. We also describe how we manage to run in less than one hour thousands of unit, integration, functional and end-to-end tests on different flavours of Windows, four versions of macOS, as well as about ten of the most used Linux distributions, taking advantage of the CERN computing infrastructure.

Piparo, Danilo [CERN]↗

cclib 2.0: An updated architecture for interoperable computational chemistry

Interoperability in computational chemistry is elusive, impeded by the independent development of software packages and idiosyncratic nature of their output files. The cclib library was introduced in 2006 as an attempt to improve this situation by providing a consistent interface to the results of various quantum chemistry programs. The shared API across programs enabled by cclib has allowed users to focus on results as opposed to output and to combine data from multiple programs or develop generic downstream tools. Initial development, however, did not anticipate the rapid progress of computational capabilities, novel methods, and new programs; nor did it foresee the growing need for customizability. Here, we recount this history and present cclib 2, focused on extensibility and modularity. We also introduce recent design pivots—the formalization of cclib’s intermediate data representation as a tree-based structure, a new combinator-based parser organization, and parsed chemical properties as extensible objects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Novel data interpretation method for DIII-D divertor retarding field energy analyzer with 3-D particle-in-cell simulations

A novel data interpretation process that utilizes comprehensive particle-in-cell (PIC) simulations is developed for the new retarding field energy analyzer (RFEA) currently being constructed at DIII-D for the lower divertor using the Divertor Material Evaluation System. Furthermore, this probe is expected to survive a heat load of up to 100 MW/m 2 for up to 5 s and reliably measure the main ion temperature (T i ) on the divertor target ranging from 10 to 200 eV. These extreme conditions posed significant engineering limitations on the probe geometry, thus extensive validation work has been performed. The conventional fitting method for the RFEA I–V characteristics is based on a simplified 1-D model without considering the ion space charge inside the probe cavity and may not be sufficient for probes designed for the DIII-D divertor environment. In this article, a more realistic description of the particle propagation process within the RFEA cavity is achieved by including both 3-D geometric effects and ion space charge in the PIC simulations, and the capability to reconstruct the ion energy distribution functions is demonstrated with reasonable consistency.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Phase-based velocity extraction method for photonic Doppler velocimetry with potential higher time resolution

We present an extension of the [Takeda et al., J. Opt. Soc. Am. 72, 156 (1982)] phase extraction method to heterodyne photonic Doppler velocimetry applications. The method yields results equivalent to those obtained by the short-time Fourier transform (STFT), while offering potential improvements in time resolution. Unlike STFT, which relies on window functions, such as the Hamming window, that emphasize central data points and diminish the influence of edges, the extended Takeda method utilizes all data uniformly. This uniform treatment allows for the derivation of empirical equations that directly relate velocity error to the actual time resolution rather than to the local analysis duration. The established equation provides a useful metric for both optimizing hardware configuration and guiding data analysis. Simulation and experimental results confirm that, for a given dataset, specifying a target time resolution yields consistent velocity errors for both methods. These findings underscore the Takeda method’s advantages, particularly its potential higher time resolution and reduced computational burden, making it a valuable tool for high-throughput applications such as laser dynamic compression experiments.

Computer simulation↗

An interactive machine learning platform for analyzing multi-particle coincidence data from cold target recoil ion momentum spectroscopy

We present SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training), a comprehensive software platform for analyzing tabulated high-dimensional multi-particle coincidence data from Cold Target Recoil Ion Momentum Spectroscopy (COLTRIMS) experiments. The software addresses critical challenges in modern momentum spectroscopy by integrating advanced machine learning techniques with physics-informed analysis in an interactive web-based environment. SCULPT implements uniform manifold approximation and projection for non-linear dimensionality reduction to reveal correlations in high-dimensional data. We also discuss potential extensions to deep autoencoders for feature learning and genetic programming for automated discovery of physically meaningful observables. A novel adaptive confidence scoring system provides quantitative reliability assessments by evaluating user-selected clustering quality metrics with predefined weights that reflect each metric’s robustness. The platform features configurable molecular profiles for different experimental systems, interactive visualization with selection tools, and comprehensive data filtering capabilities. Utilizing a subset of SCULPT’s capabilities, we analyze photo-double-ionization data measured using the COLTRIMS method for three-body dissociation of the D 2 O molecule, revealing distinct fragmentation channels and their correlations with physics parameters. The software’s modular architecture and web-based implementation make it accessible to the broader atomic and molecular physics community, significantly reducing the time required for complex multi-dimensional analyses. This opens the door to finding and isolating rare events exhibiting non-linear correlations on the fly during experimental measurements, which can help steer exploration and improve the efficiency of experiments.

Artificial neural networks↗