Search NASASearch

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

125 records · Page 7

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING

Herbaceous Feedstock 2022 State of Technology Report

The U.S. Department of Energy promotes production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the state of technology (SOT). As part of its involvement in this mission, Idaho National Laboratory completes an annual SOT report for nth-plant and 1st-plant herbaceous biomass feedstock logistics. The purpose of the SOT is to provide the status of feedstock supply system technology development for herbaceous biomass to biofuels relative to technical targets and cost goals from specific design cases, based on data and experimental results. Although conventional feedstock supply systems form the backbone of the emerging biofuels industry, they have limitations that restrict widespread implementation on a national scale. To meet the demands of the future industry, the feedstock supply system must shift from the conventional system to what has been termed “advanced” supply systems. In advanced designs, a distributed network of aggregation and processing centers, termed “depots,” are employed near the points of biomass production (i.e., the field or forest) to reduce feedstock variability and produce feedstocks of a uniform format, moving toward biomass commoditization. The 2022 Herbaceous SOT is part of a vision of achieving an implemented advanced feedstock supply system, which produces a stable, tradable commodity at the decentralized distributed depot. It utilizes feedstock fractionation by incorporating technologies that can separate the biomass into its anatomical fractions (leaves, husks, stems and cobs) to reduce impurities and produce fractions that satisfy downstream quality considerations. By using a series of air classification steps, this strategy can reduce the extrinsic ash in corn stover and produce enriched tissue fractions that can be blended to a conversion specification or converted individually in optimized biochemical conversion campaigns. Additionally, a majority of the leaves (which do not meet the quality specification) are separated out early and can be supplied to alternate markets. The 2022 Herbaceous SOT incorporates an advanced biomass fractionation and processing system to produce pellets enriched tissues from three-pass corn stover. The resulting enriched pellets are delivered to the biorefinery individually where they can be blended to a specification or converted in campaigns where the conditions are optimized for each tissue. Unused fractions can be sent to a a midstream market or to a different conversion process that is better suited to their properties to offset the cost of the delivered feedstock. The main benefits from the proposed system can be summarized as: (1) $6.86/dry ton (2016$) lower cost for the air classification due to elimination of the requirement to discard the high ash lights fraction; (2) $1.56/dry ton lower delivered cost by selling the unsuitable leaf fraction into the feed market as a midstream co-product (assuming a selling price that is 11% higher than their cost of production); (3) 0.98% increase in carbohydrate content (from 60.16% to 61.14%); and (4) 0.97% decrease in ash content (from 6.00% to 5.03%) compared to the 2021 Herbaceous SOT. Overall, the 2022 nth-plant Herbaceous SOT predicts a modeled delivered feedstock cost of $78.64/dry ton (2016$) if it is assumed that the enriched leaf fraction is sold at its production cost; this is a slight increase of $0.43/dry ton increase from the 2021 Herbaceous SOT nth-Supply case cost. The increased cost derived from a $0.38/dry ton increase in transportation and handling cost to procure more biomass (to replace the enriched leaf fraction that was not delivered to the biorefinery. The total preprocessing cost was $0.27/dry ton higher than the 2021 result because of updates to energy consumption, purchasing price and dry matter loss data for the rotary shear ($3.00/dry ton increase) and the pelleting mill ($4.52/dry ton increase). The data utilized were generated in pilot-scale tests in the Biomass Feedstock National User Facility (BFNUF) at INL and at Forest Concepts, including tests for rotary shear and pelleting of the air classified fractions. A greenhouse gas emissions analysis was performed by Argonne National Laboratory using the most up to date version of the Greenhouse Gases, Regulated Emissions, and Energy use in Transportation model (GREET®). The analysis showed an increase of 17.34 kg CO2e/dry ton from the 2021 SOT (67.71 kg CO2e/ton in the 2021 Herbaceous SOT to 85.05 kg CO2e/ton in the 2022 Herbaceous SOT). The net increase is primarily attributed to increased energy consumption in pelleting mill.

09 BIOMASS FUELS

Investigating the Effect of Water on the Mechanical Properties of Cellulose from Multiscale Molecular Dynamics Simulations

Classical molecular dynamics (MD) simulations provide insight into the structure and physicochemical properties of materials with atomic resolution. However, the length and time scales accessible to atomistic MD are orders of magnitude smaller than many relevant processes such as the response of a bulk material to experimentally accessible strain rates, which presents challenges when comparing models to experimental measurements. Bottom-up coarse-graining provides a means for systematically mapping atomistic information to lower resolution models to increase the length and time scales achievable by simulation. Cellulose is an abundant carbohydrate biopolymer with applications to many fields of research, such as materials science and renewable energy, due to its desirable mechanical properties and viability for conversion into biofuel. The effect of moisture content on the Young's modulus of cellulose is of special interest due to its native environment often being in the hydrated secondary plant cell wall and the grinding energy requirements for biomass feedstock preprocessing. The current work investigates the effects of water solvent on the Young's modulus of cellulose calculated from coarse-grained MD mechanical stress simulations. The coarse-grained model was parametrized from atomistic MD calculations of cellulose-cellulose potentials of mean force using umbrella sampling techniques under vacuum and solvated conditions. The Young's moduli of the coarse-grained cellulose assemblies parametrized from cellulose in vacuum or solvated in water were computed via mechanical stress simulations to highlight the importance of capturing solvent interactions for modeling the mechanical behavior of cellulose.

BASIC BIOLOGICAL SCIENCES,RADIATION PROTECTION AND

PhotonIDs: ML-Powered Photon Identification System for Dark Count Elimination

Reliable single photon detection is the foundation for practical quantum communication and networking. However, today's superconducting nanowire single photon detector(SNSPD) inherently fails to distinguish between genuine photon events and dark counts, leading to degraded fidelity in long-distance quantum communication. In this work, we introduce PhotonIDs, a machine learning-powered photon identification system that is the first end-to-end solution for real-time discrimination between photons and dark count based on full SNSPD readout signal waveform analysis. PhotonIDs ~demonstrates: 1) an FPGA-based high-speed data acquisition platform that selectively captures the full waveform of signal only while filtering out the background data in real time; 2) an efficient signal preprocessing pipeline, and a novel pseudo-position metric that is derived from the physical temporal-spatial features of each detected event; 3) a hybrid machine learning model with near 98% accuracy achieved on photon/dark count classification. Additionally, proposed PhotonIDs ~ is evaluated on the dark count elimination performance with two real-world case studies: (1) 20 km quantum link, and (2) Erbium ion-based photon emission system. Our result demonstrates that PhotonIDs ~could improve more than 31.2 times of signal-noise-ratio~(SNR) on dark count elimination. PhotonIDs ~ marks a step forward in noise-resilient quantum communication infrastructure.

Linne, Karl C. [Chicago U.] (ORCID:000900091870358

An Evaluation of Representation Learning Methods in Particle Physics Foundation Models

We present a systematic evaluation of representation learning objectives for particle physics within a unified framework. Our study employs a shared transformer-based particle-cloud encoder with standardized preprocessing, matched sampling, and a consistent evaluation protocol on a jet classification dataset. We compare contrastive (supervised and self-supervised), masked particle modeling, and generative reconstruction objectives under a common training regimen. In addition, we introduce targeted supervised architectural modifications that achieve state-of-the-art performance on benchmark evaluations. This controlled comparison isolates the contributions of the learning objective, highlights their respective strengths and limitations, and provides reproducible baselines. We position this work as a reference point for the future development of foundation models in particle physics, enabling more transparent and robust progress across the community.

Chen, Michael [Caltech]

LUNA: LUT-Based Neural Architecture for Fast and Low-Cost Qubit Readout

Qubit readout is a critical operation in quantum computing systems, which maps the analog response of qubits into discrete classical states. Deep neural networks (DNNs) have recently emerged as a promising solution to improve readout accuracy . Prior hardware implementations of DNN-based readout are resource-intensive and suffer from high inference latency, limiting their practical use in low-latency decoding and quantum error correction (QEC) loops. This paper proposes LUNA, a fast and efficient superconducting qubit readout accelerator that combines low-cost integrator-based preprocessing with Look-Up Table (LUT) based neural networks for classification. The architecture uses simple integrators for dimensionality reduction with minimal hardware overhead, and employs LogicNets (DNNs synthesized into LUT logic) to drastically reduce resource usage while enabling ultra-low-latency inference. We integrate this with a differential evolution based exploration and optimization framework to identify high-quality design points. Our results show up to a 10.95x reduction in area and 30% lower latency with little to no loss in fidelity compared to the state-of-the-art. LUNA enables scalable, low-footprint, and high-speed qubit readout, supporting the development of larger and more reliable quantum computing systems.

Farooq, M. A. [Arizona State U., Tempe]

A Knowledge Graph Approach to Analyze Systems and Assets Health

Nuclear power plants collect large amounts of equipment reliability data elements that contain information on the statuses of component, assets, and systems. All these data elements precisely record asset and system performance and health throughout the lifecycle of those assets and systems. However, several challenges have proved to be roadblocks to this process. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers or databases), others are conceptual in nature (i.e., data elements come in different formats, numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). This paper directly focuses on the integration of numeric and textual data elements in order to assist plant system engineers in analyzing equipment reliability data. This task begins with preprocessing the data by extracting knowledge from textual data via natural language processing methods and quantifying system, asset, and component health based on numeric data. We then employed model-based system engineering (MBSE) models of systems and assets to identify their architecture and functional (i.e., cause and effect) relations. Data elements were then associated with a single MBSE graph element, based on their nature. This bonding of MBSE models and data elements constitutes a first-of-its-kind knowledge graph of a nuclear power plants system, with data elements being organized in a structured manner that enables system engineers to identify cause-effect trends in data elements and carry out appropriate actions in response.

97 - MATHEMATICS AND COMPUTING

Air Classification of Forestry Residues for Fast Pyrolysis

Understanding critical biomass attributes through efficient fractionation is crucial for advancing sustainable pyrolysis for renewable energy and chemical production. This study investigates the intricate relationship between biomass preprocessing and pyrolysis product yields, employing the air classification technique for the treatment of loblolly pine residues with varying moisture content. A comprehensive exploration of the physicochemical properties of air-classified loblolly pine informs a sophisticated pyrolysis simulation model. Given the complex and multifaceted nature of biomass pyrolysis, operating across diverse temporal and spatial scales, a pyrolysis kinetics-based CFD–DEM simulation method is employed to predict product yields. Results showed that the elevated moisture content amplifies particle adhesiveness, necessitating augmented air velocities for effective separation, thereby influencing the efficiency of the separation process. While carbon and hydrogen contents exhibit relative stability across diverse moisture contents and blower frequencies, the oxygen content undergoes noticeable changes. For example, the oxygen contents were measured as 29.2 and 38.6 wt% in the light fraction of 30% moisture content sample at blower frequencies of 10 and 20 Hz, respectively. An intriguing finding emerges from pyrolysis simulation, indicating that a lower blower frequency in air classification moderately enhances bio-oil yield and significantly improves its quality, particularly in terms of water content. For instance, the water content in the bio-oil was about 1.5% and 10% in the heavy and light fractions, respectively from 10% moisture sample under 15 Hz blower frequency.

09 - BIOMASS FUELS

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING

Demystifying Piecewise and Localized Scatter Correction Methods

Multiplicative scatter is a common source of noise in near-infrared spectroscopy and other related instrumental techniques. A wide variety of methods are commonly used for the correction of multiplicative scatter. However, the majority of such methods assume that the parameters that describe the scatter are constant throughout the measured spectrum, which is often not the case. This work investigates a family of methods that perform scatter correction using local regions of neighboring wavelengths in order to better account for wavelength-dependent scattering. The methods in question are piecewise standard normal variate, localized standard normal variate, piecewise multiplicative scatter correction, and localized multiplicative scatter correction. This work describes the theoretical and algorithmic foundations of the family of local region-based scatter correction methods and compares their application and optimization at a qualitative and quantitative level using several datasets.

NIR spectroscopy

Enhancing the flowability of woody biomass slurries in wet biorefineries

Feeding wet lignocellulosic biomass (e.g., softwood and hardwood) slurries into high-pressure, high-temperature reactors at an industrially relevant scale presents significant challenges, such as process equipment plugging. Here, in this study, we investigate the possibility of improving the flowability of biomass slurries in an industrial scale wet biorefinery by tuning the physical and chemical characteristics of the biomass particles. To understand the effects of the chemical characteristics of woody biomass on flowability, cellulose pulp particles are produced from pulp sheets via knife milling, pelletizing, and crumbling. These cellulose pulp particles are processed for acid hydrolysis dehydration (AHDH) at one –tonne-per day pilot scale. The levulinic acid yield (a product of cellulose AHDH) and the flowability of the biomass particles are compared using market pulp, 1 mm crumbled pine softwood, a blend of 2 mm hammer-milled pine softwood with 10 wt.% bark, and 2 mm hammer-milled hardwood material with 10 wt.% bark (to represent forest residues). The results indicate that the presence of hemicellulose and lignin influences the flowability of lignocellulosic feedstocks. Crumbled wood particles show poor flowability at the pilot scale, while the presence of fine materials (less than 0.7 mm) and bark improves the flowability of biomass slurries without affecting the organic acid yields (based on C6 carbohydrate content).

09 - BIOMASS FUELS

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Overview of RFID Applications Utilizing Neural Networks

As Radio Frequency Identification (RFID) methods continue to evolve to higher levels of complexity, one form of machine learning is making its appearance. The use of Neural Networks (NN) in the RFID field is steadily increasing, and in the fields of localization and activity recognition, promising results are being shown from a variety of research. RFID applications fall primarily under two types of problems including regression and classification. We analyze RIFD localization techniques which fall under regression, and activity recognition which falls under classification. Many works don’t classify themselves as activity recognition methods, but because they fall under the classification category, we still consider them as activity recognition techniques. This research overviews the Neural Network models in the localization field based on whether they can perform independently of the environment in which they were tested. For activity recognition and accessory fields, the major methods involve tag-based and tag-free approaches. In conclusion, after the models are surveyed, a comparison study is given to examine what may be the cause for increased accuracy between different Neural Network models.

42 ENGINEERING

Interparticle Characterization of Mechanical Biomass Particle-Particle and Particle-Wall Interactions

The biomass materials industry faces significant challenges in managing material variability and its impact on storage and handling systems. Physical properties such as moisture content, particle size, and density fluctuate considerably, leading to operational issues like bridging and ratholing that disrupt material flow. These variations create a complex cascade effect throughout the process chain, affecting transportation, storage, and conversion processes. The economic consequences of this variability manifest in increased operational costs, maintenance requirements, and system downtime. Environmental factors further complicate the situation, as weather conditions and seasonal availability influence material properties and system performance. Engineers employ specialized equipment design, material characterization protocols, and pre-processing steps like size reduction and homogenization to address these challenges. A critical knowledge gap exists between continuous-level constitutive models and particle-scale behavior. This project developed a novel device to quantify interparticle mechanics between biomass particles, measuring friction and adhesion forces between particles and wall materials. The research focused on corn stover and southern pine forest residue, creating a comprehensive database of particle interactions. This breakthrough enables direct application in particle-based computational modeling, advancing the field's understanding of biomass handling characteristics and supporting the development of more reliable and efficient storage and handling systems. The project's outcomes contribute significantly to understanding biomass's mechanical and flow characteristics, particularly how variability at the particle level affects larger-scale handling operations. This knowledge is crucial for engineering feedstock supply systems that consistently meet quality and cost specifications for various conversion processes. The innovative experimental setup developed through this research represents a significant advancement in biomass characterization methodology. Providing precise measurements of particle-level interactions establishes a foundation for more accurate predictive modeling of bulk material behavior. This enhanced understanding of fundamental particle mechanics enables engineers to anticipate better and address handling challenges before they manifest in full-scale operations. This research opens new avenues for optimizing biomass handling systems through data-driven design approaches. The comprehensive database of particle interactions serves as a valuable resource for future research and development efforts, potentially leading to more efficient and cost-effective biomass processing solutions. This advancement in particle-level mechanics could revolutionize how biomass handling systems are designed and operated, contributing to more sustainable and reliable renewable energy production.

09 BIOMASS FUELS

Processing Municipal Solid Waste for Conversion to Jet Fuel (Final Technical Report)

This final report describes the successful demonstration of AI‑assisted waste sorting and a novel pressurized solids feed system to enable conversion of non‑recyclable municipal solid waste to jet fuel under DOE Award DE‑EE0009265. Results include improved feedstock purity and variability reduction, lab‑scale testing, and technoeconomic and life‑cycle assessments.

09 BIOMASS FUELS