Search NASASearch

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Rapid Optimization of Total Variation with Applications in Imaging, Additive Manufacturing, and Qualification

Total Variation optimization penalizes the gradient of a control variable or state. While this work focuses on image processing in particular, it has also found applications in inverse problems and topology optimization. In image processing, the goal is to maintain faithfulness to the original image while denoising and/or deblurring. Additionally, bilevel optimization over the spatially varying regularization weights can illuminate interfaces such as damage regions and other anomalies. We will address two fundamental challenges with TV-optimization: (i) the typical slow convergence of existing TV-optimization methods, and (ii) the selection of spatially varying TV parameters to promote interface detection. Additionally, we will apply such techniques to image data collected in additive manufacturing. In said context, stochasticity in build events induces flaws in the manufactured piece, compromising the integrity of said part. There is a critical need for in-situ monitoring to spot anomalies once they form, and in this setting we apply our total variation and hyperparameter solvers. We will develop a customized algorithm based on for extreme-scale TV-optimization that achieves super-linear or quadratic-convergence, a critical property for real-time, image-by-image analysis. A worst-case outcome is a preprocessing step that enhances image quality in-situ, specifically for out-of-focus and noisy images.

36 MATERIALS SCIENCE

Training NuGraph2 for ICARUS

This presentation describes the process of training NuGraph2, a Graphical Neural Network for event reconstruction, on simulated ICARUS neutrino event data. This began with an investigation into filtering ICARUS spacepoint data. Then NuGraph2 was repeatedly trained on three event samples, which were used for finding optimized machine-learning parameters and to find and fix the causes of several crashes in NuGraph2 s preprocessing and training scripts.

43 PARTICLE ACCELERATORS

Updates to the Hanford Soil Inventory Model (SIM Version 2) for FY 2023

The purpose of this environmental calculation file (ECF) is to document the updates made to the Soil Inventory Model version 2.1 (SIM-v2.1) (ECF-HANFORD-21-0073, Updates to the Hanford Soil Inventory Model (SIM Version 2) for FY 2021) in FY 2023. These updates to SIM-v2.1 are referred to as the SIM-v2.2 version. The following updates have been made: • Include inventory estimate for a new analyte • hexavalent chromium (Cr(VI)), separate from total chromium inventory • Revise the inventory discharged to the 216-U-10 and 216-T-4 Pond systems based on partitioning of waste streams discharges over time and space among influent ditches and ponds • Enhance the preprocessing and postprocessing of the input and output files Note that the methodology for estimating Cr(VI) inventories is based on ECF-200W-23-0040, Recommendations for Updating Liquid Discharged Inventory and Transport Modeling Parameters for Cumulative Impacts Evaluation of Hexavalent Chromium in the 200 West Area while the revision of inventory discharged to the 216-U-10 and 216-T-4 Pond systems is based on ECF-HANFORD-19-0032, Distribution of Infiltration in the 216-U-10, 216-B-3 Pond, and 216-T-4 Pond Systems 1944-1997.

54 ENVIRONMENTAL SCIENCES

Day-Ahead Probabilistic Forecasting of Net-Load and Demand Response Potentials with High Penetration of Behind-the-Meter Solar-plus-Storage

The goal of this project is to develop advanced methods for day-ahead net-load forecasting, by leveraging the state-of-the-art machine learning techniques. The developed models produce both point and probabilistic forecasts for a variety of use cases, and are versatile to work with different types of data sets. The innovation lies in the novel design of the architectures, leveraging the most recent advances in machine learning that have not been explored in power systems, accompanied by techniques in the broader artificial intelligence fields such as fuzzy systems. This project has achieved the following accomplishments: (1) preprocessing of over 10 data sets covering varying geographical regions, time horizons, and system levels, which form a robust foundation for training and evaluating forecasting models across a wide range of realistic grid scenarios; (2) development of an interactive web app that enables exploratory analysis of load and generation data, and supports better understanding of data trends, anomalies, and correlations, facilitating model development and stakeholder engagement; (3) implementation of over 10 benchmark models for point and probabilistic forecasting, which include a mix of conventional machine learning methods and state-of-the-art deep learning approaches, providing a comprehensive baseline for performance comparison and validation of the proposed models; (4) development of a fuzzy system based gradient boosting model, tailored for small (less than 3 years) data sets, which achieves a mean absolute percentage error (MAPE) of 4% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (5) development of a Transformer (a state-of-the-art deep learning architecture) based neural network model, tailored for large (3 years or more) data sets, which achieves a MAPE of 2% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (6) development of a methodology for quantifying DR potential, and extensions of the previous models for multi-target forecasting of net load and DR potential, which achieve a MAPE of 10% for DR potential.

24 POWER TRANSMISSION AND DISTRIBUTION

Grain2Mesh: Mesh Generation for Grain-Scale Nonlinear Elasticity Modeling

The nonlinear hysteretic behavior of rocks under cyclic loading is a crucial area of study in geomechanics. The macroscopic response of a variety of materials has been found to be contingent upon the behavior of the micro-scale structure. This project aims to develop a functional and maintainable software package for generating a multi-phase numerical mesh and accompanying simulation files for finite element modeling used in computational mechanics solvers. Meshes generated from images often lack key preprocessing that reduces noise and prevents mesh element distortion that can increase computational cost. By incorporating user feedback throughout, grain2mesh ensures a high-fidelity mesh that can be used to model grain-scale interactions such as shearing, crack propagation, and interfacial material contrast. Scientific applications of this software include material fracturing, stress-strain analysis for natural and engineered materials, and nonlinear meso-scale analysis.

54 ENVIRONMENTAL SCIENCES

Scalable Unit Commitment with Security Constrained AC Power Flow via ADMM and Hybrid Modeling Strategies

This research introduces a more efficient way to optimize power grid operations, breaking the problem into manageable steps and using advanced mathematical techniques to speed up calculations. By incorporating smart heuristics, improved preprocessing, and contingency analysis, the approach allows operators to make better decisions faster. These innovations enhance our understanding of how to optimize energy generation, making it possible to anticipate failures before they happen, reduce system costs, and improve overall grid performance. Ultimately, this research helps bridge the gap between theoretical models and real-world applications, paving the way for a smarter, more resilient power grid. This research directly benefits the public by making electricity more affordable, reliable, and sustainable. By improving how power grids schedule and distribute electricity, the project helps energy providers reduce operational costs, which can lead to lower electricity prices for consumers. Additionally, the ability to predict and prevent power system failures enhances grid reliability, reducing the likelihood of blackouts that can disrupt homes, businesses, and critical infrastructure such as hospitals. From an environmental perspective, optimizing power generation reduces energy waste and lowers carbon emissions, contributing to cleaner air and a more sustainable energy system. Furthermore, with extreme weather events becoming more frequent, these advancements make the power grid more resilient, ensuring communities are better prepared for emergencies and natural disasters. By strengthening the nation's energy infrastructure, this research plays a crucial role in improving economic stability, public safety, and environmental sustainability.

24 POWER TRANSMISSION AND DISTRIBUTION

Zero and Finite Temperature Quantum Simulations Powered by Quantum Magic

We introduce a quantum information theory-inspired method to improve the characterization of many-body Hamiltonians on near-term quantum devices. We design a new class of similarity transformations that, when applied as a preprocessing step, can substantially simplify a Hamiltonian for subsequent analysis on quantum hardware. By design, these transformations can be identified and applied efficiently using purely classical resources. In practice, these transformations allow us to shorten requisite physical circuit-depths, overcoming constraints imposed by imperfect near-term hardware. Importantly, the quality of our transformations is t u n a b l e : we define a 'ladder' of transformations that yields increasingly simple Hamiltonians at the cost of more classical computation. Using quantum chemistry as a benchmark application, we demonstrate that our protocol leads to significant performance improvements for zero and finite temperature free energy calculations on both digital and analog quantum hardware. Specifically, our energy estimates not only outperform traditional Hartree-Fock solutions, but this performance gap also consistently widens as we tune up the quality of our transformations. In short, our quantum information-based approach opens promising new pathways to realizing useful and feasible quantum chemistry algorithms on near-term hardware.

Physics

The Double-edged Sword of Data-driven Super-Resolution: Adversarial Super-resolution Models

Data-driven super-resolution (SR) methods are often integrated into imaging pipelines as preprocessing steps to improve downstream tasks such as classification and detection. However, these SR models introduce a previously unexplored attack surface into imaging pipelines. In this paper, we present AdvSR, a framework demonstrating that adversarial behavior can be embedded directly into SR model weights during training, requiring no access to inputs at inference time. Unlike prior attacks that perturb inputs or rely on backdoor triggers, AdvSR operates entirely at the model level. By jointly optimizing for reconstruction quality and targeted adversarial outcomes, AdvSR produces models that appear benign under standard image quality metrics while inducing downstream misclassification. We evaluate AdvSR on three SR architectures (SRCNN, EDSR, SwinIR) paired with a YOLOv11 classifier and demonstrate that AdvSR models can achieve high attack success rates with minimal quality degradation. These findings highlight a new model-level threat for imaging pipelines, with implications for how practitioners source and validate models in safety-critical applications.

Sullivan, Haley [ORNL] (ORCID:0000000274069217)

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun

Applications of LIF to Document Natural Variability of Chlorophyll Content and Cu Uptake in Moss

Chlorophyll has long been used as a natural indicator of plant health and photosynthetic efficiency. Laser-induced fluorescence (LIF) is an emerging technique for understanding broad spectrum organic processes and has more recently been used to monitor chlorophyll response in plants. Previous work has focused on developing a LIF technique for imaging moss mats to identify metal contamination with the current focus shifting toward application to moss fronds and aiding sample collection for chemical analysis. Two laser systems (CoCoBi a Nd:YGa pulsed laser system and Chl-SL with two blue continuous semiconductor diodes) were used to collect images of moss fronds exposed to increasing levels of Cu (1, 10, and 100 nmol/cm 2 ) using a CMOS camera. The best methods for the preprocessing of images were conducted before the analysis of fluorescence signatures were compared to a control. The Chl-SL system performed better than the CoCoBi, with dynamic time warping (DTW) proving the most effective for image analysis. Manual thresholding to remove lower decimal code values improved the data distributions and proved whether using one or two fronds in an image was more advantageous. A higher DTW difference from the control correlated to lower chlorophyll a/b ratios and a higher metal content, indicating that LIF, with the aid of image processing, can be an effective technique for identifying Cu contamination shortly after an event.

59 BASIC BIOLOGICAL SCIENCES

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING

H I Depletion Begins Well Beyond the Virial Radius: A FAST Stacking Study of 36 Galaxy Clusters to 5 × R 200

Abstract We present a stacking study of the neutral atomic hydrogen (H i ) content in and around 36 local galaxy clusters at z < 0.07, using a combination of the FAST All Sky H i survey (FASHI) and the extensive spectroscopic catalog mainly from the Dark Energy Spectroscopic Instrument (DESI). We employ spectral stacking techniques to probe the average H i mass and HI-to-stellar mass ratio ( M HI / M * ) for member galaxies down to stellar masses of M * ∼ 10 9 M ⊙ , spanning a projected cluster-centric distance of up to 5 R 200 . Our analysis reveals a pronounced environmental effect; both M HI and M HI / M * decrease steadily toward the cluster center, dropping by ∼0.5 dex on average from the outskirts to the core. Crucially, we find that M HI / M * of galaxies remain lower than the field galaxies even at the 5 R 200 . This provides direct, statistical evidence for substantial gas stripping and preprocessing in the cluster outskirts, likely occurring in infalling groups and large-scale filaments. By further splitting the sample by g − r color, we show that the H i deficiency persists at fixed galaxy color; even the bluest cluster members exhibit ∼0.5 dex lower M HI / M * than field galaxies of similar color, reflecting environmental effects on the cold gas reservoir prior to full optical transformation. The total H i mass within clusters and their outskirts agrees broadly with predictions from cosmological simulations. Our results underscore the critical role of the extended cluster environment in quenching galaxies by depleting their cold gas reservoirs well before they enter the dense cluster core.

Cheng, Cheng [Chinese Academy of Sciences South Am

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING

Preparation and evaluation of Apollo 14 composite experiments

An account is given of the work aimed at flight experiments on Apollo 14, in relation to space manufacturing processes. Evaluation of suitable materials, definition of in-flight processing procedures, preparation of preprocessed materials and delivery, and evaluation of the space-processed samples after return from the Apollo 14 flight are presented.

Steurer, W. H.