Search NASA⌕ Search

SEARCH · Search NASA

Results for “Batch Correction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

malbacR: A Package for Standardized Implementation of Batch Correction Methods for Omics Data

Mass spectrometry is a powerful tool for identifying and analyzing small molecules, such as metabolites and lipids, in com-plex biological samples. Liquid chromatography and gas chromatography mass spectrometry studies quite commonly in-volve large numbers of samples, which can require significant time for sample preparation and analyses. To accommodate such studies, the samples are commonly split into batches. Inevitably, variations in sample handling, temperature fluctua-tion, imprecise timing, column degradation and other factors result in systematic errors or biases of the measured abundances between the batches. Numerous methods are available via R packages to assist with batch correction for small molecule om-ics data; however, since these methods were developed by different research teams, the algorithms are available in separate R packages, each with different data input and output formats. We introduce the malbacR package which consolidates eleven common batch effect correction methods for small molecule omics data into one place so users can easily implement and compare: pareto scaling, power scaling, range scaling, ComBat, EigenMS, NOMIS, RUV-random, QC-RLSC, WaveI-CA2.0, TIGER, and SERRF. The malbacR package standardizes data input and output formats across these batch correction methods. The package works in conjunction with the pmartR package, allowing users to seamlessly include batch effect cor-rection in a pmartR workflow without needing any additional data manipulation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data from a multi-year targeted proteomics study of a longitudinal birth cohort of type 1 diabetes

The deployment of liquid chromatography-mass spectrometry-based plasma proteomics experiments in a large cohort is sparse, leading to a lack of data available for benchmarking, method development or validation. Comprised of 6,426 plasma analyses, The Environmental Determinants of Diabetes in the Young (TEDDY) proteomics validation study constitutes one of the largest targeted proteomics experiments in the literature to date. The proteomics data from this study were generated over the course of 2.5 years from over 900 study subjects, each providing up to 29 longitudinal samples. The data also includes 916 quality control samples. The targeted mass spectrometry assay was comprised of 694 peptides mapping to 167 proteins and the panel was measured in each subject and QC sample. The targeted proteomic dataset presented here can be used as a resource for new computational method development, such as for batch correction, as well as for benchmarking and comparing the performance of different methods/tools.

60 APPLIED LIFE SCIENCES↗

PNNL-Predictive-Phenomics/ProteoMeter

ProteoMeter is a Python package that assists in the statistical analysis of global proteomics, protein post-translation modification (PTM), and limited proteolysis (LiP) data. It contains batch correction, normalization, and statistical testing methods, as well as functions that "roll up" peptide-level data to the single-site level. It has a robust user configuration system, allowing it to flexibly integrate different types of experiment designs. For basic usage, a simple configuration file provides the essential functionality. Advanced users have access to the entire statistical pipeline for fine-tuning analyses. Processed data is easily exported to many common spreadsheet and data-frame formats.

Rozum, Jordan [Pacific Northwest National Lab]↗

In‐situ Analysis of Paste Properties in Resonant Acoustic Mixers for Quality Monitoring

Formulation control is key to achieving consistent target properties of energetic materials, as feedstock variations and slight deviations in the ratios of different ingredients can have major effects on final product properties, particularly in dense pastes with high particle loading >65 vol.%. In large‐scale operations, it is imperative to either correct or remove batches of material that perform outside baseline property specifications as early as possible to avoid unnecessary processing of suboptimal material. Quality monitoring is the practice of measuring material properties during processing using process analytical technologies as opposed to only testing the properties of the final product; it is a key principle in the quality‐by‐design frameworks used for designing formulations and manufacturing processes. Herein, a process analytical technology method for correlating material properties of dense pastes directly after mixing in a Resonant Acoustic Mixer to motor data is developed and used to detect differences in the particle content of dense paste formulations. This method was also capable of detecting variations in powder feedstock properties, such as particle packing efficiency, and is sensitive enough to detect changes of 2 wt.% in the total solids content of the formulation. The techniques presented herein show excellent promise for use as a process analytical technology capable of quantifying formulation effects on material movement modes during resonant acoustic mixing.

Materials science↗

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Real-Time Automated pH Control within Batch Processes Relying on Raman pH Measurement

Nuclear fission is an energy source that can provide consistent power with very low associated carbon emissions. However, management of the used nuclear fuel is an important aspect of the application of nuclear power. Recycling of useful components from used fuel is an attractive option, but this involves chemical processing of the fuel. Possible chemical separation technologies that might be used in this regard are sensitive to solution pH. Raman spectroscopy is a promising technique for monitoring the pH of solutions in real time. Classical pH probes are too fragile to be used in the harsh environments encountered in nuclear fuel processing. Raman probes are robust and can withstand these harsh environments to track pH. Coupled with chemometric analysis, the demonstration of the use of Raman spectroscopy to track and predict the pH in carboxylate-buffered systems is made possible. Utilizing this spectroscopy in conjunction with Programmable Logic Controllers mimics industrial control systems used in many modern industrial settings. This showcases a pragmatic approach toward leveraging Raman spectroscopy and chemometric model outputs as inputs for a real-time control system. The model to predict pH created by chemometrics proved to be successful in tracking pH. The optimal pH for TALSPEAK extraction of lanthanides and actinides from aqueous solution is known to proceed in a narrow pH range of around pH = 2.8 ± 0.1. This study uses Raman optical monitoring and automated control to return and maintain solution pH within this range after acid or base perturbations move the solution pH well outside this region. Root-mean-square errors show that pH changes measured using Raman spectroscopy on the batch process solution are reliably measured and used to automatically correct and maintain solution pH. Measurement of solution pH tracks favorably with electrochemical pH probe comparison measurements. As a result, the ability to showcase Raman spectroscopy paired with chemometrics analysis acts as a durable, better alternative data source compared to traditional pH probes to optimize the separation efficiency in the used nuclear fuel processing.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerated 133 Xe Quantification in Samples Containing Significant 133 mXe

The quantification of 133 Xe in the presence of its mother radionuclide 133 mXe requires the full quantification of both to perform the ingrowth correction for 133 Xe. Due to the nature of both of these radionuclides, the 133 mXe requires significantly more time to quantify by High Purity Germanium (HPGe) detectors due to lower production yields, lower gamma emission probabilities, and lower detection efficiencies. This work shows that 133 Xe and 133 mXe quantification can be accelerated by measuring the 133m:133 activity ratio for a large batch of material and applying this activity ratio to assays of lower activity subsamples of the same batch of material. Included in this report are derivations of the required decay correction equations, and experiments using actual samples to validate the performance of these equations. A detector calibration method is also shown that leverages this method as an alternative to existing calibration methods for 133 mXe quantification.

133mXe↗

Identifying snowfall elevation patterns by assimilating satellite-based snow depth retrievals

Precipitation in mountain regions is highly variable and poorly measured, posing important challenges to water resource management. Traditional methods to estimate precipitation include in-situ gauges, Doppler weather radars, satellite radars and radiometers, numerical modeling and reanalysis products. Each of these methods is unable to adequately capture complex orographic precipitation. Here, we propose a novel approach to characterize orographic snowfall over mountain regions. We use a particle batch smoother to leverage satellite information from Sentinel-1 derived snow depth retrievals and to correct various gridded precipitation products. This novel approach is tested using a simple snow model for an alpine basin located in Trentino Alto Adige, Italy. Here, we quantify the precipitation biases across the basin and found that the assimilation method (i) corrects for snowfall biases and uncertainties, (ii) leads to cumulative snowfall elevation patterns that are consistent across precipitation products, and (iii) results in overall improved basin-wide snow variables (snow depth and snow cover area) and basin streamflow estimates.

54 ENVIRONMENTAL SCIENCES↗

(U) Correlated Sampling Using Batch Statistics to Reduce the Uncertainty of Combinations of KSEN Outputs with MCNP6

The relative sensitivity of k eff to the densities of nuclides in a material are combined to compute relative sensitivities to other inputs. When computed in a single Monte Carlo run, the nuclide density sensitivities are correlated, and the statistical uncertainties propagated to other inputs will be incorrect unless those correlations are accounted for. Equations are presented to apply correlated sampling using batch statistics for the sum of an arbitrary number of random tallies and the difference of two random tallies when each is multiplied by a different constant. When correlated sampling is used on a recent benchmark evaluation, the correct statistical uncertainties for certain combinations of sensitivities are dramatically smaller than the incorrect uncertainties. A user-controlled, regular output of MCNP6’s KSEN sensitivities that allows batch statistics to be applied to combinations would be extremely valuable.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SuperNu Version 4.x

We seek to release SuperNu, Version 4.x, as a continuation of development for the open source SuperNu software. The SuperNu, Version 3.x Monte Carlo radiative transfer code for astrophysical transients was released with GPLv3 copyright, asserted by LANL in 2015. For the next release we have features planned for development , including: opacity implementation (including non-local thermodynamic equilibrium effects), generalized source implementation (e.g. for emulating shock heating in Type II supernovae), 3T (electron, ion, radiation) internal energy update, special relativity corrections through O(v^2/c^2), light polarization (e.g. for comparison to spectroplarimetry observations of supernovae and kilonovae), and infrastructure features (checkpoint and restart of simulations, and tools including setup and Slurm batch scripts and simulation post-processing/analysis scripts). These features are intended to improve the fidelity and/or better understand uncertainty in supernova and kilonova light curve calculations.

Wollaeger, Ryan↗

Humate Amendment Injection Viability Testing – Supplementary Batch Mixture and Soil Column Studies

Groundwater in the Lost Lake Aquifer Zone (LLAZ) in the Southern Sector of the M-Area Hazardous Waste management Facility (HWMF) is contaminated with chlorinated ethenes, including trichloroethylene (TCE) and tetrachloroethylene (PCE). Treatment of contaminated groundwater with humic acid is being evaluated as a potential corrective action for these volatile organic compounds (VOCs) in the LLAZ in Southern Sector (SRNS, 2019). Pilot scale injections of Huma-K brand humic acid for groundwater treatment were previously performed in M-Area between 2017 and 2020 (Amidon, 2023). Groundwater was extracted from the Lost Lake Aquifer Zone (LLAZ), then a solution of humate was mixed with recovered groundwater and subsequently re-injected into the same aquifer unit. Operations of the historical humate pilot testing were discontinued in 2020 due to COVID-19 restrictions, low injection rates, and recirculation fluid spillage. No further field scale or laboratory research to investigate humate amended groundwater injection was performed at the SRS until 2024. At the request of Area Completion Projects (ACP), the Savannah River National Laboratory (SRNL) performed supplementary testing to help determine the viability of further injection of humate as a remedial option for VOCs within the LLAZ in the Southern Sector of the M-Area HWMF.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Tetrafluoroboric Acid Digestion for Accurate Determination of Rare Earth Elements in Coal and Fly Ash by ICP-MS Analysis

Coal and coal-related fly ash often contain rare earth elements (REEs) that have the potential to be utilized as valuable mineral resources. Accurately determining the REE content in coal and fly ash is crucial for resource evaluation. The conventional approach involves using hydrofluoric acid (HF) to dissolve silicates and release REEs, which, however, prolongs the digestion process due to the additional step of complexing fluoride ions (F−) with boric acid (H3BO3). Determining the correct amount of H3BO3 for neutralization can be challenging, and in some instances, the binding of fluoride ions with certain lanthanides (Lns) hampers the accurate determination of all 14 naturally occurring rare earth elements in a single digestion batch by inductively coupled plasma mass spectrometry (ICP-MS). In this study, we present an alternative method that achieves the accurate determination of all 14 naturally occurring REEs using tetrafluoroboric acid (HBF4) followed by ICP-MS analysis. This approach eliminates the need for an F− complexing step. We tested this method on certified REE reference materials, including NIST 1632e (coal) and NIST 1633c (fly ash), as well as the REE geological reference material USGS AGV-1 (andesite). Our results demonstrated excellent recovery rates (relative standard deviation, RSD < ±10%), with a correlation coefficient (r2) exceeding 0.99. Using this method, we investigated the concentrations of all 14 REEs in coal and fly ash samples collected from various locations in the southwestern USA. This improved digestion technique streamlines the analysis process and enhances the accuracy of REE determination, facilitating a more comprehensive evaluation of REE-rich coal and fly ash deposits for resource exploration.

58 GEOSCIENCES↗

A Zero-Emission Process for Direct Reduction of Iron by Hydrogen Plasma in a Rotary Kiln Reactor

This project’s goal was to demonstrate a hydrogen plasma (H-plasma)-rotary kiln process for reducing iron ore to iron as part of the steel manufacturing process. The H-plasma provides a greater thermodynamic driving force for reducing iron ores than thermal processes such as the DRI process, enabling lower reaction temperatures. We estimated that our process technology can reduce energy consumption by 45% compared to the blast furnace process and ~15% compared to the DRI process. Steel manufacturing produces about 1.8 tons of CO2/ton of steel with iron ore reduction accounting for about one-third of the CO2 produced in the overall manufacturing process. We estimated our process has the potential to reduce GHG emissions from ironmaking by 35% with today’s grid and by up to 88% with a future low-carbon grid while being cost competitive with the current blast furnace route. We demonstrated reduction of hematite and magnetite rich materials at temperatures from 600 to 800°C. We achieved 90-95% metallization on 100 gr samples in batch reduction experiments in the H-plasma rotary kiln furnace at 600-650°C. Attempts to perform tests in a continuous operation mode identified problems with the ore feed mechanism. We identified solutions but there was not time nor budget to correct these for this project

36 MATERIALS SCIENCE↗

A Zero-Emission Process for Direct Reduction of Iron by Hydrogen Plasma in a Rotary Kiln Reactor

This project’s goal was to demonstrate a hydrogen plasma (H-plasma)-rotary kiln process for reducing iron ore to iron as part of the steel manufacturing process. The H-plasma provides a greater thermodynamic driving force for reducing iron ores than thermal processes such as the DRI process, enabling lower reaction temperatures. We estimated that our process technology can reduce energy consumption by 45% compared to the blast furnace process and ~15% compared to the DRI process. Steel manufacturing produces about 1.8 tons of CO2/ton of steel with iron ore reduction accounting for about one-third of the CO2 produced in the overall manufacturing process. We estimated our process has the potential to reduce GHG emissions from ironmaking by 35% with today’s grid and by up to 88% with a future low-carbon grid while being cost competitive with the current blast furnace route. We demonstrated reduction of hematite and magnetite rich materials at temperatures from 600 to 800°C. We achieved 90-95% metallization on 100 gr samples in batch reduction experiments in the H-plasma rotary kiln furnace at 600-650°C. Attempts to perform tests in a continuous operation mode identified problems with the ore feed mechanism. We identified solutions but there was not time nor budget to correct these for this project

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Efficient Parameterization of Density Functional Tight-Binding for 5 f -Elements: A Th–O Case Study

Density functional tight binding (DFTB) models for f-element species are challenging to parametrize owing to the large number of adjustable parameters. The explicit optimization of the terms entering the semiempirical DFTB Hamiltonian related to f orbitals is crucial to generating a reliable parametrization for f-block elements, because they play import roles in bonding interactions. However, since the number of parameters grows quadratically with the number of orbitals, the computational cost for parameter optimization is much more expensive for the f-elements than for the main group elements. In this work we present a set of efficient approaches for mitigating the hurdle imposed by the large size of the parameter space. A novel group-by-orbital correction functions for two-center bond integrals was developed. With this approach the number of parameters is reduced, and it grows linearly with the number of elements, maintaining the accuracy and the number of parameters, in the case of f elements, by more than 40%. The parameter optimization step was accelerated by means of the mini-batch BFGS method. This method allows parameter optimizations with much larger training sets than other single batch methods. A stochastic optimizer was employed that helped overcome shallow local minima in the objective function. The proposed algorithm was used to parametrize the DFTB Hamiltonian for the Th–O system, which was subsequently applied to the study of ThO 2 nanoparticles. The training set consisted of 6322 unique structures, which is barely feasible with conventional optimization methods. The optimized parameter set, LANL-ThO, displays good agreement with DFT-calculated properties such as energies, forces, and structures for both clusters and bulk ThO 2 . Benefiting from the fewer number of parameters and lower computational costs for objective function evaluations, this new approach shows its potential applications in DFTB parametrization for elements with high angular momentum, which present a challenge to conventional methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Direct Feed High-Level Waste APPS Model Glass Testing (DFHLW APPS) Matrix

This report summarizes the data collected during the batching and melting of the Direct Feed High-Level Waste APPS Model Glass Matrix (DFHLW APPS) to serve as a quality-assured validation of the Aspen Process Performance Simulation (APPS) formulation method. Of 15 glasses tested, 12 satisfied all target property constraints. Two glasses, APPS-05 and -06, formed nepheline on canister centerline cooling heat-treatment and failed the Product Consistency Test response limits. Glass APPS-07-2 formed unacceptably high concentrations of crystals (primarily Na3Nd(PO4)2) when heat treated at 950 °C. All other glasses were found to be satisfactory. The measured property values were compared to predicted values from a set of current models. In many cases the current models were found to be inadequate for design of DFHLW glasses. These models are being adjusted to correct for mispredictions. Other models, e.g., density, toxicity characteristic leaching procedure, and sulfur solubility, are adequate for formulation of DFHLW glasses.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗