Search NASASearch

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka

Integrating Analytical Solutions and U-Net Model for Predicting Groundwater Contaminant Plumes in Pump-and-Treat Systems

Pump-and-treat (P&T) is a common technique for groundwater remediation involving the extraction and treatment of contaminated water above ground. Optimizing the design and operation of the P&T well network is essential for maximizing the system’s effectiveness and efficiency. However, this optimization often necessitates many model evaluations, leading to computationally demanding tasks. This study introduces a novel approach that integrates analytical solutions for groundwater dynamics with the U-Net (Ronneberger et al., 2015) deep learning framework to predict groundwater contaminant plume migration under dynamic pumping conditions. By incorporating the Thiem equation (Thiem, 1906) into the input preprocessing, the U-Net model transforms sparse well data into a continuous spatial field that captures the hydraulic impacts of pumping activities. This integration enables the model to leverage both deep learning capabilities and classical physics-based groundwater theories, enhancing prediction accuracy and computational efficiency. These advancements can facilitate rapid, large-scale evaluations of P&T optimization simulations, allowing for timely and effective decision-making in well placement and system management. We demonstrate the model's robust performance across both simplified transient 2D models and a more complex 3D heterogeneous site model at the 200 West P&T facility at the Hanford Site. The U-Net-based model offers substantial computational advantages, reducing simulation times significantly compared to full physics-based models and providing a powerful tool for rapid site evaluation and P&T system optimization, such as evaluating alternative P&T well network designs. Our findings highlight the potential of advanced machine learning models to significantly enhance the efficiency and sustainability of groundwater remediation efforts, offering a novel application of U-Net architecture in environmental science.

Pump-and-treat

Transfer learning-based soybean LAI estimations by integrating PROSAIL, UAV, and PlanetScope imagery

Accurate Leaf Area Index (LAI) estimations at the soybean plot scale is achievable using high-resolution Unmanned Aerial Vehicle (UAV) imagery and field measurement samples. However, the limited coverage of UAV flights restricts large-scale remote sensing monitoring in expansive soybean fields. This study leverages the broad coverage and 3-m resolution of PlanetScope satellite imagery to extend LAI prediction from UAV to satellite scales through transfer learning, using UAV-scale LAI estimates as a benchmark to validate cross-scale consistency. To address this challenge, this study proposed the LAI-TransNet, a two-stage transfer learning framework designed for precise and scalable soybean LAI prediction across large areas, demonstrating its effectiveness in cross-scale monitoring. In Stage 1, a UAV-scale benchmark is established using PROSAIL-simulated UAV reflectance data (UAV-Sim) and field-measured soybean LAI. Traditional machine learning, deep learning, and transfer learning models are trained on a hybrid UAV-Sim and field-measured dataset (UAV-Sim_Measured), with the transfer learning model CNN-TL, fine-tuned using pre-trained weights derived from UAV-Sim, achieving the highest accuracy (R 2 = 0.81, RMSE = 0.64 m 2 /m 2 , rRMSE = 11.5 %). In Stage 2, LAI-TransNet is developed by fine-tuning the CNN-TL model on PlanetScope simulated data (PS-Sim), preprocessed via cross-domain mapping to align UAV and satellite spectral features. Real PlanetScope imagery is corrected for reflectance consistency with reference to UAV imagery spectral profiles. LAI-TransNet outperforms other deep learning models trained directly on PS-Sim (R 2 = 0.69 vs. 0.60–0.63), ensuring robust cross-scale consistency. In conclusion, by bridging UAV and satellite scales, LAI-TransNet enables large-scale soybean LAI monitoring, enhancing precision agriculture management through improved monitoring with the PlanetScope imagery.

Leaf area index (LAI)

Enhancing biomass flowability for entrained flow Gasification: The role of densification and torrefaction

Gasification presents a key strategy in addressing future energy demands while minimizing environmental impact. This has been recognized as a promising method to convert biomass to higher value products such as biofuels or hydrogen. Among gasification technologies, high-temperature and high-pressure reactors, particularly the R-GAS® system, emerge as an advanced option boasting superior conversion efficiency. However, akin to conventional high-temperature and high-pressure gasifiers, R-GAS® necessitates small particle sizes for optimal carbon conversion, a requirement yet to be fully explored for biomass. Hence, this study investigated the effectiveness of combined mechanical and thermal preprocessing techniques in modifying the physicochemical properties of biomass to suit gasification systems. Mechanical techniques including densification and pulverization, alongside thermal techniques such as torrefaction and steam explosion, were examined. The results demonstrate that torrefaction fosters producing of uniform granular material, enhancing flowability and reducing energy requirements for pulverization compared to steam explosion. Notably, torrefied corn stover exhibited lower internal friction angles and effective cohesion (40.09 ± 0.22° and 0.56 ± 0.01 kPa, respectively) compared to steam exploded corn stover (41.87 ± 0.65° and 0.83 ± 0.06 kPa, respectively), indicative of improved flowability. Additionally, pulverization of torrefied corn stover required approximately 16 % less energy than steam exploded corn stover and 91 % less energy than raw corn stover. Furthermore, the torrefaction-induced alterations in particle size, shape, and packing densities emphasize its potential to optimize flow and handling processes for gasification. These findings underline that densification followed by torrefaction effectively addresses biomass variability, leading to more efficient and sustainable energy conversion.

09 - BIOMASS FUELS

Optical image analysis of WSe 2 − thresholding for layer detection

The fast and reliable layer identification of two-dimensional transition metal dichalcogenide (TMD), such as WSe 2 , is essential to investigating their thickness-dependent electronic and optical properties. This article presents efficient optical image thresholding methodology designed to segment the mono, bi, and tri-layer regions of WSe 2 flakes mechanically exfoliated onto a SiO 2 /Si substrate. The optical images were first preprocessed to exclude the background effect and analyzed using the pixel medians and interquartile ranges for fundamental color channels—red, green, and blue (RGB). The analysis of red channel pixel intensities yielded three distinct ranges, serving as thresholds for layer segmentation: monolayer (111.0–118.0), bilayer (103.0–110.0), and tri-layer (93.0–103.0). Similarly, thresholds were established for each color channel, facilitating a comparative study of the segmentation performances. Further, the intersection-over-union ($IoU$) calculations revealed that the red and green channels demonstrated greater than 99 % and 90 % accuracy in differentiating each layer, respectively. This approach yields remarkable results without substantial data calibration that utilizes time-intensive heuristic techniques. Moreover, the proposed methodology offers the flexibility to compare performances across different color channels, expanding the applicability for other 2D material systems.

2D Materials

Adaptive continuity-preserving simplification of street networks

Street network data is widely used to study human-based activities and urban structure. Often, these data are geared towards transportation applications, which require highly granular, directed graphs that capture the complex relationships of potential traffic patterns. While this level of network detail is critical for certain fine-grained mobility models, it represents a hindrance for studies concerned with the morphology of the street network. For the latter case, street network simplification — the process of converting a highly granular input network into its most simple morphological form — is a necessary, but highly tedious preprocessing step, especially when conducted manually. In this manuscript, we develop and present a novel adaptive algorithm for simplifying street networks that is both fully automated and able to mimic results obtained through a manual simplification routine. The algorithm — available in the neatnet Python package — outperforms current state-of-the-art procedures when comparing those methods to manually, human-simplified data, while preserving network continuity.

Python

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Spatial and temporal characterization of municipal solid waste based on resource recovery pathways

This study presents a two-year, quarterly assessment of MSW across four source sectors (residential, schools, restaurants, and grocery stores) from sixteen sites across five U.S. states. MSW was manually sorted into 27 categories and aggregated into pathway fractions: high-moisture (HM) organics, low-moisture (LM) organics, recyclable (RC) materials, and residuals for disposal. Organics represented 89 % of the MSW stream. The largest fraction was HM organics consisting of food waste (31 %) and yard waste (3 %), with large coefficient of variations (CV), 79 and 278 %, respectively, reflecting high seasonal and site variability that varied significantly (p < 0.01) across sampling periods. The HM fraction showed properties favorable for anaerobic digestion, with moisture content ranging from 56 to 95 % and volatile solids ranges of 86-95 %. In contrast, the LM and RC fractions remained more stable (plastics CV = 41 %; paper CV = 53 %) with heating values up to 26.9 MJ/kg across sources, reflecting suitability for gasification. Microstructural analysis revealed less porosity in residential waste sampled at the landfill, which can influence preprocessing efficiency and microbial accessibility. Pathway informed allocations showed that 35 % of MSW is suitable for anaerobic digestion, 36 % for gasification, and 18 % for recycling, leaving 11 % requiring landfill disposal. These results provide quantitative evidence to determine feedstock allocation, waste-to-energy system design, and the development of data-driven sustainability and resource recovery strategies within a circular bioeconomy.

09 BIOMASS FUELS

Designing resilient IoT and Edge Computing with federated tinyML

The rapid growth of the Internet of Things (IoT) and Edge Computing (EC) has brought significant conveniences to modern society but has also greatly expanded the cyber attack surfaces, particularly as these technologies are being increasingly integrated into critical systems such as power grids, healthcare, and smart homes. Here, to improve IoT/EC’s cybersecurity posture, we leveraged Artificial Intelligence (AI) and Machine Learning (ML) by employing tinyML to monitor voluminous IoT data for cyber threats while addressing devices’ resource constraints, and utilizing Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. Building on our three-layer architecture combining tinyML and FL to enhance autonomous cyber attack detection, this paper demonstrated that the architecture improves detection accuracy, reduces resource consumption, and enables lightweight, secure IoT device monitoring. These results were validated using the public N-BaIoT dataset as well as real IoT network traffic data collected under multiple attack scenarios from our testbeds. Additionally, we introduced an enhanced FL methodology with a novel preprocessing stage, including federated feature selection and global preprocessor construction, to address IoT/EC data heterogeneity. We developed a physical IoT testbed for attack simulations and data collection, implemented a tinyML-powered detector for realistic model validation, and also built a virtual testbed for scalable evaluations of FL models across diverse network environments.

Cognitive cyber

Absorption Correction for Reliable Pair Distribution Functions from Low Energy X-ray Sources

This paper explores the development and testing of a simple absorption correction model for processing powder X-ray diffraction data from Debye−Scherrer geometry laboratory X-ray experiments. This may be used as a preprocessing step before using PDFGETX3 to obtain reliable pair distribution functions (PDFs). Various experimental and theoretical methods for estimating μR were explored, and the most appropriate μR values for correction were identified for different capillary diameters and X-ray beam sizes. We identify operational ranges of μR where a reasonable signal-to-noise ratio is possible after correction. A user-friendly software package, DIFFPY.LABPDFPROC, is presented that can help estimate μR and perform absorption corrections with a rapid calculation for efficient processing.

Absorption

Quantifying Impacts of Biomass Pelletization on Fast Pyrolysis Using a Single-Particle Reactor, X-ray Computed Tomography, and Computational Modeling

The pore structure and density of lignocellulosic feedstocks dictate intraparticle transport phenomena and thereby play an important role in thermochemical conversion processes such as fast pyrolysis for biofuel and biochemical production. Variations in microstructure are inherent from different biomass species and can be introduced by preprocessing techniques such as cutting and pelletization. Morphological changes also occur during conversion and lead to vastly different pore structures and behavior during pyrolysis, which impact required conversion times and product distributions. The current work presents a comprehensive comparison of fast pyrolysis of neat and pelletized pine feedstocks, which includes single-particle experiments, modeling, and 3D imaging by X-ray computed tomography (XCT). The particle-scale model included anisotropic heat and mass transport in a shrinking particle with pyrolysis reactions based on the CRECK mechanism with boundary conditions informed by reactor-scale simulations of the single-particle reactor. The models were validated by measurements of the temperature and mass loss from single-particle pyrolysis experiments of neat and pelletized pine. Quantitative analysis of XCT geometries revealed that pyrolytic conversion yielded chars with increased porosity and permeability compared to the unpyrolyzed materials, along with decreased tortuosity and anisotropy. Pelletization of the pine feedstock resulted in a much denser, less permeable material, which converted slower and produced more residual char after pyrolysis compared to neat pine. The results from particle modeling revealed that accounting for the dynamic and anisotropic heat and mass transport caused by differences in pore structure is critical to achieving agreement with experimental results. Overall, this study highlights the dramatic differences in conversion behavior imparted by pelletization and the importance of capturing microstructural attributes in computational models to guide the design and optimization of pyrolysis processes for specific biomass feedstocks.

09 BIOMASS FUELS

Process Feasibility Analysis of Waste Biomass Valorization to Biochar and Bio-Oil via Slow and Fast Pyrolysis

The United States has abundant biomass and waste feedstock to support the nation's energy addition and affordability targets. Pyrolysis, a thermochemical conversion process, decomposes lignocellulosic feedstocks into liquid, solid, and gaseous fuels that can contribute to the domestic production of biofuels, biopower, and bioproducts. Growing private sector interest in this technology is a key motivation for this comprehensive techno-economic process modeling analysis of a respective biorefinery that includes feedstock preprocessing, slow and fast pyrolysis, and product separation to bio-oil, biochar, and syngas hydrocarbons. Results show that biochar from slow pyrolysis could achieve minimum selling prices (MSPs) of $\$$188-$\$$260/t, competitive with reported market values, while bio-oil from fast pyrolysis is estimated to yield MSPs of $\$$6.49-$\$$9.68/GGE, approximately twice conventional fuel benchmarks. Sensitivity analysis identifies feedstock cost, product yield, and scale as primary cost drivers, while scenarios involving biochar carbon credits and high value applications may substantially improve economics. Overall, these results suggest that continued innovation in feedstock logistics, process integration, and market development will be critical to achieving economically viable and scalable bioproducts.

09 BIOMASS FUELS

Aggregation Methods for Quantifying PTM and Structural Changes in Bottom-Up Proteomics

Bottom-up proteomic workflows rely on sequential preprocessing steps, commonly including peptide-to-protein aggregation (“roll-up”), to enhance data reliability and interpretability. While roll-up is effective for protein-centered analyses, it may be suboptimal for applications focused on post-translational modifications (PTMs) or protein structural changes, such as limited proteolysis–mass spectrometry (LiP-MS). Here, we investigate how different roll-up strategies influence site-level quantification in PTM differential analysis. Moreover, we introduce a novel site-centric roll-up approach tailored for LiP-MS, which quantifies proteolytic fragments rather than solely tryptic peptides. We benchmark these methods through simulation studies, comparing their sensitivity and specificity in detecting structural and PTM-driven changes. We found that the median and mean roll-up methods outperform the sum method in both PTM and LiP proteomics, and site-level quantification in LiP outperforms peptide-level quantification. Our findings offer the first systematic, data-driven guidance for selecting roll-up techniques in site-level proteomic analyses, with implications for both PTM-focused and structural proteomics studies.

aggregation

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE

Explainable tokamak-agnostic forecasting of fusion plasma instability via megahertz turbulent fluctuations

Scientific applications of artificial intelligence (AI) often remain limited by device-specific training and unexplained “black-box” approaches, creating fundamental barriers to cross-system generalization. This challenge is critical for nuclear fusion, where future reactors will have limited operational data for AI training. Here, we demonstrate that our neural network, trained solely on megahertz-scale turbulence measurements from one machine (DIII-D), forecasts Type-I edge localized mode (ELM) onsets in a different tokamak (KSTAR) through zero-shot weight transfer following physics-consistent preprocessing without device-specific retraining. Through an explainable AI framework combining gradient-weighted class activation mapping with physics validation, we reveal that our network can internalize physics relationships governing the ELM instabilities rather than memorizing device-specific patterns. The network perceives spatiotemporal features that correlate consistently with independently calculated instability growth rates, magnetohydrodynamic stability limits, and pedestal structure dynamics. Statistical analyses of dimensionally-reduced saliency features reveal the identical triangular features between the saliency representations, instability growth rates, and prediction probability across tokamaks, providing evidence that our forecasting system can show tokamak-agnostic generalization. This work contributes to a foundation for explainable scientific AI systems, where cross-system developments are essential for transcending traditional domain-specific constraints.

AI

ChemPren: a new and economical technology for conversion of waste plastics to light olefins

With the ever-increasing demand for plastics, sustainable recycling methods are key necessities. Here, the current plastics industry can manage to recycle only 10% of the 400 million metric tons of plastic produced globally. Waste plastics, in the current infrastructure, land up mostly in landfills. Although a lot of research efforts have been spent on processing and recycling co-mingled mixed plastics, energy-efficient sustainable and scalable routes for plastic upcycling are still lacking. Catalytic valorization of waste plastic feedstock is one of the potential scalable routes for plastic upcycling. Silica-alumina based materials, and zeolites have shown a lot of promise. A major interest lies in restricting catalyst deactivation, and refining product selectivity and yield for such catalytic processes. This article highlights ChemPren technology as a clean energy solution to waste plastic recycling. Co-mingled, mixed plastic feedstock along with spray dried, attrition resistant, ZSM-5 containing catalysts is preprocessed with an extruder to form optimally sized particles and fed into a fluidized bed reactor for short contact times to produce selectively and in high yields ethylenes, propylenes and butylenes. This techno-economic perspective indicates that the ChemPren technology can produce propylene at $\$$0.16 per lb, whereas the current selling price of virgin propylene is $0.54 per lb. This technology can serve as a platform for mixed plastic upcycling, with more advancements necessary in the form of robust and resilient catalysts and reactor operation strategies for tuning product selectivity.

25 - ENERGY STORAGE