Search NASASearch

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Environmental Quenching of Low-surface-brightness Galaxies Near Hosts from Large Magellanic Cloud to Milky Way Mass Scales

Low-surface-brightness galaxies (LSBGs) are excellent probes of quenching and other environmental processes near massive galaxies. We study an extensive sample of LSBGs near massive hosts in the local universe that are distributed across a diverse range of environments. The LSBGs with surface-brightness ${\mu }_{\mathrm{eff},{g}}\gt 24.2\,\mathrm{mag}\,{\mathrm{arcsec}}^{-2}$ are drawn from the Dark Energy Survey Year 3 catalog while the hosts with masses $9.0\lt \mathrm{log}({{ \mathcal M }}_{\star }/{M}_{\odot })\lt 11.0$ comparable to the Milky Way and the Large Magellanic Cloud are selected from the z0MGS sample. We study the projected radial density profiles of LSBGs as a function of their color and surface brightness around hosts in both the rich Fornax–Eridanus cluster environment and the low-density field. We detect an overdensity with respect to the background density, out to 2.5 times the virial radius for both hosts in the cluster environment and the isolated field galaxies. When the LSBG sample is split by g − i color or surface brightness μ eff, g , we find the LSBGs closer to their hosts are significantly redder and brighter, like their high-surface-brightness counterparts. The LSBGs form a clear “red sequence” in both the cluster and isolated environments that is visible beyond the virial radius of the hosts. This suggests preprocessing of infalling LSBGs and a quenched backsplash population around both host samples. More so, the relative prominence of the “blue cloud” feature implies that preprocessing is ongoing near the isolated hosts compared to the cluster environment where the LSBGs are already well processed.

79 ASTRONOMY AND ASTROPHYSICS

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES

Vacuum Neutral Transport Model in UEDGE for Tokamak Far Scrape‐Off Layer

A model for neutral transport in the far scrape-off layer (SOL) vacuum region (vacuum neutral model) has been developed and implemented in UEDGE. Free-streaming neutral trajectories between the outermost UEDGE boundary and the vessel wall are preprocessed using DEGAS2 to construct a tele-transport matrix that captures non-local neutral relocation through the vacuum. Here, this matrix is then used in UEDGE as a non-local boundary condition for the neutral equations, preserving the robustness and convergence of the implicit solver without introducing statistical noise. Simulations of a DIII-D lower single-null configuration show that the vacuum neutral model relocates neutrals from the divertor to the upstream region, increasing the outer midplane separatrix density required for detachment onset by about 30%. However, the characteristic target temperature at detachment (T e, osp ~ 3 - 4eV) and the radiation front behavior remain unchanged.

DEGAS2

Laser powder bed fusion parameter estimation with k-NN

Abstract Laser powder bed fusion (L-PBF) is a technique within additive manufacturing that uses a high power density laser to build parts from fused powdered metal alloy. This technology is well equipped to produce complex parts with otherwise impossible features, such as hidden voids or lattice structures. Alongside capability, reliability and quality are key characteristics considered when choosing a manufacturing method, and these are gaining attention as this method becomes more prevalent in industry. One main indicator of a stable L-PBF process is consistent melt pool geometry, and the properties of which are likely to determine the quality of the part produced. As computing power and sensing technologies become more advanced, this melt pool geometry could be studied in real time. This work addresses the challenge by leveraging a k-nearest neighbor (k-NN) model to identify key features within melt pool imagery and predict the energy density. The k-NN model was trained on data provided by the National Institute of Standards and Technology (NIST). Data preprocessing was performed on the images to extract features that were used in the k-NN model. This approach was used to accurately infer the energy density of unseen layers within the same part. The algorithm was subsequently tested with unique scan strategies and found to reasonably estimate the energy density of different parts. A fivefold cross validation found the algorithm to be consistently predicting the class of 91.4% of the in situ melt pool images.

Jung, Patrick (ORCID:0000000267890859)

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka

Integrating Analytical Solutions and U-Net Model for Predicting Groundwater Contaminant Plumes in Pump-and-Treat Systems

Pump-and-treat (P&T) is a common technique for groundwater remediation involving the extraction and treatment of contaminated water above ground. Optimizing the design and operation of the P&T well network is essential for maximizing the system’s effectiveness and efficiency. However, this optimization often necessitates many model evaluations, leading to computationally demanding tasks. This study introduces a novel approach that integrates analytical solutions for groundwater dynamics with the U-Net (Ronneberger et al., 2015) deep learning framework to predict groundwater contaminant plume migration under dynamic pumping conditions. By incorporating the Thiem equation (Thiem, 1906) into the input preprocessing, the U-Net model transforms sparse well data into a continuous spatial field that captures the hydraulic impacts of pumping activities. This integration enables the model to leverage both deep learning capabilities and classical physics-based groundwater theories, enhancing prediction accuracy and computational efficiency. These advancements can facilitate rapid, large-scale evaluations of P&T optimization simulations, allowing for timely and effective decision-making in well placement and system management. We demonstrate the model's robust performance across both simplified transient 2D models and a more complex 3D heterogeneous site model at the 200 West P&T facility at the Hanford Site. The U-Net-based model offers substantial computational advantages, reducing simulation times significantly compared to full physics-based models and providing a powerful tool for rapid site evaluation and P&T system optimization, such as evaluating alternative P&T well network designs. Our findings highlight the potential of advanced machine learning models to significantly enhance the efficiency and sustainability of groundwater remediation efforts, offering a novel application of U-Net architecture in environmental science.

Pump-and-treat

Transfer learning-based soybean LAI estimations by integrating PROSAIL, UAV, and PlanetScope imagery

Accurate Leaf Area Index (LAI) estimations at the soybean plot scale is achievable using high-resolution Unmanned Aerial Vehicle (UAV) imagery and field measurement samples. However, the limited coverage of UAV flights restricts large-scale remote sensing monitoring in expansive soybean fields. This study leverages the broad coverage and 3-m resolution of PlanetScope satellite imagery to extend LAI prediction from UAV to satellite scales through transfer learning, using UAV-scale LAI estimates as a benchmark to validate cross-scale consistency. To address this challenge, this study proposed the LAI-TransNet, a two-stage transfer learning framework designed for precise and scalable soybean LAI prediction across large areas, demonstrating its effectiveness in cross-scale monitoring. In Stage 1, a UAV-scale benchmark is established using PROSAIL-simulated UAV reflectance data (UAV-Sim) and field-measured soybean LAI. Traditional machine learning, deep learning, and transfer learning models are trained on a hybrid UAV-Sim and field-measured dataset (UAV-Sim_Measured), with the transfer learning model CNN-TL, fine-tuned using pre-trained weights derived from UAV-Sim, achieving the highest accuracy (R 2 = 0.81, RMSE = 0.64 m 2 /m 2 , rRMSE = 11.5 %). In Stage 2, LAI-TransNet is developed by fine-tuning the CNN-TL model on PlanetScope simulated data (PS-Sim), preprocessed via cross-domain mapping to align UAV and satellite spectral features. Real PlanetScope imagery is corrected for reflectance consistency with reference to UAV imagery spectral profiles. LAI-TransNet outperforms other deep learning models trained directly on PS-Sim (R 2 = 0.69 vs. 0.60–0.63), ensuring robust cross-scale consistency. In conclusion, by bridging UAV and satellite scales, LAI-TransNet enables large-scale soybean LAI monitoring, enhancing precision agriculture management through improved monitoring with the PlanetScope imagery.

Leaf area index (LAI)

Enhancing biomass flowability for entrained flow Gasification: The role of densification and torrefaction

Gasification presents a key strategy in addressing future energy demands while minimizing environmental impact. This has been recognized as a promising method to convert biomass to higher value products such as biofuels or hydrogen. Among gasification technologies, high-temperature and high-pressure reactors, particularly the R-GAS® system, emerge as an advanced option boasting superior conversion efficiency. However, akin to conventional high-temperature and high-pressure gasifiers, R-GAS® necessitates small particle sizes for optimal carbon conversion, a requirement yet to be fully explored for biomass. Hence, this study investigated the effectiveness of combined mechanical and thermal preprocessing techniques in modifying the physicochemical properties of biomass to suit gasification systems. Mechanical techniques including densification and pulverization, alongside thermal techniques such as torrefaction and steam explosion, were examined. The results demonstrate that torrefaction fosters producing of uniform granular material, enhancing flowability and reducing energy requirements for pulverization compared to steam explosion. Notably, torrefied corn stover exhibited lower internal friction angles and effective cohesion (40.09 ± 0.22° and 0.56 ± 0.01 kPa, respectively) compared to steam exploded corn stover (41.87 ± 0.65° and 0.83 ± 0.06 kPa, respectively), indicative of improved flowability. Additionally, pulverization of torrefied corn stover required approximately 16 % less energy than steam exploded corn stover and 91 % less energy than raw corn stover. Furthermore, the torrefaction-induced alterations in particle size, shape, and packing densities emphasize its potential to optimize flow and handling processes for gasification. These findings underline that densification followed by torrefaction effectively addresses biomass variability, leading to more efficient and sustainable energy conversion.

09 - BIOMASS FUELS

Optical image analysis of WSe 2 − thresholding for layer detection

The fast and reliable layer identification of two-dimensional transition metal dichalcogenide (TMD), such as WSe 2 , is essential to investigating their thickness-dependent electronic and optical properties. This article presents efficient optical image thresholding methodology designed to segment the mono, bi, and tri-layer regions of WSe 2 flakes mechanically exfoliated onto a SiO 2 /Si substrate. The optical images were first preprocessed to exclude the background effect and analyzed using the pixel medians and interquartile ranges for fundamental color channels—red, green, and blue (RGB). The analysis of red channel pixel intensities yielded three distinct ranges, serving as thresholds for layer segmentation: monolayer (111.0–118.0), bilayer (103.0–110.0), and tri-layer (93.0–103.0). Similarly, thresholds were established for each color channel, facilitating a comparative study of the segmentation performances. Further, the intersection-over-union ($IoU$) calculations revealed that the red and green channels demonstrated greater than 99 % and 90 % accuracy in differentiating each layer, respectively. This approach yields remarkable results without substantial data calibration that utilizes time-intensive heuristic techniques. Moreover, the proposed methodology offers the flexibility to compare performances across different color channels, expanding the applicability for other 2D material systems.

2D Materials

Adaptive continuity-preserving simplification of street networks

Street network data is widely used to study human-based activities and urban structure. Often, these data are geared towards transportation applications, which require highly granular, directed graphs that capture the complex relationships of potential traffic patterns. While this level of network detail is critical for certain fine-grained mobility models, it represents a hindrance for studies concerned with the morphology of the street network. For the latter case, street network simplification — the process of converting a highly granular input network into its most simple morphological form — is a necessary, but highly tedious preprocessing step, especially when conducted manually. In this manuscript, we develop and present a novel adaptive algorithm for simplifying street networks that is both fully automated and able to mimic results obtained through a manual simplification routine. The algorithm — available in the neatnet Python package — outperforms current state-of-the-art procedures when comparing those methods to manually, human-simplified data, while preserving network continuity.

Python

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Spatial and temporal characterization of municipal solid waste based on resource recovery pathways

This study presents a two-year, quarterly assessment of MSW across four source sectors (residential, schools, restaurants, and grocery stores) from sixteen sites across five U.S. states. MSW was manually sorted into 27 categories and aggregated into pathway fractions: high-moisture (HM) organics, low-moisture (LM) organics, recyclable (RC) materials, and residuals for disposal. Organics represented 89 % of the MSW stream. The largest fraction was HM organics consisting of food waste (31 %) and yard waste (3 %), with large coefficient of variations (CV), 79 and 278 %, respectively, reflecting high seasonal and site variability that varied significantly (p < 0.01) across sampling periods. The HM fraction showed properties favorable for anaerobic digestion, with moisture content ranging from 56 to 95 % and volatile solids ranges of 86-95 %. In contrast, the LM and RC fractions remained more stable (plastics CV = 41 %; paper CV = 53 %) with heating values up to 26.9 MJ/kg across sources, reflecting suitability for gasification. Microstructural analysis revealed less porosity in residential waste sampled at the landfill, which can influence preprocessing efficiency and microbial accessibility. Pathway informed allocations showed that 35 % of MSW is suitable for anaerobic digestion, 36 % for gasification, and 18 % for recycling, leaving 11 % requiring landfill disposal. These results provide quantitative evidence to determine feedstock allocation, waste-to-energy system design, and the development of data-driven sustainability and resource recovery strategies within a circular bioeconomy.

09 BIOMASS FUELS

The Effect of Air Separations on Fast Pyrolysis Products for Forest Residue Feedstocks

This study investigates the intricate relationship between biomass preprocessing and pyrolysis product yields, employing the air classification technique for the treatment of loblolly pine residues with varying moisture content. A comprehensive exploration of the physicochemical properties of air-classified loblolly pine informs a sophisticated pyrolysis simulation model. Given the complex and multifaceted nature of biomass pyrolysis, operating across diverse temporal and spatial scales, a pyrolysis kinetics-based CFD–DEM simulation method is employed to predict product yields. Results showed that the elevated moisture content amplifies particle adhesiveness, necessitating augmented air velocities for effective separation, thereby influencing the efficiency of the separation process. While carbon and hydrogen contents exhibit relative stability across diverse moisture contents and blower frequencies, the oxygen content undergoes noticeable changes. For example, the oxygen contents were measured as 29.2 and 38.6 wt% in the light fraction of 30% moisture content sample at blower frequencies of 10 and 20 Hz, respectively. An intriguing finding emerges from pyrolysis simulation, indicating that a lower blower frequency in air classification moderately enhances bio-oil yield and significantly improves its quality, particularly in terms of water content. For instance, the water content in the bio-oil was about 1.5% and 10% in the heavy and light fractions, respectively from 10% moisture sample under 15 Hz blower frequency. In summary, a detailed understanding and strategic manipulation of critical material attributes in biomass through efficient fractionation techniques are imperative for advancing fast pyrolysis as a sustainable avenue for renewable energy and chemical production.

09 BIOMASS FUELS