Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data segmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Novel Data Segmentation Method for Data-driven Phase Identification

This paper presents a smart meter phase identification algorithm for two cases: meter-phase-label-known and meter-phase-label-unknown. To improve the identification accuracy, a data segmentation method is proposed to exclude data segments that are collected when the voltage correlation between smart meters on the same phase is weakened. Then, using the selected data segments, a hierarchical clustering method is used to calculate the correlation distances and cluster the smart meters. If the phase labels are unknown, a Connected-Triple-based Similarity (CTS) method is adapted to further improve the phase identification accuracy of the ensemble clustering method. The methods are developed and tested on both synthetic and real feeder data sets. Here, simulation results show that the proposed phase identification algorithm outperforms the state-of-the-art methods in both accuracy and robustness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Study of reference burnup steps optimization in fuel segment data file generation for NEXUS/ANC9 code system

For any two-step core design code system, the cross-section files are the primary factor to determine the accuracy of the system prediction. For a once-through cross-section system to cover all potential applicable conditions, the system may need to perform tens of thousands state points of lattice calculations. These calculations take a significant amount of CPU times and computing power, especially when a fine or ultra-fine energy group library is used. As computer power increases, so does the calculational complexity. Therefore, it is important to make the lattice calculations effective and efficient. Based on the Westinghouse core design code system NEXUS/ANC9, this work focuses on the reference burnup steps used by the lattice code when performing the calculations. Following the fundamental cross-section methodology and analyzing the contribution of each individual terms, this study provides an applicable solution of the optimized reference burnup steps. The test results show that, with the existing cross-section methodology, it is possible to significantly reduce the lattice calculation cases and cross-section file generation time without sacrificing the accuracy of system prediction. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Load Profile Inpainting for Missing Load Data Restoration and Baseline Estimation

This paper introduces a Generative Adversarial Nets (GAN) based, Load Profile Inpainting Network (Load-PIN) for restoring missing load data segments and estimating the baseline for a demand response event. The inputs are time series load data before and after the inpainting period together with explanatory variables (e.g., weather data). Here, we propose a Generator structure consisting of a coarse network and a fine-tuning network. The coarse network provides an initial estimation of the data segment in the inpainting period. The fine-tuning network consists of self-attention blocks and gated convolution layers for adjusting the initial estimations. Loss functions are specially designed for the fine-tuning and the discriminator networks to enhance both the point-to-point accuracy and realisticness of the results. We test the Load-PIN on three real-world data sets for two applications: patching missing data and deriving baselines of conservation voltage reduction (CVR) events. We benchmark the performance of Load-PIN with five existing deep-learning methods. Our simulation results show that, compared with the state-of-the-art methods, Load-PIN can handle varying-length missing data events and achieve 15-30% accuracy improvement.

14 SOLAR ENERGY↗

Hyperspectral segmentation of plants in fabricated ecosystems

Hyperspectral imaging provides a powerful tool for analyzing above-ground plant characteristics in fabricated ecosystems, offering rich spectral information across diverse wavelengths. This study presents an efficient workflow for hyperspectral data segmentation and subsequent data analytics, minimizing the need for user annotation through the use of ensembles of sparse mixed scale convolution neural networks. The segmentation process leverages the diversity of ensembles to achieve high accuracy with minimal labeled data, reducing labor-intensive annotation efforts. To further enhance robustness, we incorporate image alignment techniques to address spatial variability in the dataset. Downstream analysis focuses on using the segmented data for processing spectral data, enabling monitoring of plant health. This approach provides a scalable solution for spectral segmentation, and facilitates actionable insights into plant conditions in complex, controlled environments. Our results demonstrate the utility of combining advanced machine learning techniques with hyperspectral analytics for high-throughput plant monitoring.

Zwart, Petrus H.↗

Data for The utility of transfer learning to improve the performance of deep learning in axon segmentation

The utility of transfer learning to improve the performance of deep learning in axon segmentation Data Data: All the input and labeled volumes tf-logs: Tensorflow logs, view with command "tensorboard --logdir [name of folder]" Model Weights: model_weights: the argument list under variable combo indicate 1) no oversampling, 2) no rotation, 3) no learn scheduler, and 4) flipping on all three dimensions, and the additional values indicate 5) elastic deformation percentage, 6) rotate deformation percentage, 7) layer setting , 8) learning rate, and 9) training/validation/test data division suffix (leave '' if not using suffix). Results: Output from inference segment_total_results_validation_final: All validation results and calculations segment_total_results: All test results and calculations Authors The modified code was created for a paper by: Marjolein Oostrom, Michael A. Muniak, Rogene Eichler West, Sarah Akers, Paritosh Pande, Moses Obiri, Wei Wang, Kasey Bowyer, Zhuhao Wu, Lisa Bramer, Tianyi Mao, Bobbie Jo Webb-Robertson The work is adapted from Github TrailMap, which was created by Albert Pun and Drew Friedmann Acknowledgments MO, RMEW, SA, MO, LB, BJWR were supported by the Laboratory Directed Research and Development at Pacific Northwest National Laboratory (PNNL), a Department of Energy facility operated by Battelle under contract DE-AC05-76RLO01830. WW, KB, and ZW were supported in part by a NIH/BRAIN Initiative Grant RF1MH128969. MAM and TM were supported by two NIH/BRAIN Initiative Grants R01NS104944, RF1MH120119 and NIH R01NS081071. This research is affiliated with the Pacific northwest bioMedical Innovation Co-laboratory (PMedIC) collaboration between OHSU and PNNL.

Oostrom, Marjolein T↗

A deep learning approach for semantic segmentation of unbalanced data in electron tomography of catalytic materials

In computed TEM tomography, image segmentation represents one of the most basic tasks with implications not only for 3D volume visualization, but more importantly for quantitative 3D analysis. In case of large and complex 3D data sets, segmentation can be an extremely difficult and laborious task, and thus has been one of the biggest hurdles for comprehensive 3D analysis. Heterogeneous catalysts have complex surface and bulk structures, and often sparse distribution of catalytic particles with relatively poor intrinsic contrast, which possess a unique challenge for image segmentation, including the current state-of-the-art deep learning methods. To tackle this problem, we apply a deep learning-based approach for the multi-class semantic segmentation of a γ-Alumina/Pt catalytic material in a class imbalance situation. Specifically, we used the weighted focal loss as a loss function and attached it to the U-Net’s fully convolutional network architecture. We assessed the accuracy of our results using Dice similarity coefficient (DSC), recall, precision, and Hausdorff distance (HD) metrics on the overlap between the ground-truth and predicted segmentations. Our adopted U-Net model with the weighted focal loss function achieved an average DSC score of 0.96 ± 0.003 in the γ-Alumina support material and 0.84 ± 0.03 in the Pt NPs segmentation tasks. We report an average boundary-overlap error of less than 2 nm at the 90th percentile of HD for γ-Alumina and Pt NPs segmentations. The complex surface morphology of γ-Alumina and its relation to the Pt NPs were visualized in 3D by the deep learning-assisted automatic segmentation of a large data set of high-angle annular dark-field (HAADF) scanning transmission electron microscopy (STEM) tomography reconstructions.

36 MATERIALS SCIENCE↗

Data for Clumping Index Estimation With 30°-tilted Cameras in Row Crops: Evaluation of Methods and Segment Size Effects

The clumping index (CI) quantifies the spatial distribution of foliage elements and is essential for accurately estimating the plant area index (PAI), canopy radiative transfer, and photosynthesis. Traditionally, the finite-length averaging method (LX), the gap size distribution method (CC), and a combined approach of CC and LX (CLX) have been applied to instruments like TRAC and digital hemispherical photography to estimate CI. However, a comprehensive evaluation of these methods in row crops remains limited, especially regarding the influence of segment size on CI. Meanwhile, digital cameras offer a cost-effective and user-friendly solution for canopy measurements in row crops, yet their application in this context remains underexplored. In this study, we employed a new approach using a 30°-tilted digital camera to estimate CI in corn and soybean fields, applying the LX, CC, and CLX methods. We systematically assessed the performance of these three methods by combining field measurements in real-world fields with simulations using the LESS 3D radiative transfer model. Our results showed that CLX applied to the whole image and 45° segment offered accurate estimation of CI (bias within ±0.1, RMSE < 0.2) and PAI (bias within ±0.4, RMSE < 1) in real-world fields and LESS simulations. The accuracy of the LX method was highly sensitive to segment size, with the best performance observed at the 15° segment (PAI bias within ±0.4). In contrast, the CC method remained stable across different segment sizes, and its performance was generally comparable to that of LX, except at the 15° segment. Across view zenith angles, CI derived from CC generally showed a continuous increase, while those from LX and CLX followed a rising trend at small zenith angles but began to decline at 68°, likely due to an increasing proportion of no-gap segments. Seasonally, LX tended to show decreasing CI during early growth stages but increased as the canopy matured, whereas CC and CLX showed gradually increasing CI before plateauing at peak PAI. The 30°-tilted camera effectively captured CI variations across different angles and growth stages, making it a practical and robust instrument for row crop canopy structure analysis. Applying these CI methods to digital cameras offers a low-cost and accessible CI estimation alternative, improving canopy structure monitoring accuracy in row crops.

Modeling↗

Sister Rod Destructive Examinations (FY23) Appendix B: Segmentation, Defueling, Metallographic Data and Total Cladding Hydrogen

As a part of the DOE-NE High Burnup Spent Fuel Data Project, Oak Ridge National Laboratory (ORNL) is performing destructive examinations (DEs) of high burnup (HBU) (>45 GWd/MTU) spent nuclear fuel (SNF) rods from the North Anna Nuclear Power Station operated by Dominion Energy. The SNF rods, called sister rods or sibling rods are all HBU and include four different kinds of fuel rod cladding: standard Zircaloy-4 (Zirc-4), low-tin (LT) Zirc-4, ZIRLO ® , and M5 ® . The DEs are being conducted to obtain a baseline of the HBU rod’s condition before dry storage and are focused on understanding overall SNF rod strength and durability. Both composite fuel and defueled cladding will be tested to derive material properties. Although the data generated can be used for multiple purposes, one primary goal for obtaining the post-irradiation examination data and associated measured mechanical properties is to support SNF dry storage licensing and relicensing activities by (1) addressing identified knowledge gaps and (2) enhancing the technical basis for post-storage transportation, handling, and subsequent disposition of the SNF. This report documents the status of the ORNL Phase 1 DE activities related to: Rough segmentation (RS), Defueling (DEF), DE.02 optical microscopy (MET), and DE.03, cladding total hydrogen measurements. It is a cumulative update to the FY22 status report.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Videos, photos, and AI-derived grain size data associated with “High-throughput AI Video Surveys Enable Reproducible Multiscale Sediment Size Mapping, with Implications for Hydrobiogeochemical Parameterization”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “High-throughput AI Video Surveys Enable Reproducible Multiscale Sediment Size Mapping, with Implications for Hydrobiogeochemical Parameterization” under review. This data package includes five data types: 1) raw photos and videos from drone survey and walking smartphone surveys; 2) images derived from raw videos; 3) manual labeling of reference scales; 4) metadata for all images and photo resolution derived from artificial intelligence (AI) models or manual labels, 5) grain size data obtained from AI models for all photos, 6) metadata and grain size data after quality control, 7) summaries of sample efficiency for all data, and 8) computational fluid dynamics (CFD) data used to support hydro-biogeochemical (HBGC) parameter estimation. Such data is used to 1) demonstrate significant improvements in accuracy, efficiency, and quality control for grain size data collection with the help of AI models, 2) study the spatial heterogeneity of grain size and observation reproducibility based on tens of thousands of data points generated by the AI models, and 3) evaluate the impacts of grain size heterogeneity on key HBGC parameters across sediment-to-reach and hourly-to-yearly scales. In particular, the data package contains 116 folders and 179696 files. The files include 41 videos in .mov format, 64047 photos in .jpg format, 13541 video-derived photos in .png format, 12747 segmentation mask data in .tif format, 12747 segmentation data in .json format, 24771 .csv files that with metadata and grain size for each individual photo as well as water depth and velocity data from CFD and observation, 51791 .txt files of raw AI predicted labels, and 11 flight record data in .srt format. The summary for all metadata and grain size statistics information is included in “Scales_V3_NG.csv” and “Statistics_V3_NG.csv”. The summary for data that pass data quality control (QC) level 0-2 is included in “QCStatistics_V3_NG.csv”. The QC level 0 represents photos whose photo resolution is positive, excluding photos that miss reference scale. The QC level 1 means reference scale circularity uncertainty is less than 5% for smartphone images while representing photo resolution is larger than 0.44 mm/pixel for drone images. The QC level 2 means excluding photos whose grain number is less than 100, a minimum number of grains recommended by classic literature. The summary for each video’s name, length, frame rates, survey area, grain number, survey efficiency, etc. can be found in “QCSummary_V3_NG.csv”. The summary for site name, GPS coordinates, and number of images at each site can be found in “SitesSummary_V3_*.csv” files. Overall computational efficiency summary is reported in Table 4 of accompanying manuscript. Additionally, the nitrate concentration data used in this work was downloaded from an existing dataset published on ESS-DIVE (Boat-Dragged Sensor Hanford Reach.csv; Conner A. et al., 2020). We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Port of Benton, and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the data were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate data collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Sister Rod Destructive Examinations (FY22) Appendix B: Segmentation, Defueling, Metallographic Data and Total Cladding Hydrogen

This report documents work performed under the Spent Fuel and Waste Disposition’s Spent Fuel and Waste Science and Technology program for the US Department of Energy (DOE) Office of Nuclear Energy (NE). This work was performed to fulfill Level 2 Milestone M2SF-23OR010201024, “FY22 Report on ORNL Sibling Rod Testing Results,” within work package SF-23OR01020102 and is an update to the work reported in M2SF-22OR010201047, M2SF-21OR010201032, M2SF-19ORO010201026, and M2SF 19OR010201028. As a part of the DOE-NE High Burnup Spent Fuel Data Project, Oak Ridge National Laboratory (ORNL) is performing destructive examinations (DEs) of high burnup (HBU) (>45 GWd/MTU) spent nuclear fuel (SNF) rods from the North Anna Nuclear Power Station operated by Dominion Energy. The SNF rods, called sister rods or sibling rods are all HBU and include four different kinds of fuel rod cladding: standard Zircaloy-4 (Zirc-4), low-tin (LT) Zirc-4, ZIRLO ® , and M5 ® . Appendix B provides detailed information regarding defueling activities, metallographic imaging and measurements, and total cladding hydrogen measurements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Quantifying the robustness of deep multispectral segmentation models against natural perturbations and data poisoning

In overhead image segmentation tasks, including additional spectral bands beyond the traditional RGB channels can improve model performance. However, it is still unclear how incorporating this additional data impacts model robustness to adversarial attacks and natural perturbations. For adversarial robustness, the additional in-formation could improve the model’s ability to distinguish malicious inputs, or simply provide new attack avenues and vulnerabilities. For natural perturbations, the additional information could better inform model decisions and weaken perturbation effects or have no significant influence at all. In this work, we seek to characterize the performance and robustness of a multispectral (RGB and near infrared) image segmentation model subjected to adversarial attacks and natural perturbations. While existing adversarial and natural robustness research has focused primarily on digital perturbations, we prioritize on creating realistic perturbations designed with physical world conditions in mind. For adversarial robustness, we focus on data poisoning attacks whereas for natural robustness, we focus on extending ImageNet-C common corruptions for fog and snow that coherently and self-consistently perturbs the input data. Overall, we find both RGB and multispectral models are vulnerable to data poisoning attacks regardless of input or fusion architectures and that while physically-realizable natural perturbations still degrade model performance, the impact differs based on fusion architecture and input data.

Deep learning, multispectral images, multimodal fu↗

Global centroid moment tensor solutions in a heterogeneous earth: the CMT3D catalogue

SUMMARY For over 40 yr, the global centroid-moment tensor (GCMT) project has determined location and source parameters for globally recorded earthquakes larger than magnitude 5.0. The GCMT database remains a trusted staple for the geophysical community. Its point-source moment-tensor solutions are the result of inversions that model long-period observed seismic waveforms via normal-mode summation for a 1-D reference earth model, augmented by path corrections to capture 3-D variations in surface wave phase speeds, and to account for crustal structure. While this methodology remains essentially unchanged for the ongoing GCMT catalogue, source inversions based on waveform modelling in low-resolution 3-D earth models have revealed small but persistent biases in the standard modelling approach. Keeping pace with the increased capacity and demands of global tomography requires a revised catalogue of centroid-moment tensors (CMT), automatically and reproducibly computed using Green's functions from a state-of-the-art 3-D earth model. In this paper, we modify the current procedure for the full-waveform inversion of seismic traces for the six moment-tensor parameters, centroid latitude, longitude, depth and centroid time of global earthquakes. We take the GCMT solutions as a point of departure but update them to account for the effects of a heterogeneous earth, using the global 3-D wave speed model GLAD-M25. We generate synthetic seismograms from Green's functions computed by the spectral-element method in the 3-D model, select observed seismic data and remove their instrument response, process synthetic and observed data, select segments of observed and synthetic data based on similarity, and invert for new model parameters of the earthquake’s centroid location, time and moment tensor. The events in our new, preliminary database containing 9382 global event solutions, called CMT3D for ‘3-D centroid-moment tensors’, are on average 4 km shallower, about 1 s earlier, about 5 per cent larger in scalar moment, and more double-couple in nature than in the GCMT catalogue. We discuss in detail the geographical and statistical distributions of the updated solutions, and place them in the context of earlier work. We plan to disseminate our CMT3D solutions via the online ShakeMovie platform.

58 GEOSCIENCES↗

Automatic Segmentation of Building Envelope Point Cloud Data Using Machine Learning

About 50% of buildings in the US were constructed before energy codes were introduced. Modular overclad panel retrofits, in which a new envelope is constructed over the existing building, are a promising solution given that it minimizes occupant disruption and shortens construction time at the jobsite. Current state-of-the-art retrofit panel layout and dimensioning consists of three steps: 1) 3D point cloud data generation of the building envelope using commonly available surveying equipment, 2) manual segmentation of 3D point cloud data by a trained professional to identify and dimension window openings, door openings, and other architectural features, and 3) modular panel layout optimization and dimensioning by an architect or engineer. Among these steps, the second one remains the most difficult and costly because it is very labor-intensive. We propose a methodology to automatically label 3D point cloud data to reduce the time and expense spent in manual segmentation. Machine learning methods were employed to classify the point cloud data into distinct groups, each of which corresponds to different features of the building envelope. After classification, a segmentation algorithm was developed to perform boundary detection and separate the components of the façade. Finally, the algorithm returns the relative positions and dimensions of the features in the building envelope. The measurements obtained with the proposed automated method were compared against the actual dimensions to determine the overall algorithm accuracy. The proposed algorithm can then be used to reduce manual efforts for 3D point cloud labeling before modular panel layout optimization is performed.

Maldonado Puente, Bryan↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Improved Data Interpretation through Identification of Time Series Periodicity Changes

Analysis and interpretation of time series data is easiest when the data values occur at uniform intervals in time, but actual data may have differing data sampling frequencies, such as monthly and daily readings. Applying data analysis techniques, such as smoothing, to such a data set may not give a representative result between time segments. The ability to automatically distinguish time segments of differing data frequency would provide a means for applying data analysis independently to each segment, though a suitable blending at segment boundaries would be required. A method for detecting frequency changes was developed and applied to Gaussian and median smoothing of hydraulic head data from groundwater wells at the U.S. Department of Energy Hanford Site in southeastern Washington state. The process identifies time segments of high-frequency (daily) or low-frequency (greater than daily) data using adjusted-bandwidth Gaussian kernel density estimation and a threshold value, which are further refined to address small blocks of low-frequency data within larger blocks of high-frequency data. User-selectable levels of smoothing are then applied independently to the time segments prior to combining the segment results for a single smoothed data set. This time segment identification approach provides effective low- and high-frequency data separation, which provides a method to apply data analysis independently to each time segment.

97 MATHEMATICS AND COMPUTING↗

Statistical Performance of Forced Oscillation Detectors in the Presence of Missing Measurements

In bulk power systems, measurement-based monitoring for large oscillations can help maintain system reliability. One of the challenges encountered in a recent field demonstration was the unavailability of measurements due to underlying measurement quality or communication problems. During the demonstration, the oscillation detector ignored a measurement location if even 10 seconds of data was missing. To extend the detector's ability to operate in these conditions, this paper evaluates the impact of three methods for addressing missing data. The strengths and weaknesses of each approach are evaluated using theoretical expressions for the probability of detection along with results from simulated data and publicly available field measurements. Based on these results, a suitable approach is identified that can extend the oscillation detector's performance when large segments of data are missing.

Follum, James D.↗