Search NASA⌕ Search

SEARCH · Search NASA

Results for “Task analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB↗

Strategy to Develop a Control Scheme for Core Thermal Performance Optimization

This report outlines a three-year research plan for optimizing reactor core thermal performance through use of a Digital Twin (DT) model and a set of neutron detectors to correspondingly adapt the reactor control strategy. It identifies reactor features and achievable neutron measurements that factor into this optimization task. It considers spatial effects that are important and how they can be managed by a real-time algorithm that controls reactivity actuators. The report presents a set of tasks, along with corresponding methods, for accomplishing this objective. The final task in this set is to validate the combined optimization procedure in the Purdue University’s PUR-1 reactor, for which we have an agreement of understanding. Additionally, we describe an alternative approach for overcoming some of the limitations inherent in the above approach, which includes the computational costs of high-fidelity simulations and achievable integration of neutron measurements into the DT. The proposed algorithms lower the technology readiness level (TRL) of the overall approach as some additional analysis and the development of a new neutron detector concept are necessary. In particular, a sensor measuring the gradient of the neutron flux is required, and the prototype is expected to be validated in the PUR-1 reactor. The potential benefits and the opportunity to significantly push forward the state-of-the-art make this approach worthy of further exploration in the upcoming years.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Metallurgical Analysis of the High Flux Isotope Reactor (HFIR) Carrier Lifting Bails (Rev.1)

The dissolution rates of the aluminum alloys in the High Flux Isotope Reactor (HFIR) element carriers and the Material Test Reactor (MTR) L-bundles in the H-Canyon facility have been identified as the possible cause of extended dissolutions that result in significant time and financial expenditures. A study, carried out by Savannah River National Laboratory (SRNL) to determine relationships between the dissolution rates and the metallurgical properties of the aluminum alloy materials of construction of the HFIR carriers and the L-bundles, considered the dissolution rates of aluminum alloy (AA) series 1100, 6061, and 6063. The study determined that the aluminum alloy compositions played a principal role in the dissolution rate of the carrier/bundle components. Higher dissolution rates were correlated with lower concentrations of the minor element additions in the alloys and with specific element concentrations. Aluminum alloys 1100 and 6063 were found to have similar dissolution rates that were approximately two orders of magnitude (100X) greater than those of AA6061. Based on the results of the dissolution behavior study, a Technical Assistance Request (TAR) was first issued to determine if the replacement of AA6061-T6 with AA6063-T6 is feasible for the HFIR carrier lifting bails. A Technical Task Request was then issued to consider AA6063-T5 as well as other alloys to improve possible supply chain issues. The metallurgical properties of the L-bundle (specifically the end caps) were not evaluated in this report because L-Bundle drawings already allow for the use of AA6063-T6 in all structural components. The HFIR carriers are composed of thin-walled components with significant surface areas that allow for relatively quick overall dissolution times. Conversely, the carrier lifting bails and the supporting constituents are composed of solid bars and thick plate regions with relatively small surface areas that experience longer overall dissolution times. While the MTR L-bundle design includes allowances for the materials of construction to be either AA6061-T6 or AA6063-T6, the HFIR carriers are specified to be constructed fully with AA6061-T6 alloy. This report analyzes the recommendations of the dissolution behavior study to replace the materials of construction of the HFIR carrier lifting bails. The analysis considers the operational requirements of the lifting bail and its supporting structures. To decrease dissolution times, the analysis considers direct replacement of the material as well as reductions in the thicknesses of the components to decrease the mass of the elements. Material reductions are considered on options for using either AA6061 and/or AA6063. The calculations are based on specifications from the American Society of Mechanical Engineer (ASME) and The Aluminum Association, Inc. design codes. The analysis finds that direct replacement of the lifting bail material of construction with AA6063-T6, and AA6063-T5 as well as reductions in the dimensions of the lifting bail components are acceptable. Note that this study considers the structural suitability of the alloys. It does not consider their dissolution rates in the dissolvers.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

PaRSEC: Scalability, flexibility, and hybrid architecture support for task-based applications in ECP

This paper highlights the most significant enhancements made to PaRSEC, a scalable task-based runtime system designed for hybrid machines, during the Exascale Computing Project (ECP). The enhancements focus on expanding the capabilities of PaRSEC to address the evolving landscape of parallel computing. Notable achievements include the integration of support for three major types of accelerators (NVIDIA, AMD, and Intel GPUs), the refinement and increased flexibility of the communication subsystem, and the introduction of new programming interfaces tailored for irregular applications. Additionally, the project resulted in the development of powerful debugging and performance analysis tools aimed at assisting users in understanding and optimizing their applications. We present a comprehensive demonstration of these advancements through a series of benchmarks and applications within ECP and beyond, thereby showcasing the enhanced capabilities of PaRSEC across the diverse architectures within the ECP, providing valuable insights into the runtime system’s adaptability and performance across varied computing environments.

Bouteiller, Aurelien↗

Advances in supervisory control strategies for a heat pump centric HVAC system − a comprehensive review on applications

There is an increase in research investigating the development and deployment of supervisory controllers that enable high efficiency electrically driven vapor compression heat pumps operating within grid interactive efficient buildings to provide demand side management. This paper reviews over sixty relevant case studies within this domain that focus on commercial off-the-shelf heat pumps whose primary task is providing space conditioning. The concept of a heat pump-centric heating, ventilation, and air-conditioning system is introduced, accompanied by a detailed overview of various kinds of electric heat pumps and building-level thermal energy storage configurations. Additionally, the different types of supervisory controller designs, including rule-based control and model predictive control, are discussed, along with the various methods for communication between a supervisory controller and a downstream heat pump’s local controller. A comparative analysis is conducted in order to categorize the reviewed case studies based on their system design, supervisory control algorithm, and validation methodology. This detailed analysis allows the review to establish current research trends, identify potential gaps, and suggest future directions for the development of this technology. Overall, the authors recommend that more future research be devoted to low-cost practical retrofits that allow for easy integration of active thermal energy storage within heat pump-centric heating, ventilation, and air-conditioning system systems that utilize direct expansion heat pumps. We also suggest more research into the development and deployment of supervisory controllers that can properly communicate with commercial off-the-shelf heat pump local controller available control inputs, e.g., zone temperature setpoint. Lastly, more rigorous experimental demonstrations of advanced supervisory control within real and or closed-loop, transient/ quasi-steady state environments are necessary to reduce industry wide skepticism of this technology.

Demand side management↗

Assessment of ROI for Workforce Development Efforts National & Homeland Security

The IDEAL Professional Engagement project at Idaho National Laboratory (INL) aims to enhance the laboratory's workforce diversity and inclusivity efforts, focusing on the U.S. National and Homeland Security mission areas. This project involves researching and evaluating opportunities for INL to engage in various professional events, particularly cyber conferences, to support recruitment and professional development. The project involved several key tasks: compiling comprehensive information on national laboratories and their mission statements, developing a deep understanding of INL’s role and efforts in national and homeland security, and identifying and engaging key stakeholders. Interviews were conducted with a set of targeted questions, and the findings were analyzed to identify common themes, insights, and actionable recommendations. Additionally, relevant upcoming cyber conferences were identified, various sponsorship levels and their associated benefits were evaluated, and the recruitment potential of these conferences was assessed. A detailed cost analysis was performed, including registration fees and travel expenses, and a cost-benefit analysis was conducted to evaluate the financial viability and potential return on investment (ROI) of conference participation and sponsorship. Based on the research and analysis, actionable recommendations for conference participation and sponsorship were formulated, ensuring alignment with INL’s mission and diversity goals. Preliminary results include a comprehensive list of relevant cyber conferences, a detailed cost analysis, and a set of actionable recommendations for future conference participation and sponsorship. The analysis highlights the financial requirements and geographical distribution of these conferences, providing valuable insights for INL's engagement strategies. This project underscores the importance of strategic engagement in professional events to attract and develop a diverse and skilled workforce, ultimately supporting INL’s mission areas in U.S. National and Homeland Security.

99 GENERAL AND MISCELLANEOUS↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES↗

NSTTF Voucher Program RPPR-1 (Final Report)

Sandia issued a Request for Proposals (RFP) to solicit proposals from industry, academia, research laboratories, government agencies, and individuals on the use of the National Solar Thermal Test Facility (NSTTF) to increase CSP technology market adoption across the United States. The voucher funds will be used to cover the cost of NSTTF test facilities usage and technical staff support for analysis, design and test planning and execution. Sandia will collect submitted proposals, coordinate their review through DOE SETO, and work in partnership or under contract with the applicants to complete the funded research. Through this program, participants will be supported in their use of the world class facilities and expertise available at the NSTTF at Sandia in Albuquerque, NM to accelerate the advancement of CST technologies toward meeting 2030 SETO goals for CSP. The goals of this semi-annual reporting period were to complete all administrative tasks and contracting, begin testing on three of the vouchers, and report on initial findings. The fourth voucher (University of Michigan) is predicated on the results of an ongoing heat exchanger test that is expected to conclude by the end of FY22.

14 SOLAR ENERGY↗

Characterization of Tank 9H Annulus Sample in Support of Residual Material Inventory Determinations

The Savannah River National Laboratory (SRNL) was requested by Savannah River Mission Completion (SRMC) to provide sample preparation and characterization of the Tank 9H annulus sample in support of Residual Material Inventory Determinations. One Tank 9H sample in three vials [HTF-9-25-13, HTF-9-25-14 and HTF-9-25-15], with each vial containing approximately 200 mL of the Tank 9H annulus salt solution, were delivered to the SRNL Shielded Cells for sample preparation and characterizations in February 2025. The density of the “as-received” solution contained in each of the three Tank 9H annulus sample vials were determined followed by a solid-liquid separation on each one using 0.45-micron Nalgene® nylon filter membranes. The resulting filtrates were combined to form the Tank 9H annulus sample with a total volume of about 600 mL. The combined wet solid fractions, about a total of 4.8 grams of salt material, remaining on the filter membranes were air-dried in the Shielded Cells for 72 hours. The total weight of the air-dried solids was 2.1 grams. These air-dried solids were washed with deionized water (DI water) at a phase ratio of 60 mL DI water/gram of solids to recover insoluble solids, if any. No visible or measurable quantity of insoluble solids were recovered after DI water washing of the air-dried solids because the air-dried solids completely dissolved in the DI water. The solid fraction-wash water was not combined with the 600 mL of the filtrate solution, and the resulting solution was not screened or analyzed for radionuclides. Aliquot sample volumes of the undiluted Tank 9H annulus sample were sent to the SRNL analytical services groups for radionuclides, elementals, anions and total mercury analysis by various methods including radiochemical separations/counting methods, inductively coupled plasma-atomic emission spectroscopy (ICP-AES), and Inductively Coupled Plasma Mass Spectroscopy (ICP-MS) and special preparations. All sample analyses were performed in triplicate. This report presents the analytical characterization results for the Tank 9H annulus sample. The results are also reported where analytical methods yielded additional analytes, other than those requested by SRMC. In the characterization of the Tank 9H annulus sample, the detection limits for all the analytes, as specified in the Technical Task Request (TTR) and Task Technical and Quality Assurance Plan (TTQAP), were met.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Decentralized Microgrid Protection Through Relative Fault Direction Classification: Preprint

Protection in inverter-based resources (IBRs) dominated microgrids generally face significant challenges due to the low fault current and inconsistent fault behaviors from IBRs. Recently, machine learning-based approaches have attracted considerable attention to address these challenges. This paper introduces a novel decentralized protection strategy for microgrids. The proposed method decomposes the protection challenge into several distributed learning tasks, enabling individual relays to autonomously determine the direction of faults using a binary classification framework based on support vector machine (SVM) algorithms. Following the distributed fault direction estimation, classifier outcomes are shared among neighboring relays, facilitating a local decision-making process to ascertain the presence of faults within the neighborhood. Finally, a tripping signal is generated based on the classifier results of each relay to operate the circuit breaker. To test and validate this approach, a 100% renewable microgrid model is simulated in MATLAB/Simulink. In the numerical analysis, the application of SVM classifiers in our approach yields impressive results: an average relay classification accuracy of 98%, and a 96% accuracy in circuit breaker control. These findings highlight the potential of machine-learning-based approaches in enhancing the efficiency and reliability of microgrid protection systems.

decentralized algorithm↗

Samoa Updater: An Application of the Levenberg-Marquardt Method to Update DELFIC Predictions Using Field Measurements

The US Department of Energy (DOE) Forensics Operations (DFO) is a member of the Ground Collections Task Force (GCTF), which is responsible for sample collection of radiological debris for attribution should a nuclear detonation ever occur in the United States. The DFO runs the Defense Land Fallout Interpretive Code (DELFIC) Fallout Planning Tool to predict the deposition of fallout from a nuclear detonation. This prediction is refined using the DELFIC Updater tool, which takes ground measurements and adjusts DELFIC inputs to minimize the difference between prediction and observation, yielding improved predictions of fallout in locations both measured and not yet measured. Samoa, a framework for uncertainty analysis and optimization, is used to improve DELFIC predictive fallout modeling. This new capability using Samoa, dubbed “Samoa Updater,” is compared with the current DELFIC Updater, a brute-force sampling approach. Samoa Updater uses the Levenberg– Marquardt (LM) method, a gradient-based nonlinear least squares approach that uses the functional shape of the input space to increase optimization speed. In simulated test cases Samoa Updater yields faster and more accurate solutions than the current Updater.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Topological and Dynamical Representations for Radio Frequency Signal Classification

Radio Frequency (RF) signals are found throughout our world, carrying over-the-air information for both digital and analog uses with applications ranging from WiFi to the radio. One area of focus in RF signal analysis is determining the modulation schemes employed in these signals which is crucial in many RF signal processing domains from secure communication to spectrum monitoring. This work investigates the accuracy and noise robustness of novel Topological Data Analysis (TDA) and dynamic representation based approaches paired with a small convolution neural network for RF signal modulation classification with a comparison to state-of-the-art deep neural network approaches. We show that using TDA tools, like Vietoris-Rips and lower star filtrations, and the Takens' embedding in conjunction with a standard shallow neural network we can capture the intrinsic dynamical, geometric, and topological features of the underlying signal's manifold, offering informative representations of the RF signals. Our approach is effective in handling the modulation classification task and is notably noise robust, outperforming the commonly used deep neural network approaches in mode classification. Moreover, our fusion of dynamical and topological information is able to attain similar performance to deep neural network architectures with significantly smaller training datasets.

Myers, Audun D.↗

Deploying Adversarial Attacks in Super-Resolution Models

Reliable super-resolution methods are crucial for applications like remote sensing, grid resilience and disaster impact analysis, and standoff biometrics. These methods infuse additional high-frequency information into reconstructions, allowing for better contextualization and image intelligence. However, super-resolution models can also introduce hallucinations or other unseen vulnerabilities that could be exploited by an adversary. This is further compounded by the prominence of deep learning in these models, as models are often blindly applied on out-of-distribution images. In this work, we implement adversarial attacks in common open-source super-resolution models and examine their impact on reconstructions and downstream classification tasks. We find that an adversarially trained super-resolution model can produce high-quality reconstructions that degrade downstream classifications. Moreover, these attacks do not require access to low-resolution imagery or class labels at inference time. These results demonstrate the vulnerability of super-resolution methods to malicious actors and motivates the development of a detector for super-resolution adversarial attacks. Further exploration of adversarial attacks in this domain is required to ensure trustworthiness and robustness of super-resolution models for national security applications.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗