Search NASASearch

SEARCH · Search NASA

Results for “task work”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Large Load Integration - Task List and Overview

Large Load Integration Tasks: Task 1 – Workshops Support stakeholder engagement across industry to promote collaboration and identify solutions to challenges that will guide other work Task 2 – Ancillary Services Characterize different types of large loads to assess under what conditions they may be utilized to provide grid stability services Task 3 – Communications Explore the cybersecurity and communications infrastructure required to enable large loads to interface with grid operations to provide ancillary services Task 4 – Nuclear Integration Explore risks and methods for supporting large load energy needs with SMRs and incorporating them into the wider power system Task 5 – Decision Support and TA Provide support to stakeholders through the creation of planning tools and direct technical assistance.

24 - POWER TRANSMISSION AND DISTRIBUTION

Optimization and Evaluation of Energy Savings for Connected and Autonomous Off-Road Vehicles

Off-road vehicles, such as wheel loaders, excavators, and harvesters, are extensively utilized across a wide range of industries, including construction, agriculture, and mining. These machines have become indispensable in supporting the day-to-day operational needs of a nation, playing a critical role in various sectors' infrastructure and productivity. However, despite their utility, off-road vehicles are significant consumers of fossil fuels, resulting in substantial emissions that contribute to environmental degradation. This highlights the pressing need for research and technological advancements aimed at improving their energy efficiency and reducing their carbon footprint. There are, however, two primary challenges that must be addressed to achieve these goals. First, off-road vehicles typically perform both driving and working tasks simultaneously, which introduces a high level of complexity into their overall dynamic systems. Analysis the interactions between these functions is challenging. Second, research into off-road vehicles is inherently interdisciplinary, demanding expertise across several domains such as fluid power systems, vehicle dynamics, control theory, optimization techniques, and real-world implementation. Recognizing these challenges, we proposed the project titled "Optimization and Evaluation of Energy Savings for Connected and Autonomous Off-Road Vehicles" as a comprehensive solution to enhance fuel efficiency while simultaneously improving productivity. This project specifically focuses on autonomous off-road vehicles, with particular attention to wheel loaders, and seeks to develop novel methods to optimize energy consumption without sacrificing operational performance. The project integrates real-time control algorithms, vehicle dynamics modeling, and co-optimization of powertrain system and vehicle system to achieve these goals. Our optimization strategy dynamically co-optimizes critical parameters at both the powertrain and vehicle levels, including vehicle speed, working tool movements, powertrain dynamics, and engine operations in real-time. To streamline this optimization process, we developed a vehicle model that captures the key dynamics while significantly enhancing computational efficiency. This allows the system to intelligently minimize fuel consumption, all while maintaining or even improving productivity through real-time calculations during various off-road operations. To validate the effectiveness of this energy optimization method, we introduced a state-of-the-art Hardware-in-the-Loop (HIL) testbed. This reconfigurable testbed seamlessly integrates the actual engine with virtual models of the wheel loader's subsystems, allowing for accurate emulation of real-world operational loads and environments. By simulating these conditions, the HIL testbed enables us to evaluate the wheel loader’s performance under diverse working scenarios, ensuring the developed solution is applicable in real-world operations. This testbed proved to be instrumental in validating the optimization algorithms and demonstrating the system's practical effectiveness. During the evaluation and testing phase, we employed the HIL testbed to rigorously assess the energy savings and productivity improvements generated by the optimized system. The results were highly encouraging, revealing that the automated wheel loader achieved over 30% fuel savings compared to traditional, human-operated cycles, with comparable or even enhanced levels of productivity. The insights gained from this HIL-based testing provided critical validation of our approach and highlighted the potential for deploying these optimized autonomous technologies in real-world off-road vehicles.

33 ADVANCED PROPULSION SYSTEMS

Summary of SRNL Support Activities to the DOE-ORP Enhanced Waste Glass Program for Fiscal Year 2024

In fiscal year 2024 (FY24) Savannah River National Laboratory (SRNL) continued tasked work for the Office of River Protection (ORP) to expand glass compositional regions accessible for low-activity waste (LAW) and high-activity waste (HLW) vitrification processing. Experimental work continued in four primary technical areas focused on processing and performance of glasses relevant to the Hanford missions. The data and results from this work will be used to expand and validate the glass models being developed at Pacific Northwest National Laboratory (PNNL) for waste processing and acceptance. This report summarizes the activities and deliverables associated with work performed in FY24.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Radiation Accidents and Malicious Events – Scenarios and Scope of the Work of ICRP Task Group 120

The International Commission on Radiological Protection (ICRP) Task Group 120 (TG120) is developing ICRP recommendations for radiological protection for a wide range of radiation accidents and malicious events, complementing those given in ICRP Publication 146 (2020) for large nuclear accidents. The scope includes accidents involving criticalities, operating faults, and fires and explosions in nuclear facilities, inadvertent damage to sealed radiation sources, as well as malicious events, such as sabotage of nuclear facilities or materials, use of radiological dispersal devices, the contamination of food and drinking water supplies, and the deployment of nuclear weapons. A template has been designed to collate relevant information on a wide range of case studies and hypothetical malicious scenarios to ensure that the recommendations developed are broadly applicable and comprehensive. For all scenarios, a graded approach to protection is being taken, accepting that specific guidance may be required for some distinctive aspects, for example, protection during times of armed conflict. This paper provides an overview of the scenarios and scope of the work of TG120, including some of the radiological and non-radiological impacts of radiation emergencies, along the response and recovery timeline.

ICRP

Materials Characterization, Prediction, and Control Project: Characterization of 316L Stainless Steel after Solid Phase Processing using Ultrasonic NDE Method

The Pacific Northwest National Laboratory undertook the Materials Characterization, Prediction, and Control Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems produced via advanced manufacturing methods, such as solid phase processing, for use in national security and advanced energy applications (Smith 2021). A motivation of the Materials Characterization, Prediction, and Control Project was to demonstrate ultrasonic testing as a nondestructive evaluation method to complement traditional destructive methods for characterizing material microstructure with emphasis on grain size determination using a method that may have future applications for real-time inline process monitoring. The objective of the work described in this report is to establish the process and an analysis method for measuring grain sizes of polycrystalline metals with ultrafine grains using ultrasonic shear wave backscattering, building on prior studies on coarser-grained material. The work involves five tasks: Measured ultrasonic backscattering experimentally for a series of 316L stainless steel specimens with various grain sizes made by friction stir processing. Calculated ultrasonic backscattering coefficients from experimental data based on a physical measurement model. Measured ground truth grain sizes of the specimens from electron backscatter diffraction grain boundary images using a generalization of the ASTM E112 (ASTM 2021) intercept method. Built a curve of ultrasonic backscattering coefficients versus the ground truth intercept-based grain sizes to determine the correlation between mean grain sizes and ultrasonic measurements. Demonstrated the ability of using the correlation curve to deduce grain sizes with measured ultrasonic backscattering coefficients for a few 316L stainless steel specimens whose grain sizes were unknown beforehand but were targeted to be an extrapolation to larger grain sizes than used to formulate the correlation curves. Experimental procedures and computational algorithms are developed and validated for these tasks. This work establishes an ultrasonic technique for characterizing material microstructure with ultrafine grains that are often resulted by solid-phase processing. The technique is nondestructive, and it has the potential to be used for real time inline process monitoring. This work successfully demonstrates the viability of an ultrasonic nondestructive evaluation method for microstructural characterization of material having ultrafine grain structure (as small as 1?mm) and produced by an advanced manufacturing method. This includes a demonstration of the method to extrapolate to other conditions. While not demonstrated here, the method is expected to be viable for in-line, or near-inline, process monitoring in advanced manufacturing applications with suitable consideration for access of instrumentation to the material being manufactured.

316 L Stainless Steel

District Geothermal Heating + Cooling Deployment in a CT Environmental Justice Community

The report marks the team’s completion of all required tasks and milestones. Work completed for Task 1 (Technical and Economic Feasibility Assessment & Procurement Drafting) included development of analysis and design model; completion of technical, economic, and environmental assessments; and technical outreach and coalition design. Components for Task 2 (Outreach & Community Engagement) involved broad outreach and community-engagement efforts (including stakeholder meetings and a webinar as well as development of a formal engagement plan) and development of a web page and a case study. For Task 3 (Workforce Transition, Development, & Training Plan), the team undertook a formal statewide geothermal workforce needs assessment, developed corresponding recommendations for both the state as a whole and the Wallingford project, and held several workshops. For Task 4 (Project Management & Data Sharing), the team drafted a data-sharing plan.

15 GEOTHERMAL ENERGY

MRCI Task 3: Facilitating Data Collection, Sharing, and Analysis Final Technical Summary Report

The Midwest Regional Carbon Initiative (MRCI) Task 3.0 was defined to facilitate development of carbon capture, utilization, and storage (CCUS) in the region by collection and sharing of existing and new technical data from CCUS projects and research. The task also included support for further analysis and assessment of tools by the project team and by researchers working on programs such as National Risk Assessment Partnership (NRAP), machine learning (ML) techniques, and assessment and improvement of CCUS site assessment, operations, and monitoring aspects. Work under Task 3.0 addressed key issues related to CCUS deployment and provided foundational research and datasets to help establish CCUS projects in the MRCI. Report Authors and Principal Technical Contributors: Joel Sminchak, Laura Keister, Mackenzie Scharenberg, Priya Ravi-Ganesh, Autumn Haagsma, Srikanta Mishra, Jared Hawkins, Jared Schuetter, Amy Lang, Jaelen Lewis, Derrick James, Jorge Barrios, Stuart Skopec, and Sanjay Mawalkar (Battelle). Chris Korose, Carl Carmen, Nate Grigsby, Nathan Webb (Illinois State Geological Survey). Principal Investigators: Dr Neeraj Gupta, Dr. Chris Korose.

MRCI,NRAP,data collection,data compilation,legacy

Unpaired image translation to mitigate domain shift in liquid argon time projection chamber detector responses

Deep learning algorithms often are developed and trained on a training dataset and deployed on test datasets. Any systematic difference between the training and a test dataset may severely degrade the final algorithm performance on the test dataset—what is known as the domain shift problem . This issue is prevalent in many scientific domains where algorithms are trained on simulated data but applied to real-world datasets. Typically, the domain shift problem is solved through various domain adaptation (DA) methods. However, these methods are often tailored for a specific downstream task, such as classification or semantic segmentation, and may not easily generalize to different tasks. This work explores the feasibility of using an alternative way to solve the domain shift problem that is not specific to any downstream algorithm. The proposed approach relies on modern Unpaired Image-to-Image (UI2I) translation techniques, designed to find translations between different image domains in a fully unsupervised fashion. In this study, the approach is applied to a domain shift problem commonly encountered in Liquid Argon Time Projection Chamber (LArTPC) detector research when seeking a way to translate samples between two differently distributed LArTPC detector datasets deterministically. This translation allows for mapping real-world data into the simulated data domain where the downstream algorithms can be run with much less domain-shift-related performance degradation. Conversely, using the translation from the simulated data to a real-world domain can increase the realism of the simulated dataset and reduce the magnitude of any systematic uncertainties. To evaluate the quality of the translations, we use both pixel-wise metrics and a downstream task to measure the effectiveness of UI2I methods for mitigating the domain shift problem. We adapted several popular UI2I translation algorithms to work on scientific data and demonstrated the viability of these techniques for solving the domain shift problem with LArTPC detector data. To facilitate further development of DA techniques for scientific datasets, the ‘Simple Liquid-Argon Track Samples’ dataset used in this study is also published.

97 MATHEMATICS AND COMPUTING

Aspen Open Jets: unlocking LHC data for foundation models in particle physics

Foundation models are deep learning models pre-trained on large amounts of data which are capable of generalizing to multiple datasets and/or downstream tasks. This work demonstrates how data collected by the CMS experiment at the Large Hadron Collider can be useful in pre-training foundation models for HEP. Specifically, we introduce the AspenOpenJets (AOJs) dataset, consisting of approximately 178 M high p T jets derived from CMS 2016 Open Data. We show how pre-training the OmniJet-α foundation model on AOJs improves performance on generative tasks with significant domain shift: generating boosted top and QCD jets from the simulated JetClass dataset. In addition to demonstrating the power of pre-training of a jet-based foundation model on actual proton–proton collision data, we provide the ML-ready derived AOJs dataset for further public use.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Discovery, Design, Synthesis and Testing of High Performance Structural Alloys (Final Technical Report)

The overarching goal of this project is to understand the phase stability and mechanical behavior of non-stoichiometric multi-principal element alloy (MPEA) materials. In order to identify suitable alloys, we plan to use a combinatorial thin film screening approach, in collaboration with scientists at Lawrence Berkeley National Laboratory who are performing computational work as well as complementary experimental work. Specific tasks within the scope of this project include the fabrication, using thin film deposition from six sputtering targets, of combinatorial samples with multi-dimensional gradients in composition and microstructure. These samples are studied to screen MPEA systems for promising candidate alloys with specific composition(s), based on characterization of composition, structure and mechanical behavior across the thin film. We want to produce single-phase MPEAs with chemical homogeneity in a given thin film region, simple grain structures, and no intermetallic phases present. Gradient films facilitate first-pass screening for desirable characteristics and inform the next stage of work that involves fabrication of bulk MPEA specimens for (tensile) mechanical testing and characterization. To make the bulk alloys, metal (elemental) pieces are melted to form MPEAs, followed by heat treatment to homogenize the composition and microstructure. A subset of alloys is also cast, using vacuum arc melting, to yield larger samples (diameter ~1 cm and length ~5-10 cm) and these allow us to assess viability of scale-up for the alloys in structural applications. Further processing plans include rolling and heat treatment to recrystallize selected bulk MPEAs and grow grains to different extents, in order to investigate size effects in the mechanical behavior of MPEAs. Microspecimen testing will be performed (primarily in tension) to assess the mechanical behavior over a range of temperatures. The deformation microstructure of mechanically tested alloys will be characterized using transmission electron microscopy (TEM) to provide a scientific basis for understanding the structure-property relationships in MPEA mechanical behavior.

36 MATERIALS SCIENCE

Evaluation of Howard A. Hanson Dam Juvenile Fish Passage and Survival Study Live Fish Injury Assessment, Sensor Fish, and BioPA Modeling Tasks

The live fish injury assessment, Sensor Fish, and BioPA modeling study tasks were conducted by researchers from Pacific Northwest National Laboratory (PNNL). The four tasks were part of the larger Evaluation of Howard A. Hanson Dam (HAHD) Juvenile Fish Passage and Survival study, which had six total tasks. To achieve study objectives for each of the four tasks, field work occurred at Green Peter Dam (GPR) to evaluate the highest elevation steep slope bypass pipe, at HAHD to evaluate baseline conditions of the horseshoe tunnel, and at PNNL’s Aquatic Research Laboratory (ARL) to evaluate simulated dam passage conditions (i.e., shear forces and collision). Each of these evaluations utilized live fish injury assessment, Sensor Fish, and BioPA modeling. Live fish injury assessment and survival (tagged with and without balloon or passive integrated transponder [PIT] tags) was correlated with Sensor Fish to determine thresholds. The CFD analyses were then performed, and the computed values were compared to the corresponding measured values of Sensor Fish data. The results of the overall injury and survival of fish was also used in the validation of the CFD modeling method. Collectively, the results will aid in future modeling of fish passage at HAHD. Results from these tasks can be used by biologists, engineers, resource managers, and regional decision-makers to inform baseline conditions under current operations and the engineering design of the new FPF at HAHD. This draft report contains initial data and results from the four tasks. Table 8 1, Table 8 2, and Table 8 3, and Figure 8 1, Figure 8 2, and Figure 8 3 depict the CFD modeling findings for the GPR steep slope bypass, HAHD horseshoe tunnel, and laboratory testing. Table 8 4, Table 8 5, and Table 8 6 depict the Sensor Fish findings for the GPR steep slope bypass and HAHD horseshoe tunnel testing. The Mv values observed in the HAHD were significantly lower compared to the laboratory experiments conducted at PNNL. Currently, investigations are underway to understand the reasons for this disparity and to establish an appropriate threshold value for Mv. Survival predictions presented in the tables below should be considered preliminary and should not be used until further analyses and adjustments are completed. The next steps for modeling will include the flow regime, (i.e., density of flow regimes due to water and air mixing ) to continue to improve on the threshold value for Mv.

13 HYDRO ENERGY

DECOVALEX-2023: Task C Final Report

The Full-scale Emplacement (FE) heater experiment at the Mont Terri Underground Rock Laboratory (URL) was designed and conducted by Nagra to replicate an emplacement tunnel of Nagra’s reference repository design at 1:1 scale. Alongside testing the technical feasibility of constructing disposal tunnels, emplacing waste containers in the tunnels and then backfilling them, the main goals of the FE experiment are (1) to obtain a better understanding of the coupled effects of induced thermo-hydro-mechanical (THM) processes that may occur and (2) to validate existing coupled THM models (Müller et al., 2017). A key aspect of ensuring safety for repositories located in low-permeability rock involves minimizing any damage to the rock itself, thereby preserving its integrity and promoting a stable environment Amongst a number of processes that could damage the rock is the increase in pore pressure due to thermal loading caused by heat emitted from the waste. To reduce the potential damage of the rock, it is important to analyse the evolution of heat over time due to the heat load of the containers and assess possible consequences by coupled THM models. The aim of Task C of DECOVALEX-2023 was to build 3D numerical models of the FE experiment, focussing in particular on the heating induced pore pressure change in the Opalinus Clay. Data from a large number of sensors were available from the FE experiment for model comparison. These sensors measured temperature and relative humidity in the bentonite around the heaters, and temperature, pressure and displacement/strain in the surrounding Opalinus clay. Data were available from the start of excavation (April 2012) up to August 2020 for most sensors (more than 5 years from the start of heating in December 2014). To fulfil the overall aim of the task, the work was broken down into a number of steps, starting with simpler models to build confidence in each team’s approach and then moving to more complex models that better represent the FE experiment. Step 0 consisted of 2D benchmark models, gradually increasing the number of processes that are represented from thermal (T) only models in Step 0a, to coupled thermal hydraulic (TH) models in Step 0b with a representation of changing porosity, to coupled thermo-hydro-mechanical (THM) models in Step 0c, where porosity changes are calculated by the mechanical model. A detailed specification of processes, parameters, initial and boundary conditions was provided for this step, with the ambition that all teams would work towards close agreement in their model results, thus building confidence in the model implementations. vi It was not straightforward to achieve agreement between the teams, so additional steps (Step 0b2, 0b3, 0c2, 0c3) were added along with derivation of some analytical solutions against which the models could be compared. The reasons for the differences between teams were investigated and found to be caused primarily by different conceptual model assumptions (including temperature dependence of the thermal expansion of water), different model formulations (including porosity evolution) and differences in modelled domain sizes, boundary conditions and grid discretisation. This demonstrates that comparisons between multiple modelling teams and/or comparison with analytical results and experimental data are highly beneficial in providing an indication of uncertainty in model predictions. At the conclusion of Step 0, almost all teams had achieved a close agreement in model results and those that had not achieved an agreement knew the reason for this. Step 1 moved from 2D models to 3D models of the FE experiment without adding technical features like shotcrete or EDZ, and only considering the heating phase. Initially the 3D model was tightly specified to continue to build confidence in the model implementations (Step 1a). The results of Step 1a were compared to the data from the FE-experiment without the teams seeing the data. The teams were then provided with a sub-set of the data from the FE-experiment and invited to consider how best to use the large dataset for model comparison (Step 1b). Teams were then asked to use the data provided to calibrate their models, only changing material property values rather than adding features or processes to their models (Step 1c). In Step 1, teams were asked to only model the heating phase of the experiment, so pressure in the Opalinus Clay was reported as change in pressure since the initial conditions were specified rather than modelled. The change from 2D to 3D models was accompanied by an increase in the dispersion of results between the teams. Some of this was resolved during the task, but some remained and is potentially due to model discretisation. Calibration of parameters was useful in improving the fit of the models to the data but the remaining differences indicated that the models were missing features or processes. In Step 2, the teams were asked to update their models with additional features and processes as well as calibrating parameters to try and improve the fit of the models to the data. Teams were encouraged to represent ventilation of the open FE tunnel prior to backfilling with heaters and bentonite and in Step 2, the absolute pressure in the Opalinus Clay was compared between the teams. Teams took different approaches, but there was consideration of adding shotcrete and an EDZ into the model, representing stress change during excavation and different approaches to modelling ventilation of the FE tunnel. Overall, the documented results showed a very good agreement for temperature. The results for porewater pressure evolution showed a significant improvement for most teams compared to Step 1c with a good agreement to the measurements for several teams whereas some teams overpredicted the pressure increase and others overpredicted the drainage effect especially for the sensors close to the heater. Step 3 was an opportunity for teams to use the models developed in Step 1 and Step 2 to make predictions about the temperature and pressure changes that will be expected at the FE experiment over the next few years in light of the planned changes in thermal output of the heaters.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Transforming the bootstrap: using Transformers to compute scattering amplitudes in planar N = 4 Super Yang-Mills theory

Abstract We pursue the use of deep learning methods to improve state-of-the-art computations in theoretical high-energy physics. Planar N = 4 Super Yang-Mills theory is a close cousin to the theory that describes Higgs boson production at the Large Hadron Collider; its scattering amplitudes are large mathematical expressions containing integer coefficients. In this paper, we apply Transformers to predict these coefficients. The problem can be formulated in a language-like representation amenable to standard cross-entropy training objectives. We design two related experiments and show that the model achieves high accuracy (> 98%) on both tasks. Our work shows that Transformers can be applied successfully to problems in theoretical physics that require exact solutions.

Cai, Tianji (ORCID:0000000232359486)

Predicting Critical Transitions in Multiscale Data

Predicting the dynamics of complex nonlinear systems remains a challenging problem both in dynamical systems theory as well as real world science and engineering applications. Data-driven methods utilizing the latest advances in machine learning (ML) provide a promising new paradigm for this task. Our work centered on Reservoir Computing (RC), which has shown itself to be capable of skillfully predicting chaotic dynamics in multiscale systems. In the first part of the work, the focus is on how to improve predictions of critical transitions in a class of slow-fast metastable systems in which the equations are known. An additional goal was to determine whether a relationship exists between RC and Koopman operator theory, to improve the efficiency and broaden the applicability of the approach. In the second part of this work, a variation on the RC model known as Reconstructive Reservoir Computing (RRC) is applied to real-world data to identify anomalies.

97 MATHEMATICS AND COMPUTING

Integrity Monitoring and Assessment, Prediction, Repair, and Corrosion Control of the Hanford Storage Tanks — 25213

BACKGROUND AND PROJECT TASKS Proposed work addresses focus area 1: Waste Retrieval, Transport and Closure, with particular focus on increasing volume available for tank storage 1) Refurbish/fortify existing double shell storage tanks (Tank Refurbish – Polymer Grout) 2) Robust monitoring tools for internal corrosion (Integrity Monitoring -Reference Electrodes) 3) Mitigate external corrosion (Degradation Prevention - Cathodic Protection) 4) Determine viability of constructing new tanks, and storage room creation by way of evaporation (Cost Benefit Analysis)

Shukla, Pavan K. [Savannah River National Laborato

Detection and Association of Operational Events using DAS and Seismometers (FY 2025 Mid-Year Report)

This mid-year report summarizes ongoing work to identify anomalous vibration signals indicative of potential containment breaches. This work includes compiling continuous seismic datasets and testing and refining underground detection and geolocation techniques. In the first two quarters of FY25, we have completed two project work plan tasks: (1) creating a database of continuous waveforms and ground truth event data from multiple modalities and (2) refining and implementing a detection and association algorithm to create a catalog of anomalous underground activities. This report contains a summary of the seismic database including the continuous seismic data collected by a dense array of surface seismic stations above Pleasant Gap Mine, and continuous seismic data collected using subsurface distributed acoustic sensing (DAS) in the subsurface at Sanford Underground Research Facility (SURF) and the ground truth information gathered from both sites. This report also includes results from refining and applying a dynamic power spectral density detector to both continuous seismic datasets. Finally, the report provides an initial catalog of subsurface operational events from both sensing modalities.

58 GEOSCIENCES

Tandem Predictions for HPC Jobs

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

HPC

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING