Search NASA⌕ Search

SEARCH · Search NASA

Results for “Training Time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Simple Diesel Train Fuel Consumption Model for Real-Time Train Applications

This paper introduces a simple diesel train energy consumption model that calculates the instantaneous energy consumption using vehicle operational input variables, including the instantaneous speed, acceleration, and roadway grade, which can be easily obtained from global positioning system (GPS) loggers. The model was tested against real-world data and produced an error of -1.33% for all data and errors ranging from -12.4% to +8.0% for energy consumption of four train datasets amounting to a total of 5854 km trips. The study also validated the proposed model with separate data that were collected between Valencia and Cuenca, Spain, which had a total length of 198 km and found that the model was accurate, yielding a relative error of -1.55% for the total energy consumption. These results show that the proposed model can be used by train operators, transportation planners, policy makers, and environmental engineers to evaluate the energy consumption effects of train operational projects and train simulation within intermodal transportation planning tools.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Simultaneously improving accuracy and computational cost under parametric constraints in materials property prediction tasks

Abstract Modern data mining techniques using machine learning (ML) and deep learning (DL) algorithms have been shown to excel in the regression-based task of materials property prediction using various materials representations. In an attempt to improve the predictive performance of the deep neural network model, researchers have tried to add more layers as well as develop new architectural components to create sophisticated and deep neural network models that can aid in the training process and improve the predictive ability of the final model. However, usually, these modifications require a lot of computational resources, thereby further increasing the already large model training time, which is often not feasible, thereby limiting usage for most researchers. In this paper, we study and propose a deep neural network framework for regression-based problems comprising of fully connected layers that can work with any numerical vector-based materials representations as model input. We present a novel deep regression neural network, iBRNet, with branched skip connections and multiple schedulers, which can reduce the number of parameters used to construct the model, improve the accuracy, and decrease the training time of the predictive model. We perform the model training using composition-based numerical vectors representing the elemental fractions of the respective materials and compare their performance against other traditional ML and several known DL architectures. Using multiple datasets with varying data sizes for training and testing, We show that the proposed iBRNet models outperform the state-of-the-art ML and DL models for all data sizes. We also show that the branched structure and usage of multiple schedulers lead to fewer parameters and faster model training time with better convergence than other neural networks. Scientific contribution: The combination of multiple callback functions in deep neural networks minimizes training time and maximizes accuracy in a controlled computational environment with parametric constraints for the task of materials property prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Power System Event Identification with Transfer Learning Using Large-scale Real-world Synchrophasor Data in the United States

The lack of sufficient labeled events and long training time limit the applicability of deep neural network-based power system event identification using synchrophasor data. In this paper, we propose to leverage transfer learning technique to boost the reliability and reduce the required training time of neural classifier for power system event identification. We use the weights of a neural classifier trained on one transmission system as the initial parameters of another neural classifier for a different transmission system. Numerical tests with real-world synchrophasor data from the Eastern and Western Interconnections of the United States show that the proposed transfer learning approach is very effective in not only improving the training reliability but also reducing the training time.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enabling real-time adaptation of machine learning models at x-ray Free Electron Laser facilities with high-speed training optimized computational hardware

The emergence of novel computational hardware is enabling a new paradigm for rapid machine learning model training. For the Department of Energy’s major research facilities, this developing technology will enable a highly adaptive approach to experimental sciences. In this manuscript we present the per-epoch and end-to-end training times for an example of a streaming diagnostic that is planned for the upcoming high-repetition rate x-ray Free Electron Laser, the Linac Coherent Light Source-II. We explore the parameter space of batch size and data parallel training across multiple Graphics Processing Units and Reconfigurable Dataflow Units. We show the landscape of training times with a goal of full model retraining in under 15 min. Although a full from scratch retraining of a model may not be required in all cases, we nevertheless present an example of the application of emerging computational hardware for adapting machine learning models to changing environments in real-time, during streaming data acquisition, at the rates expected for the data fire hoses of accelerator-based user facilities.

97 MATHEMATICS AND COMPUTING↗

Object Detection and Recognition with PointPillars in LiDAR Point Clouds – Comparisions

In the field of autonomous systems, neural networks have been leveraged for object detection and recognition in 2-dimensional images captured by cameras. Other types of sensors are available for sensing surroundings, including LiDAR sensors, and corresponding networks have been developed to perform detection and recognition in the point clouds generated by these sensors. The approaches are similar, both perform convolutions, but have distinct characteristics and challenges. In designing and configuring autonomous systems, a variety of LiDAR sensors are available, along with configurable deep neural networks to leverage their data. This work presents a review of the PointPillars network, an evolution of the seminal PointNet, comparing accuracy and training time relative to different LiDAR sensors, network and training parameters, CPU and GPU hardware, and the criticality of the use of reflective intensity as a feature. The value of using reflectivity as a predictive feature is explored and quantified to determine if it makes a significant difference in accuracy of the PointPillars network. Two separate LiDAR sensors are utilized, a 16-plane and a 32-plane, and corresponding accuracies and training times with the PointPillars network are evaluated.

LiDAR, machine learning, neural network, object re↗

Accelerated Over-The-Air Neural Receiver Training Using Self-Contrastive Learning

Self-contrastive learning (SCL), a self-supervised learning method, has been shown to improve image and signal classifier accuracies and reduce the training time for neural communications receivers. In particular, prior work has shown that SCL applied as a pre-training step can improve simulated performance of OFDM in 3GPP TDL channel models by reducing the training time of the downstream classification task (demodulation and demapping). In this work a practical implementation demonstrating SCL pre-training using software defined radios (SDRs) is proposed.

Cooke, Corey [ORNL] (ORCID:0000000234263672)↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Tensor train continuous time solver for quantum impurity models

The simulation of strongly correlated quantum impurity models is a significant challenge in modern condensed matter physics that has multiple important applications. Thus far, the most successful methods for approaching this challenge involve Monte Carlo techniques that accurately and reliably sample perturbative expansions to any order. However, the cost of obtaining high precision through these methods is high. Recently, tensor train decomposition techniques have been developed as an alternative to Monte Carlo integration. In this study, we apply these techniques to the single-impurity Anderson model at equilibrium by calculating the systematic expansion in power of the hybridization of the impurity with the bath. Furthermore, we demonstrate the performance of the method in a paradigmatic application, examining the first-order phase transition on the infinite-dimensional Bethe lattice, which can be mapped to an impurity model through dynamical mean field theory. Our results indicate that using tensor train decomposition schemes allows the calculation of finite-temperature Green's functions and thermodynamic observables with unprecedented accuracy. The methodology holds promise for future applications to frustrated multiorbital systems, using a combination of partially summed series with other techniques pioneered in diagrammatic and continuous time quantum Monte Carlo.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Efficient learning of power grid voltage control strategies via model-based deep reinforcement learning

Here this article proposes a model-based deep reinforcement learning (DRL) method to design emergency control strategies for short-term voltage stability problems in power systems. Recent advances show promising results for model-free DRL-based methods in power systems control problems. But in power systems applications, these model-free methods have certain issues related to training time (clock time) and sample efficiency; both are critical for making state-of-the-art DRL algorithms practically applicable. DRL-agent learns an optimal policy via a trial-and-error method while interacting with the real-world environment. It is also desirable to minimize the direct interaction of the DRL agent with the real-world power grid due to its safety-critical nature. Additionally, the state-of-the-art DRL-based policies are mostly trained using a physics-based grid simulator where dynamic simulation is computationally intensive, lowering the training efficiency. We propose a novel model-based DRL framework where a deep neural network (DNN)-based dynamic surrogate model (SM), instead of a real-world power grid or physics-based simulation, is utilized within the policy learning framework, making the process faster and more sample efficient. However, having stable training in model-based DRL is challenging because of the complex system dynamics of large-scale power systems. We addressed these issues by incorporating imitation learning to have a warm start in policy learning, reward-shaping, and multi-step loss in surrogate model training. Finally, we achieved 97.5% reduction in samples and 87.7% reduction in training time for an application to the IEEE 300-bus test system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Language models for the prediction of SARS-CoV-2 inhibitors

The COVID-19 pandemic highlights the need for computational tools to automate and accelerate drug design for novel protein targets. We leverage deep learning language models to generate and score drug candidates based on predicted protein binding affinity. We pre-trained a deep learning language model (BERT) on ∼9.6 billion molecules and achieved peak performance of 603 petaflops in mixed precision. Our work reduces pre-training time from days to hours, compared to previous efforts with this architecture, while also increasing the dataset size by nearly an order of magnitude. For scoring, we fine-tuned the language model using an assembled set of thousands of protein targets with binding affinity data and searched for inhibitors of specific protein targets, SARS-CoV-2 Mpro and PLpro. We utilized a genetic algorithm approach for finding optimal candidates using the generation and scoring capabilities of the language model. Our generalizable models accelerate the identification of inhibitors for emerging therapeutic targets.

Blanchard, Andrew E.↗

Quantifying uncertainty in machine learning for nuclear binding energy

Techniques from artificial intelligence and machine learning are increasingly employed in nuclear theory; however, the uncertainties that arise from the complex parameter manifold encoded by the neural networks are often overlooked. Epistemic uncertainties arising from training the same network multiple times for an ensemble of initial weight sets offer a first insight into the confidence of machine learning predictions, but they often come with a high computational cost. Instead, we apply a single-model uncertainty quantification method called Δ-UQ that gives epistemic uncertainties with one-time training. Here, we demonstrate our approach on a two-feature model of nuclear binding energies per nucleon with proton and neutron number pairs as inputs. We show that Δ-UQ can produce reliable and self-consistent epistemic uncertainty estimates and can be used to assess the degree of confidence in predictions made with deep neural networks.

Huang, Mengyao [Lawrence Livermore National Labora↗

Structure–Property Linkage in Alloys Using Graph Neural Network and Explainable Artificial Intelligence

Deep learning tools have recently shown significant potential for accelerating the prediction of microstructure–property linkage in materials. While deep neural networks like convolution neural networks (CNNs) can extract physics information from 3D microstructure images, they often require a large network architecture and substantial training time. In this research, we trained a graph neural network (GNN) using phase field generated microstructures of Ni-Al alloys to predict the evolution of mechanical properties. We found that a single GNN is capable of accurately predicting the strengthening of Ni-Al alloys with microstructures of varying sizes and dimensions, which cannot otherwise be done with a CNN. Additionally, GNN requires significantly less GPU utilization than CNN and offers more interpretable explanation of predictions using saliency analysis as features are manually defined in the graph. We also utilize explainable artificial intelligence tool Bayesian Inference to determine the coefficients in the power law equation that governs coarsening of precipitates. Overall, our work demonstrates the ability of the GNN to accurately and efficiently extract relevant information from material microstructures without having restrictions on microstructure size or dimension and offers an interpretable explanation.

Chemistry↗

Automated and efficient local adaptive regression for principal component-based reduced-order modeling of turbulent reacting flows

Principal Component Analysis can be used to reduce the cost of Computational Fluid Dynamics simulations of turbulent reacting flows by reducing the dimensionality of the transported variables through projection of the thermochemical state onto a lower-dimensional manifold. However, because of the nonlinearity of the principal component source terms, nonlinear regression techniques must be utilized for the source terms in terms of the principal components. Unfortunately, widely available and utilized nonlinear regression techniques can have prohibitive computational requirements and/or accuracy that is highly dependent on user experience in ad hoc tuning of model architecture and hyperparameters. Here, in this work, a new nonlinear regression approach is proposed that is both computationally efficient and automated so does not require any user input. The approach is evaluated through a priori prediction of principal component source terms using data from a Direct Numerical Simulation of a turbulent nonpremixed n-heptane/air jet flame. In particular, the proposed framework consists of local regressions whose complexity is adapted according to the local nonlinearity of the data: local linear regression when accurate enough and local Artificial Neural Networks when nonlinear regression is required. The number of local clusters for local regression is determined automatically using the Davies-Bouldin index. In addition, Bayesian optimization is utilized for model training (i.e., to select the best architectures and hyperparameters of the nonlinear regressions in an unsupervised fashion), eliminating ad hoc hand-tuning and/or expensive grid searches. Overall, compared to a single, global neural network, the new local adaptive regression approach is shown to have comparable accuracy but 69% less training time due to the utilization of local linear regression and faster training of local neural networks.

42 ENGINEERING↗

Effect of image resolution on automated classification of chest X-rays

Deep learning (DL) models have received much attention lately for their ability to achieve expert-level performance on the accurate automated analysis of chest X-rays (CXRs). Recently available public CXR datasets include high resolution images, but state-of-the-art models are trained on reduced size images due to limitations on graphics processing unit memory and training time. As computing hardware continues to advance, it has become feasible to train deep convolutional neural networks on high-resolution images without sacrificing detail by downscaling. This study examines the effect of increased resolution on CXR classification performance. We used the publicly available MIMIC-CXR-JPG dataset, comprising 377,110 high resolution CXR images for this study. We applied image downscaling from native resolution to 2048 × 2048 pixels, 1024 × 1024 pixels, 512 × 512 pixels, and 256 × 256 pixels and then we used the DenseNet121 and EfficientNet-B4 DL models to evaluate clinical task performance using these four downscaled image resolutions. We find that while some clinical findings are more reliably labeled using high resolutions, many other findings are actually labeled better using downscaled inputs. We qualitatively verify that tasks requiring a large receptive field are better suited to downscaled low resolution input images, by inspecting effective receptive fields and class activation maps of trained models. Lastly, we show that stacking an ensemble across resolutions outperforms each individual learner at all input resolutions while providing interpretable scale weights, indicating that diverse information is extracted across resolutions.

47 OTHER INSTRUMENTATION↗

Training reinforcement learning models via an adversarial evolutionary algorithm

When training for control problems, more episodes used in training usually leads to better generalizability, but more episodes also requires significantly more training time. There are a variety of approaches for selecting the way that training episodes are chosen, including fixed episodes, uniform sampling, and stochastic sampling, but they can all leave gaps in the training landscape. In this work, we describe an approach that leverages an adversarial evolutionary algorithm to identify the worst performing states for a given model. We then use information about these states in the next cycle of training; this process can be repeated until the desired level of model performance is met. We demonstrate this approach with the OpenAI Gym cart-pole problem. With this problem, we show that the adversarial evolutionary algorithm did not reduce the number of episodes required in training needed to attain model generalizability when compared with stochastic sampling, and actually performed slightly worse.

Coletti, Mark↗