Search NASA⌕ Search

SEARCH · Search NASA

Results for “learning rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Open Architecture for Cost Savings in Advanced Nuclear Reactors

Recently, nuclear power plant build projects in the West have run over budget due to high capital costs and schedule overruns. Compared to other sources of energy, nuclear power plants have higher capital costs. Reactors are often different at every site, resulting in a lack of standardization. Nuclear is expected to compete with other low carbon sources of energy which have lower capital costs making it essential for nuclear to develop ways of reducing costs. Strategies such as standardization, learning rates, modularization, and schedule reduction in advanced reactors can reduce nuclear costs by about 40%. Standardization as a way of cutting capital costs has been explored even in large nuclear power plants. Standardization of certain plant components can result in lower component and installation costs and higher learning from experience. Standardization can be achieved by adopting a criterion of key performance indicators and general design principles for a specific system or component such as the balance of plant. Modularization allows the construction of certain components of SMRs in a factory, which saves time, increases productivity, and encourages higher learning rates. Production learning decreases the time and the cost related to an activity. The potential for modularized components of advanced reactors to be manufactured in factories makes it conducive to achieving higher learning rates. Developing large-capacity nuclear programs through sequential builds cultivates a higher learning rate, which in effect may reduce schedule overruns. Open architecture has been identified as a way to drive standardization among advanced reactor designs and result in cost savings. Open architecture (OA) is defined as a design enabling a diverse supply chain by defining and publishing requirements of systems or equipment in functional and/or interface terms, utilizing technical standards in widespread use. Currently, the nuclear industry’s approach is to use closed architecture, making most designs proprietary. However, collaboration between various advanced reactor vendors and suppliers utilizing the concept of open architecture can result in modular and standardized architecture of subsystems or subcomponents of a nuclear power plant. Completely standardizing nuclear power plants may be impossible, however, certain common subsystems amongst the various reactor designs could be standardized and/or access a wider supply chain and leverage existing learning from other sectors. Open architecture will save time and allocate resources to the parts of the plants that have the most unique features. A key advantage of open architecture is its ability to improve production learning across advanced reactors (AR) types in the industry, by providing and utilizing the same kind of component. Sodium fast reactor (SFR), High Temperature Gas Reactor (HTGR) and Molten Salt Reactor (MSR) are the advanced reactors considered for this project. This paper aims to determine the cost savings in advanced reactor programs due to open architecture learning rate. This work is an extension of work done on light water reactor small modular reactors; the cost methodology was utilized to investigate the impact of open architecture on advanced reactors with a particular focus on sodium fast reactors. The cost data on sodium fast reactors used in the model presented the most adequate information required for the analysis.

Advanced Nuclear Reactors↗

Meta-Analysis of Advanced Nuclear Reactor Cost Estimations

Supporting Data can be downloaded at: https://gain.inl.gov/content/uploads/4/2024/06/INL-RPT-24-77048-R1.xlsx Nuclear energy is a critical cornerstone of the current United States clean energy supply and may play a larger role in the future in support of a transition to a net-zero economy. The current fleet of nuclear reactors predominantly consists of large light-water reactors (LWRs), while many of the reactor designs under consideration are smaller and/or different technologies. Because these new designs have not yet been built, there is a high degree of uncertainty associated with their cost. This complicates energy-planning efforts because cost projections are not always standardized, consistent, and centralized in an easily accessible location. To help support energy planning in the US, this report provides advanced nuclear cost ranges using a transparent methodology along with other relevant information that can be used to help support decision making and energy planning. The purpose of this work was to conduct a methodical process for cost evaluation using only public information that was vetted with the end-goal to provide reference cost projections for nuclear energy. To provide a solid basis for these values, the approach and assumptions are explicitly laid out throughout the report allowing any user of the data to challenge or reconsider them. Because future US nuclear-reactor costs are still unknown due to little recent observed data, the report opted to compile a comprehensive list of bottom-up estimates and evaluate averages/trends within the data to identify reference ranges. This was deemed preferable to opining on the robustness or validity of one cost estimation versus another. To that end, the work evaluated thousands of lines of cost subaccounts from several bottom-up cost estimates. A wide variety of different reactor types captured in the data are of various sizes and technologies. Some of these reactors will be representative of advanced reactors under development while others will not. Thus, the results here are dependent on the data that are available and the accuracy of the estimates that are used. Each bottom-up estimate was reviewed to determine whether it was complete. Incomplete data sets were corrected to ensure an adequate basis of cross-comparison. The report is not without limitations and should be interpreted as an initial step to develop cost ranges for nuclear technology. Ultimately, future work can build upon the methodology with refined cost estimates to reduce uncertainty. US-based overnight capital cost (OCC) estimates were compiled from extensive data sets into ranges for both large and small reactor sizes for 2030. To project the cost declines over time, learning rates were sampled from literature sources. No SMRs were previously built; hence, learning rates based on bottom-up approaches (e.g., by quantifying the impact stemming from fabrication of different components, modular work, site construction, commissioning) were prioritized. For larger reactors, actual learning rates from deployments were used to project future costs (adjusted to account for standardization or lack thereof between designs). Other costs included are fixed and variable operations and maintenance costs. The final variables were capacity factors and ramp rates to support energy planning.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A Systematic Framework for Projecting the Future Cost of Offshore Wind Energy

Offshore wind costs are expected to decline rapidly in the short and medium term future as the industry grows and gains experience in manufacturing, installing, and operating commercial scale projects. Estimating the future costs of offshore wind energy is critical for evaluating the technology's economic performance, how it can fit into a broader clean energy economy, and how R&D investment can be allocated to advance the technology. We present a newly developed approach for forecasting these costs which focuses on an empirically-derived learning rate for capital costs and prescribed cost and performance improvements for operational costs and capacity factor. We establish baseline costs for a series of reference fixed-bottom and floating projects in 2021 and project cost trajectories to 2035, presenting both an average cost trajectory as well as describing the range of potential future costs associated with site-specific cost variations and uncertainty in the estimate of the learning rate. We also conduct sensitivity analyses showing the impact of different global deployments by 2035 and variations in the prescribed operational costs and capacity factors. The results show that fixed-bottom and floating offshore wind capital costs could decrease to around $\$$2,400/kW and $\$$3,300/kW by 2035, respectively, with ranges of $\$$2,100/kW - $\$$2,750/kW for fixed-bottom projects and $\$$2,850/kW - $\$$5,500/kW for floating projects. The levelized cost of energy of fixed-bottom and floating wind projects could decrease to $\$$53.1/MWh and $\$$63.9/MWh by 2035, with ranges of $\$$48.4/MWh - $\$$59.7/MWh for fixed-bottom projects and $\$$46.5/MWh - $\$$99.9/MWh for floating projects. By presenting the uncertainty associated with the forecast we provide a transparent description of the spectrum of potential cost trajectories for offshore wind.

17 WIND ENERGY↗

BUTTER - Empirical Deep Learning Dataset

The BUTTER Empirical Deep Learning Dataset represents an empirical study of the deep learning phenomena on dense fully connected networks, scanning across thirteen datasets, eight network shapes, fourteen depths, twenty-three network sizes (number of trainable parameters), four learning rates, six minibatch sizes, four levels of label noise, and fourteen levels of L1 and L2 regularization each. Multiple repetitions (typically 30, sometimes 10) of each combination of hyperparameters were preformed, and statistics including training and test loss (using a 80% / 20% shuffled train-test split) are recorded at the end of each training epoch. In total, this dataset covers 178 thousand distinct hyperparameter settings ("experiments"), 3.55 million individual training runs (an average of 20 repetitions of each experiments), and a total of 13.3 billion training epochs (three thousand epochs were covered by most runs). Accumulating this dataset consumed 5,448.4 CPU core-years, 17.8 GPU-years, and 111.2 node-years.

Array↗

Gravity Well Commercial Economics Assessment: Potential Revenue and Cost: Cooperative Research and Development (Final Report)

In the Gravity Well Revenue Study, we evaluate the potential revenue from energy storage using historical energy-only electricity prices, forward-looking projections of hourly electricity prices, and actual reported revenue. This analysis examines the impact of storage duration and round-trip efficiency, as well as the location of the storage, on storage revenue within the current and projected U.S. power system. We also investigated the impact of round-trip efficiency on storage revenue. We found that the relationship between storage revenue and round-trip efficiency is nonlinear. The value of improved round-trip efficiency declines as round-trip efficiency increases. In the Gravity Well Future Cost Study, we applied learning curves to predict the future cost trajectory of Gravity Wells (GrWs). Two types of analysis were implemented. The first was a bottom-up analysis that used historical learning rates for cost components, such as motors and gearboxes, and cost categories (e.g., engineering and design, etc.) to determine the learning-by-doing based single-factor learning curve. The single factor learning curve expresses the relationship between the cost of GrW and the number of units deployed (or the cumulative capacity). In the second analysis, we predicted future GrW costs via a top-down approach. This approach accounts for historical cost trends in other renewable energy and storage technologies, which have similarities with GrWs. Using a multifactor learning curve that accounts for both intrinsic (cumulative capacity) and extrinsic (the elasticity in the price of steel) factors, we estimated the future cost of GrWs.

25 ENERGY STORAGE↗

Towards provably efficient quantum algorithms for large-scale machine-learning models

Large machine learning models are revolutionary technologies of artificial intelligence whose bottlenecks include huge computational expenses, power, and time used both in the pre-training and fine-tuning process. In this work, we show that fault-tolerant quantum computing could possibly provide provably efficient resolutions for generic (stochastic) gradient descent algorithms, scaling as $\mathcal{O}$(T 2 x polylog($n$)), where n is the size of the models and T is the number of iterations in the training, as long as the models are both sufficiently dissipative and sparse, with small learning rates. Based on earlier efficient quantum algorithms for dissipative differential equations, we find and prove that similar algorithms work for (stochastic) gradient descent, the primary algorithm for machine learning. In practice, we benchmark instances of large machine learning models from 7 million to 103 million parameters. We find that, in the context of sparse training, a quantum enhancement is possible at the early stage of learning after model pruning, motivating a sparse parameter download and re-upload scheme. Our work shows solidly that fault-tolerant quantum algorithms could potentially contribute to most state-of-the-art, large-scale machine-learning problems.

97 MATHEMATICS AND COMPUTING↗

Surrogate Model Based Optimization for Finding Robust Deep Learning Model Architectures

Deep Learning (DL) models are increasingly used throughout the sciences. However, their performance and usefulness depend greatly on their architecture which is defined by hyperparameters such as the number of nodes, layers, the learning rate, etc. Tuning these hyperparameters is time-consuming because evaluating their performance requires a lengthy training step. Stochastic optimizers used in training lead to performance variability and potentially prediction reliability issues. In this talk, we will describe an automated optimization method based on surrogate models and active learning strategies for tuning DL model architectures. We take into account the prediction variability with the goal to identify architectures that make reliable and robust predictions. We demonstrate our developments on an application arising in particle physics.

deep learning↗

Unraveling the impact of initial choices and in-loop interventions on learning dynamics in autonomous scanning probe microscopy

The current focus in Autonomous Experimentation (AE) is on developing robust workflows to conduct the AE effectively. This entails the need for well-defined approaches to guide the AE process, including strategies for hyperparameter tuning and high-level human interventions within the workflow loop. This paper presents a comprehensive analysis of the influence of initial experimental conditions and in-loop interventions on the learning dynamics of Deep Kernel Learning (DKL) within the realm of AE in scanning probe microscopy. We explore the concept of the “seed effect,” where the initial experiment setup has a substantial impact on the subsequent learning trajectory. Additionally, we introduce an approach of the seed point interventions in AE allowing the operator to influence the exploration process. Using a dataset from Piezoresponse Force Microscopy on PbTiO 3 thin films, we illustrate the impact of the “seed effect” and in-loop seed interventions on the effectiveness of DKL in predicting material properties. The study highlights the importance of initial choices and adaptive interventions in optimizing learning rates and enhancing the efficiency of automated material characterization. This work offers valuable insights into designing more robust and effective AE workflows in microscopy with potential applications across various characterization techniques.

47 OTHER INSTRUMENTATION↗

Improving Deep Neural Networks’ Training for Image Classification With Nonlinear Conjugate Gradient-Style Adaptive Momentum

Momentum is crucial in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs). In deep learning practice, the momentum is usually weighted by a well-calibrated constant. However, tuning the hyperparameter for momentum can be a significant computational burden. In this article, we propose a novel adaptive momentum for improving DNNs training; this adaptive momentum, with no momentum-related hyperparame- ter required, is motivated by the nonlinear conjugate gradient (NCG) method. Stochastic gradient descent (SGD) with this new adaptive momentum eliminates the need for the momentum hyperparameter calibration, allows using a significantly larger learning rate, accelerates DNN training, and improves the final accuracy and robustness of the trained DNNs. For example, SGD with this adaptive momentum reduces classification errors for training ResNet110 for CIFAR10 and CIFAR100 from 5.25% to 4.64% and 23.75% to 20.03%, respectively. Furthermore, SGD, with the new adaptive momentum, also benefits adversarial training and, hence, improves the adversarial robustness of the trained DNNs.

97 MATHEMATICS AND COMPUTING↗

Modeling injection-induced fault slip using long short-term memory networks

Stress changes due to changes in fluid pressure and temperature in a faulted formation may lead to the opening/shearing of the fault. This can be due to subsurface (geo)engineering activities such as fluid injections and geologic disposal of nuclear waste. Such activities are expected to rise in the future making it necessary to assess their short- and long-term safety. Here, a new machine learning (ML) approach to model pore pressure and fault displacements in response to high-pressure fluid injection cycles is developed. The focus is on fault behavior near the injection borehole. To capture the temporal dependencies in the data, long short-term memory (LSTM) networks are utilized. To prevent error accumulation within the forecast window, four critical measures to train a robust LSTM model for predicting fault response are highlighted: (i) setting an appropriate value of LSTM lag, (ii) calibrating the LSTM cell dimension, (iii) learning rate reduction during weight optimization, and (iv) not adopting an independent injection cycle as a validation set. Several numerical experiments were conducted, which demonstrated that the ML model can capture peaks in pressure and associated fault displacement that accompany an increase in fluid injection. The model also captured the decay in pressure and displacement during the injection shut-in period. Further, the ability of an ML model to highlight key changes in fault hydromechanical activation processes was investigated, which shows that ML can be used to monitor risk of fault activation and leakage during high pressure fluid injections.

58 GEOSCIENCES↗

CyBERT: Cybersecurity Claim Classification by Fine-Tuning the BERT Language Model

We introduce CyBERT, a cybersecurity feature claims classifier based on bidirectional encoder representations from transformers and a key component in our semi-automated cybersecurity vetting for industrial control systems (ICS). To train CyBERT, we created a corpus of labeled sequences from ICS device documentation collected across a wide range of vendors and devices. This corpus provides the foundation for fine-tuning BERT’s language model, including a prediction-guided relabeling process. We propose an approach to obtain optimal hyperparameters, including the learning rate, the number of dense layers, and their configuration, to increase the accuracy of our classifier. Fine-tuning all hyperparameters of the resulting model led to an increase in classification accuracy from 76% obtained with BertForSequenceClassification’s original architecture to 94.4% obtained with CyBERT. Furthermore, we evaluated CyBERT for the impact of randomness in the initialization, training, and data-sampling phases. CyBERT demonstrated a standard deviation of ±0.6% during validation across 100 random seed values. Finally, we also compared the performance of CyBERT to other well-established language models including GPT2, ULMFiT, and ELMo, as well as neural network models such as CNN, LSTM, and BiLSTM. The results showed that CyBERT outperforms these models on the validation accuracy and the F1 score, validating CyBERT’s robustness and accuracy as a cybersecurity feature claims classifier.

97 MATHEMATICS AND COMPUTING↗

Assessments of epistemic uncertainty using Gaussian stochastic weight averaging for fluid-flow regression

Here, we use Gaussian stochastic weight averaging (SWAG) to assess the epistemic uncertainty associated with neural-network-based function approximation relevant to fluid flows. SWAG approximates a posterior Gaussian distribution of each weight, given training data, and a constant learning rate. Having access to this distribution, it is able to create multiple models with various combinations of sampled weights, which can be used to obtain ensemble predictions. The average of such an ensemble can be regarded as the 'mean estimation', whereas its standard deviation can be used to construct 'confidence intervals', which enable us to perform uncertainty quantification (UQ) with regard to the training process of neural networks. We utilize representative neural-network-based function approximation tasks for the following cases: (i) a two-dimensional circular-cylinder wake; (ii) the DayMET dataset (maximum daily temperature in North America); (iii) a three-dimensional square-cylinder wake; and (iv) urban flow, to assess the generalizability of the present idea for a wide range of complex datasets. SWAG-based UQ can be applied regardless of the network architecture, and therefore, we demonstrate the applicability of the method for two types of neural networks: (i) global field reconstruction from sparse sensors by combining convolutional neural network (CNN) and multi-layer perceptron (MLP); and (ii) far-field state estimation from sectional data with two-dimensional CNN. We find that SWAG can obtain physically-interpretable confidence-interval estimates from the perspective of epistemic uncertainty. This capability supports its use for a wide range of problems in science and engineering.

97 MATHEMATICS AND COMPUTING↗

AI‐Driven Robot Enables Synthesis‐Property Relation Prediction for Metal Halide Perovskites in Humid Atmosphere

Materials Acceleration Platforms (MAPs) – also known as self-driving laboratories– present a new paradigm for materials science and promise an order of magnitude accelerated materials discovery compared to the traditional trial-and-error approach. Metal halide perovskites (MHPs) are an emerging class of materials for optoelectronic applications but are plagued by irreproducible optoelectronic quality, particularly for films fabricated in a humid atmosphere. Here, in this work, a machine learning (ML)-guided closed-loop platform is developed with a multimodal data fusion approach to predict synthesis–property relations for the optical quality of MHP thin films in relative humidities (RHs) ranging from 5–55%. The efficiency of this approach is confirmed by the fast-dropping learning rate to 2% after experimentally sampling less than 1% of the possible 5,000+ combinations. The prediction of synthesis–property relations is done by optical and imaging characterizations. In situ photoluminescence characterization revealed the origin of thin film quality variation at different RH. These insights provide an avenue for controlling the MHP crystallization by fine-tuning the synthesis parameters and RH for a given chemistry, thus lifting the need for stringent atmosphere control. The MAP enables an accelerated screening and understanding of the synthesis design space, facilitating rational synthesis recipe choice for a wide range of materials.

AI-driven robot↗

A phase transition for finding needles in nonlinear haystacks with LASSO artificial neural networks

To fit sparse linear associations, a LASSO sparsity inducing penalty with a single hyperparameter provably allows to recover the important features (needles) with high probability in certain regimes even if the sample size is smaller than the dimension of the input vector (haystack). More recently learners known as artificial neural networks (ANN) have shown great successes in many machine learning tasks, in particular fitting nonlinear associations. Small learning rate, stochastic gradient descent algorithm and large training set help to cope with the explosion in the number of parameters present in deep neural networks. Yet few ANN learners have been developed and studied to find needles in nonlinear haystacks. Driven by a single hyperparameter, our ANN learner, like for sparse linear associations, exhibits a phase transition in the probability of retrieving the needles, which we do not observe with other ANN learners. To select our penalty parameter, we generalize the universal threshold of Donoho and Johnstone (Biometrika 81(3):425–455, 1994) which is a better rule than the conservative (too many false detections) and expensive cross-validation. In the spirit of simulated annealing, we propose a warm-start sparsity inducing algorithm to solve the high-dimensional, non-convex and non-differentiable optimization problem. We perform simulated and real data Monte Carlo experiments to quantify the effectiveness of our approach.

97 MATHEMATICS AND COMPUTING↗

A Kaczmarz-inspired approach to accelerate the optimization of neural network wavefunctions

Neural network wavefunctions optimized using the variational Monte Carlo method have been shown to produce highly accurate results for the electronic structure of atoms and small molecules, but the high cost of optimizing such wavefunctions prevents their application to larger systems. We propose the Subsampled Projected-Increment Natural Gradient Descent (SPRING) optimizer to reduce this bottleneck. SPRING combines ideas from the recently introduced minimum-step stochastic reconfiguration optimizer (MinSR) and the classical randomized Kaczmarz method for solving linear least-squares problems. We demonstrate that SPRING outperforms both MinSR and the popular Kronecker-Factored Approximate Curvature method (KFAC) across a number of small atoms and molecules, given that the learning rates of all methods are optimally tuned. For example, on the oxygen atom, SPRING attains chemical accuracy after forty thousand training iterations, whereas both MinSR and KFAC fail to do so even after one hundred thousand iterations.

97 MATHEMATICS AND COMPUTING↗

Insights from Initial Engineering Designs of Point Source Capture at Industrial Facilities

Initial engineering design studies examining the application of state-of-the-art carbon capture technology at industrial plants contain generally overlooked real-world design considerations for near-term deployment of point source capture (PSC). The implementation of PSC across a wide range of industrial applications presents unique challenges associated with fluctuating CO<sub>2</sub> concentrations, flue gas composition, and utility and land availability. In this article, seven recent industrial retrofit PSC projects are reviewed to investigate the impact of site-specific factors on project design and cost. Across the seven projects, three capture technology classes and four industrial applications are considered, allowing insight into industry-specific opportunities for PSC technology synergy. Common challenges across projects and proposed design solutions are highlighted to propagate ideas and solutions to close the technology gaps and accelerate learning rates.

FEED Studies↗