Search NASA⌕ Search

SEARCH · Search NASA

Results for “batch size”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Sensitivity Study of Mini-Batch Size on a Long Short-Term Memory Network for In-situ Sensing of Core-to-shell Ratio of Microencapsulated Phase Change Materials

Microencapsulated phase change materials are being studied for applications for thermal energy storage in concentrated solar fields. During fabrication, the thickness of the encapsulation cannot be readily measured for real-time control. Therefore, a machine learning network, specifically a Long Short-Term Memory network, is being developed to estimate the ratio of the shell radius to core radius based on a one second temperature history. The mini-batch size determines how often the algorithm weights are updated during network training, and shuffle indicates whether the training data is shuffled during training. A general factorial design is used to analyze the effects of varying mini-batch size and shuffle, along with the core-to-shell ratio, on the RMSE of the response from the Long Short-Term Memory network. It was found that the network performed better for smaller core to shell ratios (less than 0.6) and had the lowest RMSE when the minibatch size was 128. The minimum RMSE found was 0.00501.

Shannon, Rebecca↗

SineKAN: Kolmogorov-Arnold Networks using sinusoidal activation functions

Recent work has established an alternative to traditional multi-layer perceptron neural networks in the form of Kolmogorov-Arnold Networks (KAN). The general KAN framework uses learnable activation functions on the edges of the computational graph followed by summation on nodes. The learnable edge activation functions in the original implementation are basis spline functions (B-Spline). Here, we present a model in which learnable grids of B-Spline activation functions are replaced by grids of re-weighted sine functions (SineKAN). We evaluate numerical performance of our model on a benchmark vision task. We show that our model can perform better than or comparable to B-Spline KAN models and an alternative KAN implementation based on periodic cosine and sine functions representing a Fourier Series. Further, we show that SineKAN has numerical accuracy that could scale comparably to dense neural networks (DNNs). Compared to the two baseline KAN models, SineKAN achieves a substantial speed increase at all hidden layer sizes, batch sizes, and depths. Current advantage of DNNs due to hardware and software optimizations are discussed along with theoretical scaling. Additionally, properties of SineKAN compared to other KAN implementations and current limitations are also discussed.

Reinhardt, Eric↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Autonomous organic synthesis for redox flow batteries via flexible batch Bayesian optimization

Traditional trial-and-error methods for materials discovery are inefficient to meet the urgent demands posed by the rapid progression of climate change. This urgency has driven the increasing interest in integrating robotics and machine learning into materials research to accelerate experimental learning. However, idealized decision-making frameworks to achieve maximum sampling efficiency are not always compatible with high-throughput experimental workflows inside a laboratory. For multi-step chemical processes, differences in hardware capacities can complicate the digital framework by introducing constraints on the maximum number of samples in each step of the experiment, hence causing varying batch sizes in variable selection within the same batch. Therefore, designing flexible sampling algorithms is necessary to accommodate the multi-step synthesis with practical constraints unique to each high-throughput workflow. In this work, we designed and employed three strategies on a high-throughput robotic platform to optimize the sulfonation reaction of redox-active molecules used in flow batteries. Our strategies adapt to the multi-step experimental workflow, where their formulation and heating steps are separate, causing varying batch size requirements. By strategically sampling using clustering and mixed-variable batch Bayesian optimization, we were able to iteratively identify optimal conditions that maximize the yields. Our work presents a flexible approach that allows tailoring the machine learning decision-making to suit the practical constraints in individual high-throughput experimental platforms, followed by performing resource-efficient yield optimization using available open-source Python libraries.

Tamura, Clara [Univ. of Washington, Seattle, WA (U↗

Active Learning‐Driven Inkless Additive Nanomanufacturing for Printed Electronics

Inkless additive nanomanufacturing for printed electronics promises broad material and substrate versatility, yet the high-dimensional print parameter space makes tuning print parameters time-intensive. We present a Bayesian optimization study that constructs a digital twin from printed-silver data to benchmark surrogate models, acquisition functions, and batch sizes head-to-head to achieve user-specified target resistance. Tested surrogate models included Gaussian process, random forest, and Bayesian neural network surrogates with expected improvement and confidence bound acquisition functions. In total, we evaluate 48 unique model configurations alongside a random sampling baseline for comparison. For printed silver, the Bayesian neural network with a batch size of one achieved the lowest average cumulative regret, approximately four times more efficient on average than random sampling. To balance performance and substrate space, a random forest model with expected improvement and a batch size of four was chosen as the model for validation testing. Applying this chosen configuration to copper with an additional print parameter, the model achieved a resistance within 0.15 Ω of a 1 Ω target in fewer than 30 printed lines across five validation sets. Altogether, the workflow yields a tuned and validated model that efficiently guides experiments toward the target while simultaneously learning the parameter space.

Bevel, Colton [Auburn University, AL (United State↗

spdlayers

Symmetric Positive Definite (SPD) enforcement layers for PyTorch. Regardless of the input, the output of these layers will always be a SPD tensor! The `Cholesky` layer uses a cholesky factorization to enforce SPD, and the `Eigen` layer uses an eigendecomposition to enforce SPD. Both layers take in some tensor of shape `[batch_size, input_shape]` and output a SPD tensor of shape`[batch_size, output_shape, output_shape]`. The relationship between input and output is defined by the following. ```python input_shape = sum([i for i in range(output_shape + 1)]) ``` The layers have no learnable parameters, and merely serve to transform a vector space to a SPD matrix space.

Jekel, CharlesF.↗

EdgeAI: Machine learning via direct attached accelerator for streaming data processing at high shot rate x-ray free-electron lasers

We present a case for low batch-size inference with the potential for adaptive training of a lean encoder model. We do so in the context of a paradigmatic example of machine learning as applied in data acquisition at high data velocity scientific user facilities such as the Linac Coherent Light Source-II x-ray Free-Electron Laser. We discuss how a low-latency inference model operating at the data acquisition edge can capitalize on the naturally stochastic nature of such sources. We simulate the method of attosecond angular streaking to produce representative results whereby simulated input data reproduce high-resolution ground truth probability distributions. By minimizing the mean-squared error between the decoded output of the latent representation and the ground truth distributions, we ensure that the encoding layers and resulting latent representation maintains full fidelity for any downstream task, be it classification or regression. We present throughput results for data-parallel inference of various batch sizes, some with throughput exceeding 100 k images per second. We also show in situ training below 10 s per epoch for the full encoder–decoder model as would be relevant for streaming and adaptive real-time data production at our nation’s scientific light sources.

97 MATHEMATICS AND COMPUTING↗

Active learning for SNAP interatomic potentials via Bayesian predictive uncertainty

Bayesian inference with a simple Gaussian error model is used to efficiently compute prediction variances for energies, forces, and stresses in the linear SNAP interatomic potential. Here, the prediction variance is shown to have a strong correlation with the absolute error over approximately 24 orders of magnitude. Using this prediction variance, an active learning algorithm is constructed to iteratively train a potential by selecting the structures with the most uncertain properties from a pool of candidate structures. The relative importance of the energy, force, and stress errors in the objective function is shown to have a strong impact upon the trajectory of their respective net error metrics when running the active learning algorithm. Batched training of different batch sizes is also tested against singular structure updates, and it is found that batches can be used to significantly reduce the number of retraining steps required with only minor impact on the active learning trajectory.

97 MATHEMATICS AND COMPUTING↗

Demonstration of the Reproducibility Challenges in the Sintering Behavior of Lithium‐Stuffed Garnets in Scaling up Synthesis

Lithium-stuffed garnets, such as Li 7 La 3 Zr 2 O 12 (LLZO), are promising candidates for next-generation solid-state batteries because of their high room-temperature ionic conductivity and chemical stability against lithium metal anodes, which are crucial for achieving higher energy density. However, realizing LLZO's potential in practical devices requires synthesis methods that can be scaled reliably to large batch sizes for manufacturing. Herein, we investigate the sintering reproducibility of LLZO synthesized at larger scales using ultrasonic spray pyrolysis, a cost-effective and scalable synthesis route. Two 100 g batches of Al-doped LLZO are prepared and their sintering behavior is examined in detail. Both Al-LLZO batches contain over 90 wt.% cubic-phase LLZO, and both batches exhibit room temperature conductivities greater than 1 × 10 −4 S cm −1 at a relative density above 0.8. However, variations in secondary phases and subtle differences in Al content lead to significant differences in densification and microstructure. These results demonstrate that LLZO's sintering behavior is highly sensitive to small changes in secondary phases and Al content, creating reproducibility challenges when moving from laboratory- to manufacturing-scale synthesis.

36 MATERIALS SCIENCE↗

Finding MIDDLE Ground: Scalable and Secure Distributed Learning

Edge computing methods allow devices to efficiently train a high-performing, robust, and personalized model for predictive tasks. However, these methods succumb to privacy and scalability concerns such as adversarial data recovery and expensive model communication. Furthermore, edge computing methods unrealistically assume that all devices train an identical model. In practice, edge devices have varying computational and memory constraints which may not allow certain devices to have the space or speed to train a specific model. To overcome these issues, we propose MIDDLE: a model independent distributed learning algorithm which allows heterogeneous edge devices to assist each other’s training while communicating only non-sensitive information. MIDDLE unlocks the ability for edge devices, regardless of computational or memory constraints, to assist each other even with completely different model architectures. Furthermore, MIDDLE does not require model or gradient communication which greatly reduces communication size and time. We prove that MIDDLE attains the optimal convergence rate O(1/sqrt(TM)) of stochastic gradient descent for convex and non-convex smooth optimization (for total iterations T and batch size M). Finally, our experimental results demonstrate that MIDDLE (even in non-IID data settings) attains robust and high-performing models without model or gradient communication.

Bornstein, Marc I.↗

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong↗

Custom Equipment Development for Processing of Surplus Plutonium

The Strategic Laboratory Assessment (SLA), a collaborative team of SRNL and ORNL personnel, has been established to advance the objectives of the Surplus Plutonium Disposition (SPD) Project, by identifying and developing technologies to accelerate disposition, reduce life cycle costs, minimize worker radiation exposure, improve worker safety, and minimize Surplus Plutonium Disposition Program risks. [1] The SLA team has identified can cutting and plutonium (Pu) oxide size reduction as two glovebox processes where technology enhancements would be valuable. The DOESTD-3013 package currently in use for Pu downblending requires cutting two nested cans before the inner convenience can that holds the Pu oxide may be accessed for further processing. A rotary tubing-style cutter is used for opening the 3013 packages within the glovebox. Collet changeouts are required between cutting of the outer and inner cans. The SLA team is currently developing and testing an adjustable-clamp can cutter design that eliminates collet changeouts and allows cutting of the outer and inner can at the same time, resulting in significant reduction of radiological dose and process time, as well as improved ergonomics. To meet the Pu oxide particle size requirement, size reduction of Pu oxide agglomerations must be performed within the process gloveboxes. The SLA team has identified jaw crushing technology as an alternative to the currently employed rotary mill. Jaw crusher advantages include reduced dust within the glovebox, increased batch sizes, and easier integration with other glovebox processes due to the flow-through nature of jaw crushing. Commercially manufactured jaw crushers are either too large and/or too heavy for implementation in the SPD gloveboxes, so the SLA team is developing and testing a custom jaw crusher to meet the needs of the SPD Project.

Krementz, Daniel [Savannah River National Laborato↗

An Accelerated Clip Algorithm for Unstructured Meshes: A Batch-Driven Approach

The clip technique is a popular method for visualizing complex structures and phenomena within 3D unstructured meshes. Meshes can be clipped by specifying a scalar isovalue to produce an output unstructured mesh with its external surface as the isovalue. Similar to isocontouring, the clipping process relies on scalar data associated with the mesh points, including scalar data generated by implicit functions such as planes, boxes, and spheres, which facilitates the visualization of results interior to the grid. In this paper, we introduce a novel batch-driven parallel algorithm based on a sequential clip algorithm designed for high-quality results in partial volume extraction. Our algorithm comprises five passes, each progressively processing data to generate the resulting clipped unstructured mesh. The novelty lies in the use of fixed-size batches of points and cells, which enable rapid workload trimming and parallel processing, leading to a significantly improved memory footprint and run-time performance compared to the original version. On a 32-core CPU, the proposed batch-driven parallel algorithm demonstrates a run-time speed-up of up to 32.6x and a memory footprint reduction of up to 4.37x compared to the existing sequential algorithm. The software is currently available under an open-source license in the VTK visualization system.

Tsalikis, Spiros↗

Latency Analysis of the Nexus Digital Twin Framework

Real-time digital catalogs are increasingly relied upon to track metadata and connect disparate data sources for cloud-based data integration efforts. One such tool, Deeplynx Nexus is supporting real-time digital twin efforts through event-driven data integration and time-series queries. Nexus’s usefulness for these applications depends critically on how quickly individual records can be uploaded and downloaded, since delays directly affect the responsiveness of any system built on top of it. However, the actual latency a user should expect from Nexus has not been systematically measured before, particularly for the small, frequent transactions typical of live sensor feeds. Here we show that single-record round-trip latency is 61.1 ms on a local Nexus instance and 391.7 ms on the hosted production infrastructure, a roughly 6.4x difference driven primarily by fixed per-request overhead rather than data volume. This overhead dominates at small scale: comparing single-record and ten-record trials suggests approximately 56 ms of each single-record request is fixed connection and authentication cost rather than data-transfer time, meaning batching even a handful of records is substantially more efficient than transmitting them individually. At large batch sizes, this pattern reverses for uploads, which converge to near parity between local and hosted environments by 25,000-50,000 records, while download latency remains persistently 5.7-6.4x slower on hosted infrastructure even at scale. These results suggest that Nexus deployments intended for real-time digital twin applications should prioritize record batching over single-record transactions, and that download-path optimization on hosted infrastructure offers the largest remaining opportunity to reduce latency at scale. We anticipate these baseline measurements will serve as a reference point for future digital twin projects evaluating whether Nexus’s latency profile meets their real-time requirements, and as a benchmark for tracking the effect of future infrastructure or API changes.

99 - GENERAL AND MISCELLANEOUS↗

ZeoNet: 3D convolutional neural networks for predicting adsorption in nanoporous zeolites

Zeolites are one of the most widely used materials in the chemical industry due to their nanometer-sized pores that can adsorb and react upon molecules selectively. With hundreds of known framework topologies and hundreds of thousands of computationally predicted structures, the ability to rapidly predict zeolite performance allows researchers to prioritize their efforts on the most promising structures for a given application. Although the accuracy of forcefield-based atomistic simulations has advanced significantly in the past two decades, these simulations can be computationally expensive, especially for long-chain, complex molecules. Here, we present ZeoNet, a representation learning framework using convolutional neural networks (ConvNets) and 3D volumetric representations for predicting adsorption in zeolites. ZeoNet was trained on the task of predicting Henry's constants for adsorption, k H , of n-octadecane in more than 330 000 known and predicted zeolite materials. Employing a 3D grid based on the distances to solvent-accessible surfaces, a volumetric representation that can be generated efficiently, the best-performing ZeoNet achieved a correlation coefficient r 2 = 0.977 and a mean-squared error MSE = 3.8 in ln k H , which corresponds to an error of 9.3 kJ mol -1 in adsorption free energy. In comparison, a model based on hand-designed geometric features has values of r 2 = 0.783 and MSE = 35.7. ZeoNet is also relatively efficient and can process ≈8 structures per second on an Nvidia RTX 2080TI GPU, orders of magnitude faster than forcefield-based simulations. A systematic analysis was conducted to investigate how the choice of ConvNet architectures, the linear dimension (L) and spatial resolution (Δd) of the distance grids, batch size, optimizer, and learning rate impact the model performance. We found that ConvNets based on the ResNet architecture offer the best tradeoff between expressiveness and efficiency. The performance for all models reaches a plateau at L = 30–45 Å and depends less sensitively on grid resolution, with a small benefit around Δd = 0.30–0.45 Å. Finally, saliency maps were visualized to identify which regions of the materials contributed the most to model predictions. It was found, interestingly, that the predictions are driven primarily by the accessible pore volume rather than the region occupied by the framework atoms.

36 MATERIALS SCIENCE↗

Universal method for the optimization of HDC coating uniformity on non-planar, non-stationary substrates for inertial confinement fusion targets

The thickness uniformity of chemical vapor deposited (CVD) diamond coatings on non-planar, non-stationary substrates depends on both the intrinsic instantaneous coating thickness distribution (ICTD) of the coating conditions used and, if applicable, on the frequency of substrate reorientation. While important for many CVD diamond applications, the relative impact of the ICTD and substrate reorientation on the coating thickness uniformity has not been studied. In this work, we systematically investigate the effect of these factors for microwave-plasma chemical vapor deposition (MPCVD) of diamond (referred to as high density carbon (HDC) in the inertial confinement fusion (ICF) community) coatings on spherical, rolling substrates. This coating technique is used to fabricate capsules for ICF experiments, which require extreme coating uniformity with <0.3 % thickness variation (so-called Mode 1 or M1) to ensure symmetric compression of imploding targets. To extract the otherwise unobservable reorientation timescale (Δt), Monte Carlo simulations were performed using experimental ICTD data as input. This combined approach confirms scaling relationships between the substrate reorientation timescale as well as coating thickness and coating uniformity, as expected from a 3D random walk. Simulations confirm that M1 is Rayleigh-distributed and scales as (Δt) 1/2 , consistent with the randomization of two angles that determine orientation of a sphere. We also demonstrate that, under the conditions studied, Δt is the dominant factor in determining thickness uniformity while the intrinsic ICTD has minimal impact. Finally, experiments show that Δt can be affected by total batch size under constant agitation conditions due to space constraints that limit the capsule reorientation kinetics. In conclusion, this study highlights the utility of a combined experiment-simulation approach as a general methodology for understanding and improving coating uniformity on non-planar, non-stationary substrates.

Capsule↗

A strategy for automated core design to increase economic viability and minimize fuel fragmentation, relocation, and dispersal susceptibility in high-burnup cores

The nuclear industry aims to increase the cycle length of pressurized water reactors from 18 to 24 months to increase power plant capacity factors and economic viability. These cycle length extensions will inherently require fuel rods to exceed the current peak rod average burnup limit of 62 GWd/MTU. A chief concern of operating beyond the current burnup limit is the fuel fragmentation, relocation, and dispersal (FFRD) phenomenon in which pulverized fuel fragments can axially relocate and escape through a burst in the cladding formed during a loss-of-coolant accident. In this work, we demonstrate an approach for automating core design employing an optimization tool based on a penalty-free, parallel simulated annealing algorithm to produce pressurized water reactor core designs with two different optimization objectives. The two objectives were to produce core designs with (1) mitigated FFRD susceptibility while achieving 24-month cycle lengths (2) maximum cycle length with no regard for the likelihood of FFRD. Batch size was considered in tandem with both cases to maximize economic viability. The PARCS nodal model was the primary reactor physics tool used in the optimizations and used nuclear cross sections calculated with 2D Polaris lattice physics models. Reactor performance and safety characteristics of the optimized cores were verified using high-fidelity Virtual Environment for Reactor Applications models. The core designs produced by the optimization tool are compared with each other and to a high-burnup core design produced and analyzed in previous works to highlight the fuel management strategies that may enhance high-burnup reactor safety and economic viability. The optimized cores satisfied their respective objective functions, producing a maximum cycle length of 720 effective full-power days in one core design and one that may reduce FFRD susceptibility by up to 50% based on the first-order approximation to FFRD risk formulated in this work. The optimized cores met most constraints but exceeded the hot channel factor limit, especially in FFRD cases where fresh fuel carried more power. Furthermore, this highlights the need for future lattice-level optimizations and broader assembly options.

Cycle length↗