Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelize algorithm computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59

DIMPLES: Distributed Influence Maximization for Pandemic pLanning on Exascale Systems

We study exascale parallel algorithms for the selection of intervention or monitoring strategies in massive realistic socio-technical networks through scalable Influence Maximization (InfMax) algorithms. We employ novel techniques to enable efficient scaling on up to 8k nodes of OLCF Frontier, with 65k AMD GPUs and 458k AMD CPU cores. Current state-of-the-art InfMax tools are limited to networks with only a few million actors (vertices) and a few hundred million interactions (edges). By overcoming these limitations, we show that our approach is capable of processing a realistic social contact network of the United States with 285 million nodes and about 8 billion edges. This two orders-of-magnitude improvement over the previous state-of-the-art is obtained by leveraging algorithmic advancements for the InfMax problem and designing several problem-specific approaches to overlap communication with computation, improve GPU efficiency, and lower the application’s memory requirements. We evaluate strong scaling for computing 10k most influential seeds using up to 8k nodes of an exascale system, and weak scaling from 128 to 8k system nodes for seed sets ranging from 625 to 40k seeds. We achieve the fastest-known runtime of 25 minutes while performing 48 million diffusion simulations totaling 2.31 petabytes to identify 40k influential seeds using 8k nodes, and take 5.75 minutes to identify 10k seeds while using 4k nodes.

Minutoli, Marco [Pacific Northwest National Labora↗

Task scheduling in dataflow computer architectures

Dataflow computers provide a platform for the solution of a large class of computational problems, which includes digital signal processing and image processing. Many typical applications are represented by a set of tasks which can be repetitively executed in parallel as specified by an associated dataflow graph. Research in this area aims to model these architectures, develop scheduling procedures, and predict the transient and steady state performance. Researchers at NASA have created a model and developed associated software tools which are capable of analyzing a dataflow graph and predicting its runtime performance under various resource and timing constraints. These models and tools were extended and used in this work. Experiments using these tools revealed certain properties of such graphs that require further study. Specifically, the transient behavior at the beginning of the execution of a graph can have a significant effect on the steady state performance. Transformation and retiming of the application algorithm and its initial conditions can produce a different transient behavior and consequently different steady state performance. The effect of such transformations on the resource requirements or under resource constraints requires extensive study. Task scheduling to obtain maximum performance (based on user-defined criteria), or to satisfy a set of resource constraints, can also be significantly affected by a transformation of the application algorithm. Since task scheduling is performed by heuristic algorithms, further research is needed to determine if new scheduling heuristics can be developed that can exploit such transformations. This work has provided the initial development for further long-term research efforts. A simulation tool was completed to provide insight into the transient and steady state execution of a dataflow graph. A set of scheduling algorithms was completed which can operate in conjunction with the modeling and performance tools previously developed. Initial studies on the performance of these algorithms were done to examine the effects of application algorithm transformations as measured by such quantities as number of processors, time between outputs, time between input and output, communication time, and memory size.

Katsinis, Constantine↗

High Resolution Aerospace Applications using the NASA Columbia Supercomputer

This paper focuses on the parallel performance of two high-performance aerodynamic simulation packages on the newly installed NASA Columbia supercomputer. These packages include both a high-fidelity, unstructured, Reynolds-averaged Navier-Stokes solver, and a fully-automated inviscid flow package for cut-cell Cartesian grids. The complementary combination of these two simulation codes enables high-fidelity characterization of aerospace vehicle design performance over the entire flight envelope through extensive parametric analysis and detailed simulation of critical regions of the flight envelope. Both packages. are industrial-level codes designed for complex geometry and incorpor.ats. CuStomized multigrid solution algorithms. The performance of these codes on Columbia is examined using both MPI and OpenMP and using both the NUMAlink and InfiniBand interconnect fabrics. Numerical results demonstrate good scalability on up to 2016 CPUs using the NUMAIink4 interconnect, with measured computational rates in the vicinity of 3 TFLOP/s, while InfiniBand showed some performance degradation at high CPU counts, particularly with multigrid. Nonetheless, the results are encouraging enough to indicate that larger test cases using combined MPI/OpenMP communication should scale well on even more processors.

Mavriplis, Dimitri J.↗

Design and Implementation of a Parallel Multivariate Ensemble Kalman Filter for the Poseidon Ocean General Circulation Model

A multivariate ensemble Kalman filter (MvEnKF) implemented on a massively parallel computer architecture has been implemented for the Poseidon ocean circulation model and tested with a Pacific Basin model configuration. There are about two million prognostic state-vector variables. Parallelism for the data assimilation step is achieved by regionalization of the background-error covariances that are calculated from the phase-space distribution of the ensemble. Each processing element (PE) collects elements of a matrix measurement functional from nearby PEs. To avoid the introduction of spurious long-range covariances associated with finite ensemble sizes, the background-error covariances are given compact support by means of a Hadamard (element by element) product with a three-dimensional canonical correlation function. The methodology and the MvEnKF configuration are discussed. It is shown that the regionalization of the background covariances; has a negligible impact on the quality of the analyses. The parallel algorithm is very efficient for large numbers of observations but does not scale well beyond 100 PEs at the current model resolution. On a platform with distributed memory, memory rather than speed is the limiting factor.

Keppenne, Christian L.↗

A New Objective Technique for Verifying Mesoscale Numerical Weather Prediction Models

This report presents a new objective technique to verify predictions of the sea-breeze phenomenon over east-central Florida by the Regional Atmospheric Modeling System (RAMS) mesoscale numerical weather prediction (NWP) model. The Contour Error Map (CEM) technique identifies sea-breeze transition times in objectively-analyzed grids of observed and forecast wind, verifies the forecast sea-breeze transition times against the observed times, and computes the mean post-sea breeze wind direction and speed to compare the observed and forecast winds behind the sea-breeze front. The CEM technique is superior to traditional objective verification techniques and previously-used subjective verification methodologies because: It is automated, requiring little manual intervention, It accounts for both spatial and temporal scales and variations, It accurately identifies and verifies the sea-breeze transition times, and It provides verification contour maps and simple statistical parameters for easy interpretation. The CEM uses a parallel lowpass boxcar filter and a high-order bandpass filter to identify the sea-breeze transition times in the observed and model grid points. Once the transition times are identified, CEM fits a Gaussian histogram function to the actual histogram of transition time differences between the model and observations. The fitted parameters of the Gaussian function subsequently explain the timing bias and variance of the timing differences across the valid comparison domain. Once the transition times are all identified at each grid point, the CEM computes the mean wind direction and speed during the remainder of the day for all times and grid points after the sea-breeze transition time. The CEM technique performed quite well when compared to independent meteorological assessments of the sea-breeze transition times and results from a previously published subjective evaluation. The algorithm correctly identified a forecast or observed sea-breeze occurrence or absence 93% of the time during the two- month evaluation period from July and August 2000. Nearly all failures in CEM were the result of complex precipitation features (observed or forecast) that contaminated the wind field, resulting in a false identification of a sea-breeze transition. A qualitative comparison between the CEM timing errors and the subjectively determined observed and forecast transition times indicate that the algorithm performed very well overall. Most discrepancies between the CEM results and the subjective analysis were again caused by observed or forecast areas of precipitation that led to complex wind patterns. The CEM also failed on a day when the observed sea- breeze transition affected only a very small portion of the verification domain. Based on the results of CEM, the RAMS tended to predict the onset and movement of the sea-breeze transition too early and/or quickly. The domain-wide timing biases provided by CEM indicated an early bias on 30 out of 37 days when both an observed and forecast sea breeze occurred over the same portions of the analysis domain. These results are consistent with previous subjective verifications of the RAMS sea breeze predictions. A comparison of the mean post-sea breeze winds indicate that RAMS has a positive wind-speed bias for .all days, which is also consistent with the early bias in the sea-breeze transition time since the higher wind speeds resulted in a faster inland penetration of the sea breeze compared to reality.

Case, Jonathan L.↗

A Survey of New Trends in Symbolic Execution for Software Testing and Analysis

Symbolic execution is a well-known program analysis technique which represents values of program inputs with symbolic values instead of concrete (initialized) data and executes the program by manipulating program expressions involving the symbolic values. Symbolic execution has been proposed over three decades ago but recently it has found renewed interest in the research community, due in part to the progress in decision procedures, availability of powerful computers and new algorithmic developments. We provide a survey of some of the new research trends in symbolic execution, with particular emphasis on applications to test generation and program analysis. We first describe an approach that handles complex programming constructs such as input data structures, arrays, as well as multi-threading. We follow with a discussion of abstraction techniques that can be used to limit the (possibly infinite) number of symbolic configurations that need to be analyzed for the symbolic execution of looping programs. Furthermore, we describe recent hybrid techniques that combine concrete and symbolic execution to overcome some of the inherent limitations of symbolic execution, such as handling native code or availability of decision procedures for the application domain. Finally, we give a short survey of interesting new applications, such as predictive testing, invariant inference, program repair, analysis of parallel numerical programs and differential symbolic execution.

Pasareanu, Corina S.↗

Knowledge-based vision for space station object motion detection, recognition, and tracking

Computer vision, especially color image analysis and understanding, has much to offer in the area of the automation of Space Station tasks such as construction, satellite servicing, rendezvous and proximity operations, inspection, experiment monitoring, data management and training. Knowledge-based techniques improve the performance of vision algorithms for unstructured environments because of their ability to deal with imprecise a priori information or inaccurately estimated feature data and still produce useful results. Conventional techniques using statistical and purely model-based approaches lack flexibility in dealing with the variabilities anticipated in the unstructured viewing environment of space. Algorithms developed under NASA sponsorship for Space Station applications to demonstrate the value of a hypothesized architecture for a Video Image Processor (VIP) are presented. Approaches to the enhancement of the performance of these algorithms with knowledge-based techniques and the potential for deployment of highly-parallel multi-processor systems for these algorithms are discussed.

Symosek, P.↗

High Rate Digital Demodulator ASIC

The architecture of High Rate (600 Mega-bits per second) Digital Demodulator (HRDD) ASIC capable of demodulating BPSK and QPSK modulated data is presented in this paper. The advantages of all-digital processing include increased flexibility and reliability with reduced reproduction costs. Conventional serial digital processing would require high processing rates necessitating a hardware implementation in other than CMOS technology such as Gallium Arsenide (GaAs) which has high cost and power requirements. It is more desirable to use CMOS technology with its lower power requirements and higher gate density. However, digital demodulation of high data rates in CMOS requires parallel algorithms to process the sampled data at a rate lower than the data rate. The parallel processing algorithms described here were developed jointly by NASA's Goddard Space Flight Center (GSFC) and the Jet Propulsion Laboratory (JPL). The resulting all-digital receiver has the capability to demodulate BPSK, QPSK, OQPSK, and DQPSK at data rates in excess of 300 Mega-bits per second (Mbps) per channel. This paper will provide an overview of the parallel architecture and features of the HRDR ASIC. In addition, this paper will provide an over-view of the implementation of the hardware architectures used to create flexibility over conventional high rate analog or hybrid receivers. This flexibility includes a wide range of data rates, modulation schemes, and operating environments. In conclusion it will be shown how this high rate digital demodulator can be used with an off-the-shelf A/D and a flexible analog front end, both of which are numerically computer controlled, to produce a very flexible, low cost high rate digital receiver.

Ghuman, Parminder↗

Multinode reconfigurable pipeline computer

A multinode parallel-processing computer is made up of a plurality of innerconnected, large capacity nodes each including a reconfigurable pipeline of functional units such as Integer Arithmetic Logic Processors, Floating Point Arithmetic Processors, Special Purpose Processors, etc. The reconfigurable pipeline of each node is connected to a multiplane memory by a Memory-ALU switch NETwork (MASNET). The reconfigurable pipeline includes three (3) basic substructures formed from functional units which have been found to be sufficient to perform the bulk of all calculations. The MASNET controls the flow of signals from the memory planes to the reconfigurable pipeline and vice versa. the nodes are connectable together by an internode data router (hyperspace router) so as to form a hypercube configuration. The capability of the nodes to conditionally configure the pipeline at each tick of the clock, without requiring a pipeline flush, permits many powerful algorithms to be implemented directly.

Nosenchuck, Daniel M.↗

On the spectral stability of time integration algorithms for a class of constrained dynamics problems

Incomplete field formulations have recently been the subject of intense research because of their potential in coupled analysis of independently modeled substructures, adaptive refinement, domain decomposition, and parallel processing. This paper discusses the design and analysis of time-integration algorithms for these formulations and emphasizes the treatment of their inter-subdomain constraint equations. These constraints are shown to introduce a destabilizing effect in the dynamic system that can be analyzed by investigating the behavior of the time-integration algorithm at infinite and zero frequencies. Three different approaches for constructing penalty-free unconditionally stable second-order accurate solution procedures for this class of hybrid formulations are presented, discussed and illustrated with numerical examples. The theoretical results presented in this paper also apply to a large family of nonlinear multibody dynamics formulations. Some of the algorithms outlined herein are important alternatives to the popular technique consisting of transforming differential/algebraic equations into ordinary differential equations via the introduction of a stabilization term that depends on arbitrary constants and that influences the computed so1ution.

Farhat, Charbel↗

Experimental Investigation of Pool Boiling Heat Transfer Enhancement in Microgravity in the Presence of Electric Fields

In boiling high heat fluxes are possible driven by relatively small temperature differences, which make its use increasingly attractive in aerospace applications. The objective of the research is to develop ways to overcome specific problems associated with boiling in the low gravity environment by substituting the buoyancy force with the electric force to enhance bubble removal from the heated surface. Previous studies indicate that in terrestrial applications nucleate boiling heat transfer can be increased by a factor of 50, as compared to values obtained for the same system without electric fields. The goal of our research is to experimentally explore the mechanisms responsible for EHD heat transfer enhancement in boiling in low gravity conditions, by visualizing the temperature distributions in the vicinity of the heated surface and around the bubble during boiling using real-time holographic interferometry (HI) combined with high-speed cinematography. In the first phase of the project the influence of the electric field on a single bubble is investigated. Pool boiling is simulated by injecting a single bubble through a nozzle into the subcooled liquid or into the thermal boundary layer developed along the flat heater surface. Since the exact location of bubble formation is known, the optical equipment can be aligned and focused accurately, which is an essential requirement for precision measurements of bubble shape, size and deformation, as well as the visualization of temperature fields by HI. The size of the bubble and the frequency of bubble departure can be controlled by suitable selection of nozzle diameter and mass flow rate of vapor. In this approach effects due to the presence of the electric field can be separated from effects caused by the temperature gradients in the thermal boundary layer. The influence of the thermal boundary layer can be investigated after activating the heater at a later stage of the research. For the visualization experiments a test cell was developed. All four vertical walls of the test cell are transparent, and they allow transillumination with laser light for visualization experiments by HI. The bottom electrode is a copper cylinder, which is electrically grounded. The copper block is heated with a resistive heater and it is equipped with 6 thermocouples that provide reference temperatures for the measurements with HI. The top electrode is a mesh electrode. Bubbles are injected with a syringe into the test cell through the bottom electrode. The working fluids presently used in the interferometric visualization experiments, water and PF 5052, satisfy requirements regarding thermophysical, optical and electrical properties. A 30kV power supply equipped with a voltmeter allows to apply the electric field to the electrodes during the experiments. The magnitude of the applied voltage can be adjusted either manually or through the LabVIEW data acquisition and control system connected to a PC. Temperatures of the heated block are recorded using type-T thermocouples, whose output is read by a data acquisition system. Images of the bubbles are recorded with 35mm photographic and 16mm high-speed cameras, scanned and analyzed using various software packages. Visualized temperature fields HI allows the visualization of temperature fields in the vicinity of bubbles during boiling in the form of fringes. Typical visualized temperature distributions around the air bubbles injected into the thermal boundary layer in PF5052 are shown. The temperature of the heated surface is 35 C. The temperature difference for a pair of fringes is approximately 0.05 C. The heat flux applied to the bottom surface is moderate, and the fringe patterns are regular. In the image a bubble penetrating the thermal boundary layer is visible. Because of the axial symmetry of the problem, simplified reconstruction techniques can be applied to recover the temperature field. The thermal plume developing above the heated surface for more intensive heating is shown. The temperature distribution in the liquid is clearly 3D, and tomographic techniques have to be applied to recover the temperature distribution in such a physical situation. A sequence of interferometric images showing the temperature distribution around the rising bubble, recorded with a high-speed camera is shown. Again, the temperature distribution is 3D, and a more complex approach to the evaluation, the tomographic reconstruction has to be taken. Measurement of the temperature distribution from the fringe pattern temperature distributions that yield important information regarding heat transfer are determined. Two algorithms that allow the quantitative evaluation of interferometric fringe patterns and the reconstruction of temperature fields during boiling have been developed at the Heat Transfer Laboratory of the Johns Hopkins University. In the first algorithm the bubble is assumed to be axially symmetrical, which significantly reduces the computational effort for quantifying the temperature distribution around the bubble. For this purpose the thermal boundary layer around the bubble is divided into equidistant concentric shells, and the refractive index is assumed to be constant in each of the shells. Since large temperature gradients are expected in the vicinity of the bubble during boiling, the deflection of the light beam cannot be neglected in boiling experiments. Since the exit angle of the light beam is known, this allows to account for the deflections and phase shifts outside the boundary layer (in the bulk fluid and in the windows of the test cell). Three dimensional temperature distributions in the vicinity of the bubble are reconstructed using tomographic techniques. In tomography, the measurement volume is sliced into 2D planes. In the present study these planes are parallel to the heated surface. The objective is to determine the values of the field parameter of interest in form of the field function in these 2D planes. The field parameter is the change of the refractive index of the liquid in the measurement volume caused by temperature changes. By superimposing data for many 2D planes recorded at the same time instant, the 3D temperature distribution in the measurement volume is recovered.

Herman, Cila↗

An Approach to Solving Enclosure Radiation Problems in A Multi-Physics Context

Thermal protection system analysis of complex features or damage sites can sometimes require modeling of high temperature enclosures. Implementing efficient and accurate view-factor algorithms required to model such problems is complex. The current work leverages the Non-equilibrium Radiation (NERO) software, which solves the radiation transport equation in a finite-volume scheme, to alleviating challenges often faced with view-factor calculations. By assuming heat transfer occurs only between grey bodies and that the medium is non-participating, computational cost of the method is significantly reduced. The enclosure physics are modeled through emitting and reflecting boundary conditions in NERO. The emitted radiative flux is dependent on the wall temperature which is a solution to the material response, obtained from Icarus, in this context. The Ares framework manages the time-advancement and exchange of the necessary data between the solvers. The surface energy balance is modified to account for the enclosure terms within the material response boundary condition. The methodology was verified against analytical solutions including radiating parallel plates, a hollow cylinder (shown in Fig. 1), and a hemisphere. Application of the methodology to inform the design of components of the Dragonfly system will be shown.

Ablation↗

Parallel-vector computation for linear structural analysis and non-linear unconstrained optimization problems

Several parallel-vector computational improvements to the unconstrained optimization procedure are described which speed up the structural analysis-synthesis process. A fast parallel-vector Choleski-based equation solver, pvsolve, is incorporated into the well-known SAP-4 general-purpose finite-element code. The new code, denoted PV-SAP, is tested for static structural analysis. Initial results on a four processor CRAY 2 show that using pvsolve reduces the equation solution time by a factor of 14-16 over the original SAP-4 code. In addition, parallel-vector procedures for the Golden Block Search technique and the BFGS method are developed and tested for nonlinear unconstrained optimization. A parallel version of an iterative solver and the pvsolve direct solver are incorporated into the BFGS method. Preliminary results on nonlinear unconstrained optimization test problems, using pvsolve in the analysis, show excellent parallel-vector performance indicating that these parallel-vector algorithms can be used in a new generation of finite-element based structural design/analysis-synthesis codes.

Nguyen, D. T.↗

The Use of Dynamic Visual Acuity as a Functional Test of Gaze Stabilization Following Space Flight

After prolonged exposure to a given gravitational environment the transition to another is accompanied by adaptations in the sensorimotor subsystems, including the vestibular system. Variation in the adaptation time course of these subsystems, and the functional redundancies that exist between them make it difficult to accurately assess the functional capacity and physical limitations of astro/cosmonauts using tests on individual subsystems. While isolated tests of subsystem performance may be the only means to address where interventions are required, direct measures of performance may be more suitable for assessing the operational consequences of incomplete adaptation to changes in the gravitational environment. A test of dynamic visual acuity (DVA) is currently being used in the JSC Neurosciences Laboratory as part of a series of measures to assess the efficacy of a countermeasure to mitigate postflight locomotor dysfunction. In the current protocol, subjects visual acuity is determined using Landolt ring optotypes presented sequentially on a computer display. Visual acuity assessments are made both while standing and while walking at 1.8 m/s on a motorized treadmill. The use of a psychophysical threshold detection algorithm reduces the required number of optotype presentations and the results can be presented immediately after the test. The difference between the walking and standing acuity measures provides a metric of the change in the subject s ability to maintain gaze fixation on the visual target while walking. This functional consequence is observable regardless of the underlying subsystem most responsible for the change. Data from 15 cosmo/astronauts have been collected following long-duration (approx. 6 months) stays in space using a visual target viewing distance of 4.0 meters. An investigation of the group mean shows a change in DVA soon after the flight that asymptotes back to baseline approximately one week following their return to earth. The performance of some subjects nicely parallels the stereotypical recovery curve observed in the group mean data. Others show dramatic changes in DVA from one test day to another. These changes may be indicative of a re-adaptation process that is not characterized by a steady improvement with the passage of time, but is instead a dynamic search for appropriate coordinative strategy to achieve the desired gaze stabilization goal. Ground-based data have been collected in our lab using DVA with one of the goals being to improve the DVA test itself. In one of these studies, the DVA test was repeated using a visual target viewing distance of 0.5 meters. While walking, the relative contributions of the otoliths and semi-circular canals that are required to stabilize gaze are affected by visual target viewing distance. It may be possible to exploit this using the current treadmill DVA test to differentially assess changes in these vestibular subsystems. The postflight DVA evaluations currently used have been augmented to include the near target version of the test. Preliminary results from these assessments, as well as the results from the ground-based tests will also be reported. DVA provides a direct measure of a subject's ability to see clearly in the presence of self-motion. The use of the current tests for providing a functionally relevant metric is evident. However, it is possible to expand the scope of DVA testing to include scenarios other than walking. A facility for measuring DVA in the presence of passive movements is being created. Using a mechanized platform to provide the perturbation, it should be possible to simulate aircraft and automobile vibration profiles. Used in conjunction with the far and near visual displays this facility should be able to assess a subject s ability to clearly see distant objects as well as those that appear on the dashboard or instrument control panel during functionally relevant situations.

Peters, B. T.↗

Rolling Horizon with K-Position Search Method for Strategic Deconfliction of Package Delivery UAS

This research focuses on the strategic deconfliction of unmanned aircraft systems (UAS) in an urban package delivery environment with two depots and multiple drop-off locations. Since the formulated mixed-integer nonlinear programming (MINLP) problem is non-deterministic polynomial-time (NP) hard, a heuristic algorithm called "rolling horizon with k-position search (KPS)" is used to compute the departure sequence and scheduled time of departure (STD) of each UAS at a depot, considering temporal constraints at en-route crossing waypoints and depots for strategic deconfliction. The simulation studies show that an increase in the value of k (local neighborhood search) in the KPS reduces the average ground delay at the cost of an increase in the computation time for a given number of UAS, size of the rolling horizon window, and number of depots involved in the local neighborhood search. The studies also show that for a given rolling horizon window, the computation time increases exponentially with an increase in the total number of UAS flights when serial processing the local neighborhood search of KPS (with k > 1) and drops by an order of magnitude upon performing the local neighborhood search of KPS using parallel processing instead of serial processing. The computation time drops with the reduction in air traffic complexity of a scenario for a given number of flights, k (local neighborhood search), and rolling horizon window.

UTM↗

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization↗

Communication Studies of DMP and SMP Machines

Understanding the interplay between machines and problems is key to obtaining high performance on parallel machines. This paper investigates the interplay between programming paradigms and communication capabilities of parallel machines. In particular, we explicate the communication capabilities of the IBM SP-2 distributed-memory multiprocessor and the SGI PowerCHALLENGEarray symmetric multiprocessor. Two benchmark problems of bitonic sorting and Fast Fourier Transform are selected for experiments. Communication-efficient algorithms are developed to exploit the overlapping capabilities of the machines. Programs are written in Message-Passing Interface for portability and identical codes are used for both machines. Various data sizes and message sizes are used to test the machines' communication capabilities. Experimental results indicate that the communication performance of the multiprocessors are consistent with the size of messages. The SP-2 is sensitive to message size but yields a much higher communication overlapping because of the communication co-processor. The PowerCHALLENGEarray is not highly sensitive to message size and yields a low communication overlapping. Bitonic sorting yields lower performance compared to FFT due to a smaller computation-to-communication ratio.

Sohn, Andrew↗

Fiber Bragg Grating Sensor System for Monitoring Smart Composite Aerospace Structures

Lightweight, electromagnetic interference (EMI) immune, fiber-optic, sensor- based structural health monitoring (SHM) will play an increasing role in aerospace structures ranging from aircraft wings to jet engine vanes. Fiber Bragg Grating (FBG) sensors for SHM include advanced signal processing, system and damage identification, and location and quantification algorithms. Potentially, the solution could be developed into an autonomous onboard system to inspect and perform non-destructive evaluation and SHM. A novel method has been developed to massively multiplex FBG sensors, supported by a parallel processing interrogator, which enables high sampling rates combined with highly distributed sensing (up to 96 sensors per system). The interrogation system comprises several subsystems. A broadband optical source subsystem (BOSS) and routing and interface module (RIM) send light from the interrogation system to a composite embedded FBG sensor matrix, which returns measurand-dependent wavelengths back to the interrogation system for measurement with subpicometer resolution. In particular, the returned wavelengths are channeled by the RIM to a photonic signal processing subsystem based on powerful optical chips, then passed through an optoelectronic interface to an analog post-detection electronics subsystem, digital post-detection electronics subsystem, and finally via a data interface to a computer. A range of composite structures has been fabricated with FBGs embedded. Stress tensile, bending, and dynamic strain tests were performed. The experimental work proved that the FBG sensors have a good level of accuracy in measuring the static response of the tested composite coupons (down to submicrostrain levels), the capability to detect and monitor dynamic loads, and the ability to detect defects in composites by a variety of methods including monitoring the decay time under different dynamic loading conditions. In addition to quasi-static and dynamic load monitoring, the system can capture acoustic emission events that can be a prelude to structural failure, as well as piezoactuator-induced ultrasonic Lamb-waves-based techniques as a basis for damage detection.

Moslehi, Behzad↗