Search NASA⌕ Search

SEARCH · Search NASA

Results for “Error Mitigation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Exploring the Utility of Machine Learning-Based Passive Microwave Brightness Temperature Data Assimilation over Terrestrial Snow in High Mountain Asia

This study explores the use of a support vector machine (SVM) as the observation operator within a passive microwave brightness temperature data assimilation framework (herein SVM-DA) to enhance the characterization of snow water equivalent (SWE) over High Mountain Asia (HMA). A series of synthetic twin experiments were conducted with the NASA Land Information System (LIS) at a number of locations across HMA. Overall, the SVM-DA framework is effective at improving SWE estimates (~70% reduction in RMSE relative to the Open Loop) for SWE depths less than 200 mm during dry snowpack conditions. The SVM-DA framework also improves SWE estimates in deep, wet snow (~45% reduction in RMSE) when snow liquid water is well estimated by the land surface model, but can lead to model degradation when snow liquid water estimates diverge from values used during SVM training. In particular, two key challenges of using the SVM-DA framework were observed over deep, wet snowpacks. First, variations in snow liquid water content dominate the brightness temperature spectral difference (TB) signal associated with emission from a wet snowpack, which can lead to abrupt changes in SWE during the analysis update. Second, the ensemble of SVM-based predictions can collapse (i.e., yield a near-zero standard deviation across the ensemble) when prior estimates of snow are outside the range of snow inputs used during the SVM training procedure. Such a scenario can lead to the presence of spurious error correlations between SWE and TB, and as a consequence, can result in degraded SWE estimates from the analysis update. These degraded analysis updates can be largely mitigated by applying rule-based approaches. For example, restricting the SWE update when the standard deviation of the predicted TB is greater than 0.05 K helps prevent the occurrence of filter divergence. Similarly, adding a thin layer (i.e., 5 mm) of SWE when the synthetic TB is larger than 5 K can improve SVM-DA performance in the presence of a precipitation dry bias. The study demonstrates that a carefully constructed SVM-DA framework cognizant of the inherent limitations of passive microwave-based SWE estimation holds promise for snow mass data assimilation.

Kwon, Yonghwan↗

Human Performance Contributions to Safety in Commercial Aviation

Every day in aviation, pilots, air traffic controllers, and other front-line personnel perform countless correct judgments and actions in a variety of operational environments. These judgments and actions are often the difference between an accident and a non-event. Ironically, data on these behaviors are rarely collected or analyzed. Data-driven decisions about safety management and design of safety-critical systems are limited by the available data, which influence how decision makers characterize problems and identify solutions. Large volumes of data are collected on the failures and errors that result in infrequent incidents and accidents, but in the absence of data on behaviors that result in routine successful outcomes, safety management and system design decisions are based on a small sample of nonrepresentative safety data. This assessment aimed to find and document “safety successes” made possible by human operators. With many Aeronautics Research Mission Directorate (ARMD) Programs and Projects focusing on increased automation and autonomy and decreased human involvement, failure to fully consider the human contributions to successful system performance in civil aviation represents a significant risk — a risk that has not been recognized to date. Without understanding how humans contribute to safety, any estimate of predicted safety of autonomous capabilities is incomplete and inherently suspect. Furthermore, understanding the ways in which humans contribute to safety can promote strategic interactions among safety technologies, functions, procedures and the people using them. Without this understanding, the full benefits of an integrated, optimized human/technology or autonomous system will not be realized. Historically, safety has been consistently defined in terms of the occurrence of accidents or recognized risks (i.e., in terms of things that go wrong). These adverse outcomes are explained by identifying their causes, and safety is restored by eliminating or mitigating these causes. An alternative to this approach is to focus on what goes right and identify how to replicate that process. Focusing on the rare cases of failures attributed to “human error” provides little information about why human performance routinely prevents adverse events. Hollnagel has proposed that things go right because people continuously adjust their work to match their operating conditions. These adjustments become increasingly important as systems continue to grow in complexity. Thus, the definition of safety should reflect not only “avoiding things that go wrong” but “ensuring that things go right.” The basis for safety management requires developing an understanding of everyday activities. However, few mechanisms to monitor everyday work exist in the aviation domain, which limits opportunities to learn how designs function in reality. This concept of safety thinking and safety management is reflected in the emerging field of resilience engineering. According to Hollnagel, a system is resilient if it can sustain required operations under expected and unexpected conditions by adjusting its functioning prior to, during, or following changes, disturbances, and opportunities. To explore “positive” behaviors that contribute to resilient performance in commercial aviation, the assessment team examined a range of existing sources of data about pilot and air traffic control (ATC) tower controller performance, including subjective interviews with domain experts and objective aircraft flight data records. These data were used to identify strategies that support resilient performance, methods for exploring and refining those strategies in existing data, and proposed methods for capturing and analyzing new data.

Null, Cynthia H.↗

Accuracy of Center of Pressure Determination via Motion Capture

BACKGROUND: This study was conducted to support the stability assessment for tasks in lunar gravity and exercises on a Vibration Isolation and Stabilization (VIS) system in microgravity based on the dynamic feasibility criterion of whether the calculated position of the center of pressure (COP) falls within the base of support (BOS) which outlines the subject’s feet. Motion capture data combined with biomechanical modeling and simulation allows the forces and moments between the human and the VIS platform to be computed and the position of the COP as well as the location and shape of the BOS to be determined. The goal of this study was to assess the accuracy of the COP trajectory calculated using motion capture-based data. METHODS AND RESULTS: To obtain the dynamic quantities from which COP is calculated, motion capture data is first collected in the 1g lab environment by recording the trajectories of passive retroreflective markers placed on a subject during exercise or performance of a given task. The OpenSim [1] inverse kinematics (IK) tool is used to fit a scaled subject model to recorded marker trajectories while minimizing marker error to obtain joint angles. Then, a custom OpenSim plugin [2] is used to determine the subject’s time-varying moment of inertia and its time derivative, center of mass (CM) position, velocity, and acceleration, as well as the angular momentum and its time derivative relative to the subject’s CM. Some of these quantities are not needed for modeling tasks performed on a stationary lunar surface but, due to the moving exercise platform, are needed to model VIS response to the subject’s motion. Hand positions, used in calculating a cable force if present, are recorded as well. These quantities are used to calculate the total force (F ⃗^((plate) )) and moment (M ⃗^((plate) )) exerted by the lunar surface or the VIS plate on the subject’s shoe soles. COP is then calculated from the following equations: r_x^((cop) )= M_z^((plate) )/F_y^((plate) ) and r_z^((cop) )= 〖-M〗_x^((plate) )/F_y^((plate) ), where the y axis is normal to the surface. COP accuracy for feasibility assessments is then determined by whether it falls within the BOS, which is also computed by the plugin. To study the accuracy of COP calculated from motion capture, we first investigated whether COP remained within the BOS, as it must, for exercises performed in the 1g lab environment. Standard exercises such as back squat and deadlift were analyzed, as well as more explosive exercises including hang clean and press. Cases in which the COP exited the BOS indicated that COP accuracy required further investigation. In this study, an exercise device with cables was used, so cable force modeling accuracy should also be considered. In a separate study, we collected motion capture and force plate data for twenty-seven motions not involving an exercise device. About a third were genuine countermeasures exercises (e.g., hang clean and press), some were relevant for lunar tasks (e.g., object pick up), and the rest were of a “unit test” nature (e.g., swaying back and forth or side to side). Motion capture-based COP positions were compared with force plate measured COP. We found that while force plate measured COP remained within the BOS, motion capture-based COP was observed to briefly exit the BOS on occasion. Techniques to mitigate IK artifacts and filtering of calculated data could be used to improve the agreement of calculated and measured results, resolving excursions from the BOS within this dataset. The mean error between calculated and measured COP was found to be less than 6 mm. Additionally, we derived and investigated equations for the COP in terms of the cable force, cable location, as well as the human CM position, acceleration, and angular momentum with respect to the CM, and analyzed them for sensitivity to errors in individual quantities. Several were found, but the most significant one was that when the vertical force on the feet approaches zero, indicating a near-detachment or ‘jump off’ condition, errors are amplified. This is consistent with the observation that in the absence of pressure, the concept of the center of pressure would become meaningless.

C A Bell↗

A tone-aided dual vestigial sideband system for digital communications on fading channels

A spectrally efficient tone-aided dual vestigial sideband (TA/DVSB) system for digital data communications on fading channels is presented and described analytically. This PSK (phase-shift-keying) system incorporates a feed-forward, tone-aided demodulation technique to compensate for Doppler frequency shift and channel- induced, multipath fading. In contrast to other tone-in-band-type systems, receiver synchronization is derived from the complete data VSBs. Simulation results for the Rician fading channel are presented. These results demonstrate the receiver's ability to mitigate performance degradation due to fading and to obtain proper data carrier synchronization, suggesting that the proposed TA/DVSB system has promise for this application. Simulated BER (bit-error rate) data indicate that the TA/DVSB system effectively alleviates the channel distortions of the land mobile satellite application.

Hladik, Stephen M.↗

Polarization Performance Simulation for the GeoXO Atmospheric Composition Instrument: NO2 Retrieval Impacts

NOAA's Geostationary Extended Observations (GeoXO) constellation will continue and expand on the capabilities of the current generation of geostationary satellite systems to support US weather, ocean, atmosphere, and climate operations. It is planned to consist of a dedicated atmospheric composition instrument (ACX) to support air quality forecasting and monitoring by providing capabilities similar to missions such as TEMPO (Tropospheric Emission: Monitoring Pollution), currently planned to launch in 2023, as well as OMI (Ozone Monitoring Instrument), TROPOMI (TROPOspheric Monitoring Instrument), and GEMS (Geostationary Environment Monitoring Spectrometer) currently in operation. As the early phases of ACX development are progressing, design trade-offs are being considered to understand the relationship between instrument design choices and trace gas retrieval impacts. Some of these choices will affect the instrument polarization sensitivity (PS), which can have radiometric impacts on environmental satellite observations. We conducted a study to investigate how such radiometric impacts can affect NO2 retrievals by exploring their sensitivities to time of day, location, and scene type with an ACX instrument model that incorporates PS. The study addresses the basic steps of operational NO2 retrievals: the spectral fitting step and the conversion of slant column to vertical column via the air mass factor (AMF). The spectral fitting step was performed by generating at-sensor radiance from a clear-sky scene with a known NO2 amount, the application of an instrument model including both instrument PS and noise, and a physical retrieval. The spectral fitting step was found to mitigate the impacts of instrument PS. The AMF-related step was considered for clear-sky and partially cloudy scenes, for which instrument PS can lead to errors in interpreting the cloud content, propagating to AMF errors and finally to NO2 retrieval errors. For this step, the NO2 retrieval impacts were small but non-negligible for high NO2 amounts; we estimated that a typical high NO2 amount can cause a maximum retrieval error of 0.25×1015 molec. cm−2 for a PS of 5 %. These simulation capabilities were designed to aid in the development of a GeoXO atmospheric composition instrument that will improve our ability to monitor and understand the Earth's atmosphere.

Aaron Pearlman↗

Astronaut Biography Project for Countermeasures of Human Behavior and Performance Risks in Long Duration Space Flights

This final report will summarize research that relates to human behavioral health and performance of astronauts and flight controllers. Literature reviews, data archival analyses, and ground-based analog studies that center around the risk of human space flight are being used to help mitigate human behavior and performance risks from long duration space flights. A qualitative analysis of an astronaut autobiography was completed. An analysis was also conducted on exercise countermeasure publications to show the positive affects of exercise on the risks targeted in this study. The three main risks targeted in this study are risks of behavioral and psychiatric disorders, risks of performance errors due to poor team performance, cohesion, and composition, and risks of performance errors due to sleep deprivation, circadian rhythm. These three risks focus on psychological and physiological aspects of astronauts who venture out into space on long duration space missions. The purpose of this research is to target these risks in order to help quantify, identify, and mature countermeasures and technologies required in preventing or mitigating adverse outcomes from exposure to the spaceflight environment

Banks, Akeem↗

Impact of Pilot Delay and Non-Responsiveness on the Safety Performance of Airborne Separation

Assessing the safety effects of prediction errors and uncertainty on automationsupported functions in the Next Generation Air Transportation System concept of operations is of foremost importance, particularly safety critical functions such as separation that involve human decision-making. Both ground-based and airborne, the automation of separation functions must be designed to account for, and mitigate the impact of, information uncertainty and varying human response. This paper describes an experiment that addresses the potential impact of operator delay when interacting with separation support systems. In this study, we evaluated an airborne separation capability operated by a simulated pilot. The experimental runs are part of the Safety Performance of Airborne Separation (SPAS) experiment suite that examines the safety implications of prediction errors and system uncertainties on airborne separation assistance systems. Pilot actions required by the airborne separation automation to resolve traffic conflicts were delayed within a wide range, varying from five to 240 seconds while a percentage of randomly selected pilots were programmed to completely miss the conflict alerts and therefore take no action. Results indicate that the strategicAirborne Separation Assistance System (ASAS) functions exercised in the experiment can sustain pilot response delays of up to 90 seconds and more, depending on the traffic density. However, when pilots or operators fail to respond to conflict alerts the safety effects are substantial, particularly at higher traffic densities.

Consiglio, Maria↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Robust Damage-Mitigating Control of Aircraft for High Performance and Structural Durability

This paper presents the concept and a design methodology for robust damage-mitigating control (DMC) of aircraft. The goal of DMC is to simultaneously achieve high performance and structural durability. The controller design procedure involves consideration of damage at critical points of the structure, as well as the performance requirements of the aircraft. An aeroelastic model of the wings has been formulated and is incorporated into a nonlinear rigid-body model of aircraft flight-dynamics. Robust damage-mitigating controllers are then designed using the H(infinity)-based structured singular value (mu) synthesis method based on a linearized model of the aircraft. In addition to penalizing the error between the ideal performance and the actual performance of the aircraft, frequency-dependent weights are placed on the strain amplitude at the root of each wing. Using each controller in turn, the control system is put through an identical sequence of maneuvers, and the resulting (varying amplitude cyclic) stress profiles are analyzed using a fatigue crack growth model that incorporates the effects of stress overload. Comparisons are made to determine the impact of different weights on the resulting fatigue crack damage in the wings. The results of simulation experiments show significant savings in fatigue life of the wings while retaining the dynamic performance of the aircraft.

Caplin, Jeffrey↗

Ultracoherent SRF Cavity-Based Multi-Qudit Platform with Error-Resilient Control

Superconducting radio-frequency (SRF) cavities offer a promising platform for quantum computing due to their long coherence times, yet integrating nonlinear elements like transmons for control often introduces additional loss. We report a multimode quantum system based on a 2-cell elliptical shaped SRF cavity, comprising two cavity modes weakly coupled to an ancillary transmon circuit, designed to preserve coherence while enabling efficient control of the cavity modes. We mitigate the detrimental effects of the transmon decoherence through careful design optimization that reduces transmon-cavity couplings and participation in the dielectric substrate and lossy interfaces, to achieve single-photon lifetimes of 20.6\,ms and 15.6\,ms for the two modes, and a pure dephasing time exceeding 40\,ms. This marks an order-of-magnitude improvement over prior 3D multimode memories. Leveraging sideband interactions and novel error-resilient protocols, including measurement-based correction and post-selection, we achieve high-fidelity control over quantum states. This enables the preparation of Fock states up to $N = 20$ with fidelities exceeding 95\%, the highest reported to date to the authors' knowledge, as well as two-mode entanglement with an estimated coherence-limited fidelities of 99.9\% after post-selection. These results establish our platform as a robust foundation for quantum information processing, allowing for future extensions to high-dimensional qudit encodings.

Kim, T. [Northwestern U.]↗

Ultracoherent superconducting cavity-based multiqudit platform with error-resilient control

Superconducting radio-frequency (SRF) cavities offer a promising platform for quantum computing due to their long coherence times, yet integrating nonlinear elements like transmons for control often introduces additional loss. We report a multimode quantum system based on a 2-cell elliptical shaped SRF cavity, comprising two cavity modes weakly coupled to an ancillary transmon circuit, designed to preserve coherence while enabling efficient control of the cavity modes. We mitigate the detrimental effects of the transmon decoherence through careful design optimization that reduces transmon-cavity couplings and participation in the dielectric substrate and lossy interfaces, to achieve single-photon lifetimes of 20.6 ms and 15.6 ms for the two modes, and a pure dephasing time exceeding 40 ms. This marks an order-of-magnitude improvement over prior 3D multimode memories. Leveraging sideband interactions and novel error-resilient protocols, including measurement-based correction and post-selection, we achieve high-fidelity control over quantum states. This enables the preparation of Fock states up to $N = 20$ with fidelities exceeding 95%, the highest reported to date to the authors' knowledge, as well as two-mode entanglement with an estimated coherence-limited fidelities of 99.9% after post-selection. These results establish our platform as a robust foundation for quantum information processing, allowing for future extensions to high-dimensional qudit encodings.

Kim, Taeyoon [Fermilab; Northwestern U.]↗

Solid State Transformer Controls for Mitigation of E3a High-Altitude Electromagnetic Pulse Insults

This paper explores the use of a solid state transformer (SST) to mitigate the 𝐸 3𝐴 component of a high-altitude electromagnetic pulse (HEMP) insult using external energy storage optimal control techniques. In lieu of conventional passive blocking devices or feedback-controlled energy storage devices, a novel implementation of Hamiltonian error tracking is utilized to develop a feedback control law for the variable converter ratio in an SST. The findings of the simulations performed in this paper suggest that additional energy storage is not necessary to protect an individual load from a HEMP insult. The simulations performed examine the response of a single-phase SST connected to a single voltage source on a long transmission line on the one side and a single linear resistor on the other. The control law is specifically developed for the late-time, low-frequency portion of a HEMP insult, namely the 𝐸 3𝐴 components. The Hamiltonian error-based converter ratio control law is compared with nonlinear optimal feedforward controls to show that the HSSPFC is an external energy storage optimal controller.

HEMP mitigation↗

Ultracoherent superconducting cavity-based multiqudit platform with error-resilient control

Superconducting radio-frequency (SRF) cavities offer a promising platform for quantum computing due to their long coherence times, yet integrating nonlinear elements like transmons for control often introduces additional loss. We report a multimode quantum system based on a 2-cell elliptical-shaped SRF cavity, comprising two cavity modes weakly coupled to an ancillary transmon circuit, designed to preserve coherence while enabling efficient control of the cavity modes. We mitigate the detrimental effects of the transmon decoherence through careful design optimization that reduces transmon-cavity couplings and participation in the dielectric substrate and lossy interfaces, to achieve single-photon lifetimes of 20.6 ms and 15.6 ms for the two modes, and a pure dephasing time exceeding 40 ms. This marks an order-of-magnitude improvement over prior 3D multimode memories. Leveraging sideband interactions and novel error-resilient protocols, including measurement-based correction and post-selection, we achieve high-fidelity control over quantum states. This enables the preparation of Fock states up to N = 20 with fidelities exceeding 95%, the highest reported to date to the authors' knowledge, as well as two-mode entanglement with an estimated coherence-limited fidelities of 99.9% after post-selection. These results establish our platform as a robust foundation for quantum information processing, allowing for future extensions to high-dimensional qudit encodings.

Lu, Yao [Fermilab] (ORCID:000000020413698X)↗

Ultracoherent superconducting cavity-based multiqudit platform with error-resilient control

Superconducting radio-frequency (SRF) cavities offer a promising platform for quantum computing due to their long coherence times, yet integrating nonlinear elements like transmons for control often introduces additional loss. We report a multimode quantum system based on a 2-cell elliptical-shaped SRF cavity, comprising two cavity modes weakly coupled to an ancillary transmon circuit, designed to preserve coherence while enabling efficient control of the cavity modes. We mitigate the detrimental effects of the transmon decoherence through careful design optimization that reduces transmon-cavity couplings and participation in the dielectric substrate and lossy interfaces, to achieve single-photon lifetimes of 20.6 ms and 15.6 ms for the two modes, and a pure dephasing time exceeding 40 ms. This marks an order-of-magnitude improvement over prior 3D multimode memories. Leveraging sideband interactions and novel error-resilient protocols, including measurement-based correction and post-selection, we achieve high-fidelity control over quantum states. This enables the preparation of Fock states up to N = 20 with fidelities exceeding 95%, the highest reported to date to the authors' knowledge, as well as two-mode entanglement with an estimated coherence-limited fidelities of 99.9% after post-selection. These results establish our platform as a robust foundation for quantum information processing, allowing for future extensions to high-dimensional qudit encodings.

Lu, Yao [Fermilab] (ORCID:000000020413698X)↗

Validation and Scaling of Soil Moisture in a Semi-Arid Environment: SMAP Validation Experiment 2015 (SMAPVEX15)

The NASA SMAP (Soil Moisture Active Passive) mission conducted the SMAP Validation Experiment 2015 (SMAPVEX15) in order to support the calibration and validation activities of SMAP soil moisture data products. The main goals of the experiment were to address issues regarding the spatial disaggregation methodologies for improvement of soil moisture products and validation of the in situ measurement upscaling techniques. To support these objectives high-resolution soil moisture maps were acquired with the airborne PALS (Passive Active L-band Sensor) instrument over an area in southeast Arizona that includes the Walnut Gulch Experimental Watershed (WGEW), and intensive ground sampling was carried out to augment the permanent in situ instrumentation. The objective of the paper was to establish the correspondence and relationship between the highly heterogeneous spatial distribution of soil moisture on the ground and the coarse resolution radiometer-based soil moisture retrievals of SMAP. The high-resolution mapping conducted with PALS provided the required connection between the in situ measurements and SMAP retrievals. The in situ measurements were used to validate the PALS soil moisture acquired at 1-km resolution. Based on the information from a dense network of rain gauges in the study area, the in situ soil moisture measurements did not capture all the precipitation events accurately. That is, the PALS and SMAP soil moisture estimates responded to precipitation events detected by rain gauges, which were in some cases not detected by the in situ soil moisture sensors. It was also concluded that the spatial distribution of the soil moisture resulted from the relatively small spatial extents of the typical convective storms in this region was not completely captured with the in situ stations. After removing those cases (approximately10 of the observations) the following metrics were obtained: RMSD (root mean square difference) of0.016m3m3 and correlation of 0.83. The PALS soil moisture was also compared to SMAP and in situ soil moisture at the 36-km scale, which is the SMAP grid size for the standard product. PALS and SMAP soil moistures were found to be very similar owing to the close match of the brightness temperature measurements and the use of a common soil moisture retrieval algorithm. Spatial heterogeneity, which was identified using the high-resolution PALS soil moisture and the intensive ground sampling, also contributed to differences between the soil moisture estimates. In general, discrepancies found between the L-band soil moisture estimates and the 5-cm depth in situ measurements require methodologies to mitigate the impact on their interpretations in soil moisture validation and algorithm development. Specifically, the metrics computed for the SMAP radiometer-based soil moisture product over WGEW will include errors resulting from rainfall, particularly during the monsoon season when the spatial distribution of soil moisture is especially heterogeneous.

SMAPVEX15↗

Mitigating Photon Jitter in Optical PPM Communication

A theoretical analysis of photon-arrival jitter in an optical pulse-position-modulation (PPM) communication channel has been performed, and now constitutes the basis of a methodology for designing receivers to compensate so that errors attributable to photon-arrival jitter would be minimized or nearly minimized. Photon-arrival jitter is an uncertainty in the estimated time of arrival of a photon relative to the boundaries of a PPM time slot. Photon-arrival jitter is attributable to two main causes: (1) receiver synchronization error [error in the receiver operation of partitioning time into PPM slots] and (2) random delay between the time of arrival of a photon at a detector and the generation, by the detector circuitry, of a pulse in response to the photon. For channels with sufficiently long time slots, photon-arrival jitter is negligible. However, as durations of PPM time slots are reduced in efforts to increase throughputs of optical PPM communication channels, photon-arrival jitter becomes a significant source of error, leading to significant degradation of performance if not taken into account in design. For the purpose of the analysis, a receiver was assumed to operate in a photon- starved regime, in which photon counts follow a Poisson distribution. The analysis included derivation of exact equations for symbol likelihoods in the presence of photon-arrival jitter. These equations describe what is well known in the art as a matched filter for a channel containing Gaussian noise. These equations would yield an optimum receiver if they could be implemented in practice. Because the exact equations may be too complex to implement in practice, approximations that would yield suboptimal receivers were also derived.

Moision, Bruce↗

A Risk Analysis Tool for Estimating the Risk of Electrical Failures Due to Human Induced Defects

Aerospace electrical systems are required to withstand and adequately operate in extremely harsh environments that include, for example, high radiation exposure, temperature extremes, intense vibrational stress and drastic temperature cycling. The nature of aerospace electronics also demands high reliability since, with very few exceptions, there is no chance for hardware servicing or repairs. Common risk mitigation techniques for this type of situation are to perform a Reliability Analysis of the system throughout the development cycle, and to use electrical components that are regarded as “high reliability” because of additional controls and requirements applied in their design, manufacturing and testing. Unfortunately, studies have shown that even though these techniques are used, many systems fail to meet mission requirements well before the predicted lifetimes. This paper presents the analysis of failures of electrical parts, experienced during various stages of system development, at NASA Goddard Space Flight Center, Greenbelt MD, between the years 2001 and 2013. These components were subjected to qualification, screening and testing in which the goal was to ensure that the components would survive the stresses of the mission. The analysis categorizes failures by part type and failure mechanisms. One of the results of the analysis was the realization that a surprising proportion of failures experienced during system integration and testing were caused by human error (i.e. human induced defect). Further analysis included the determination of root failure mechanisms and any influencing factors contributing to these failures. The major causes of these defects were attributed to electrostatic damage (ESD), electrical overstress (EOS), mechanical overstress (MOS), and thermal overstress (TOS). Finally, the study proposes a risk analysis tool which incorporates these major causes for the failures, termed error-producing conditions (EPCs), and a proportionality factor representing the number of each type of failure that has occurred at the facility under study. These factors are quantified and used to communicate the risk of human induced defects for the assembly, integration and testing of space hardware based on the system’s electrical parts list. The new risk identification can trigger risk-mitigating actions more effectively, based on the presence of component categories or other hazardous conditions that have a history of failure due to human error.

Majewicz, Peter J.↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗