Detection of Random Test Failures and their Impact on CI Testing Pipelines
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Cardinal is a wrapping of the GPU-oriented spectral element Computational Fluid Dynamics (CFD) code NekRS and the Monte Carlo particle transport code OpenMC within the Multiphysics Object-Oriented Simulation Environment (MOOSE). Cardinal provides high-resolution thermal-hydraulics and/or radiation transport feedback to MOOSE multiphysics simulations. Multiphysics feedback is implemented in a geometry-agnostic manner which eliminates the need for rigid one-to-one mappings. A generic data transfer implementation also allows NekRS and OpenMC to couple to any MOOSE application, enabling a broad set of multiphysics capabilities. Cardinal simulations can also leverage combinations of MPI, OpenMP, and GPU resources. Cardinal continuous development and improvement efforts have led to the software being considered as a high-fidelity design and licensing tool for key areas of nuclear reactor relevant physics, including neutron transport, fluid flow, heat transfer, and mechanical processes. The fast development and expansion of the software from a pure R&D framework towards its application in the nuclear industry and regulation require a focus on developing, enhancing and, maintaining Cardinal’s software quality through strict adherence to a Software Quality Assurance (SQA) framework and SQA program. To facilitate compliance with SQA standards, the Cardinal SQA Program has been initiated during Fiscal Year 2023 (FY23). During the development of the Cardinal SQA Program, multiple gaps have been identified. These gaps are primarily related to model verification and code pedigree as they relate to the use of Cardinal as a safety analysis tool. These gaps have been captured in a report published in 2023. A second report highlighted the progress made during Fiscal Year 2024 (FY24) and described Argonne’s effort to document and integrate software verification within Cardinal’s software development process. This report documents a snapshot of the verification test cases currently available for Cardinal and NekRS in their assimilation into a Continuous Integration (CI) platform. Following the CI practice permits the integrating of source code changes frequently and ensuring that the integrated codebase clears the verification testing for the software. It should be noted that the SQA program itself, including the program plans, procedures, configuration management, and testing strategies, need to be developed in a future step of this task.
This report documents the results of spiral bevel gear rig tests performed under a NASA Space Act Agreement with the Federal Aviation Administration (FAA) to support validation and demonstration of rotorcraft Health and Usage Monitoring Systems (HUMS) for maintenance credits via FAA Advisory Circular (AC) 29-2C, Section MG-15, Airworthiness Approval of Rotorcraft (HUMS) (Ref. 1). The overarching goal of this work was to determine a method to validate condition indicators in the lab that better represent their response to faults in the field. Using existing in-service helicopter HUMS flight data from faulted spiral bevel gears as a "Case Study," to better understand the differences between both systems, and the availability of the NASA Glenn Spiral Bevel Gear Fatigue Rig, a plan was put in place to design, fabricate and test comparable gear sets with comparable failure modes within the constraints of the test rig. The research objectives of the rig tests were to evaluate the capability of detecting gear surface pitting fatigue and other generated failure modes on spiral bevel gear teeth using gear condition indicators currently used in fielded HUMS. Nineteen final design gear sets were tested. Tables were generated for each test, summarizing the failure modes observed on the gear teeth for each test during each inspection interval and color coded based on damage mode per inspection photos. Gear condition indicators (CI) Figure of Merit 4 (FM4), Root Mean Square (RMS), +/- 1 Sideband Index (SI1) and +/- 3 Sideband Index (SI3) were plotted along with rig operational parameters. Statistical tables of the means and standard deviations were calculated within inspection intervals for each CI. As testing progressed, it became clear that certain condition indicators were more sensitive to a specific component and failure mode. These tests were clustered together for further analysis. Maintenance actions during testing were also documented. Correlation coefficients were calculated between each CI, component, damage state and torque. Results found test rig and gear design, type of fault and data acquisition can affect CI performance. Results found FM4, SI1 and SI3 can be used to detect macro pitting on two more gear or pinion teeth as long as it is detected prior to progressing to other components or transitioning to another failure mode. The sensitivity of RMS to system and operational conditions limit its reliability for systems that are not maintained at steady state. Failure modes that occurred due to scuffing or fretting were challenging to detect with current gear diagnostic tools, since the damage is distributed across all the gear and pinion teeth, smearing the impacting signatures typically used to differentiate between a healthy and damaged tooth contact. This is one of three final reports published on the results of this project. In the second report, damage modes experienced in the field will be mapped to the failure modes created in the test rig. The helicopter CI data will then be re-processed with the same analysis techniques applied to spiral bevel rig test data. In the third report, results from the rig and helicopter data analysis will be correlated. Observations, findings and lessons learned using sub-scale rig failure progression tests to validate helicopter gear condition indicators will be presented.
The Controller-Pilot Data Link Communications (CPDLC) and Air Traffic Control workstation research was conducted as part of the 1997 NASA Low Visibility Landing and Surface Operations (LVLASO) demonstration program at Atlanta Hartsfield airport. Research activity under this grant increased the sophistication of the Controllers' Communication and Situational Awareness Terminal (C-CAST) and developed a VHF Data Link -Mode 2 communications platform. The research culminated with participation in the 2000 NASA Aviation Safety Program's Synthetic Vision System (SVS) / Runway Incursion Prevention System (RIPS) flight demonstration at Dallas-Fort Worth Airport.
In this talk we will discuss how we leverage the Cloud for HDF software daily regression testing including testing of the HDF5 parallel library on the Cloud cluster using Orange FS.
A "seeded fault test" in support of a rotorcraft condition based maintenance program (CBM), is an experiment in which a component is tested with a known fault while health monitoring data is collected. These tests are performed at operating conditions comparable to operating conditions the component would be exposed to while installed on the aircraft. Performance of seeded fault tests is one method used to provide evidence that a Health Usage Monitoring System (HUMS) can replace current maintenance practices required for aircraft airworthiness. Actual in-service experience of the HUMS detecting a component fault is another validation method. This paper will discuss a hybrid validation approach that combines in service-data with seeded fault tests. For this approach, existing in-service HUMS flight data from a naturally occurring component fault will be used to define a component seeded fault test. An example, using spiral bevel gears as the targeted component, will be presented. Since the U.S. Army has begun to develop standards for using seeded fault tests for HUMS validation, the hybrid approach will be mapped to the steps defined within their Aeronautical Design Standard Handbook for CBM. This paper will step through their defined processes, and identify additional steps that may be required when using component test rig fault tests to demonstrate helicopter CI performance. The discussion within this paper will provide the reader with a better appreciation for the challenges faced when defining a seeded fault test for HUMS validation.
Cardinal is a wrapping of the GPU-oriented spectral element Computational Fluid Dynamics (CFD) code NekRS and the Monte Carlo particle transport code OpenMC within the Multiphysics Object-Oriented Simulation Environment (MOOSE). Cardinal provides high-resolution thermal-hydraulics and/or radiation transport feedback to MOOSE multiphysics simulations. Multiphysics feedback is implemented in a geometry-agnostic manner which eliminates the need for rigid one-to-one mappings. A generic data transfer implementation also allows NekRS and OpenMC to couple to any MOOSE application, enabling a broad set of multiphysics capabilities. Cardinal simulations can also leverage combinations of MPI, OpenMP, and GPU resources. Cardinal continuous development and improvement efforts have led to the software being considered as a high-fidelity design and licensing tool for key areas of nuclear reactor relevant physics, including neutron transport, fluid flow, heat transfer, and mechanical processes. The fast development and expansion of the software from a pure R&D framework towards its application in the nuclear industry and regulation require a focus on developing, enhancing,and maintaining Cardinal’s software quality through strict adherence to a Software Quality Assurance (SQA) framework and SQA program. To facilitate compliance with SQA standards, the Cardinal SQA Program was initiated during Fiscal Year 2023 (FY23). During the development of the Cardinal SQA Program, multiple gaps have been identified. These gaps are primarily related to model verification and code pedigree as they relate to the use of Cardinal as an analysis tool. These gaps were captured in a report published in 2023. A second report highlighted the progress made during Fiscal Year 2024 (FY24) and described Argonne’s effort to document and integrate software verification within Cardinal’s software development process. This report documents the progress made towards NQA-1 for Cardinal in the Fiscal Year 2025 (FY25). All cases in the expanded Continuous Integration (CI) suite of NekRS are included in this report which test the solvers and modules available in NekRS exhaustively. The NekRS tests are integrated with the Cardinal CI suite and made available in publicly accessible Github documentation. Following the CI practice permits integrating of source code changes frequently and ensuring that the integrated codebase clears the verification testing for the software. Also in this report is a brief overview of the development of the Cardinal Software Quality Assurance Plan (SQAP) that was done in FY25, though it should be noted that the rest of the documentation for the SQA program needs to be developed in a future step of this task.
While the HEDs provide an extremely useful basis for interpreting data from the Dawn mission, there is no guarantee that they provide a complete vision of all possible crustal (and possibly mantle) lithologies that are exposed at the surface of Vesta. With this in mind, an alternative approach is to identify plausible bulk compositions and use mass-balance and geochemical modelling to predict possible internal structures and crust/mantle compositions and mineralogies. While such models must be consistent with known HED samples, this approach has the potential to extend predictions to thermodynamically plausible rock types that are not necessarily present in the HED collection. Nine chondritic bulk compositions are considered (CI, CV, CO, CM, H, L, LL, EH, EL). For each, relative proportions and densities of the core, mantle, and crust are quantified. This calculation is complicated by the fact that iron may occur in metallic form (in the core) and/or in oxidized form (in the mantle and crust). However, considering that the basaltic crust has the composition of Juvinas and assuming that this crust is in thermodynamic equilibrium with the residual mantle, it is possible to calculate a single solution to this problem for a given bulk composition. Of the nine bulk compositions tested, solutions corresponding to CI and LL groups predicted a negative metal fraction and were not considered further. Solutions for enstatite chondrites imply significant oxidation relative to the starting materials and these solutions too are considered unlikely. For the remaining bulk compositions, the relative proportion of crust to bulk silicate is typically in the range 15 to 20% corresponding to crustal thicknesses of 15 to 20 km for a porosity-free Vesta-sized body. The mantle is predicted to be largely dominated by olivine (greater than 85%) for carbonaceous chondrites, but to be a roughly equal mixture of olivine and pyroxene for ordinary chondrite precursors. All bulk compositions have a significant core, but the relative proportions of metal and sulphide can be widely different. Using these data, total core size (metal+ sulphide) and average core densities can be calculated, providing a useful reference frame within which to consider geophysical/gravity data of the Dawn mission. Further to these mass-balance calculations, the MELTS thermodynamic calculator has been used to assess to what extent chondritic bulk compositions can produce Juvinas-like liquids at relevant degrees of partial melting/crystallization. This work will refine acceptable bulk compositions and predict the mineralogy and composition of the associated solid and liquid products over wide ranges of partial melting and crystallization, providing a useful and self-consistent reference frame for interpretation of the data from the VIR and GRaND instruments onboard the Dawn spacecraft.
The focus of this paper was to compare the performance of HUMS condition indicators (CI) when detecting a bearing fault in a test stand or on a helicopter. This study compared data from two sources: first, CI data collected from accelerometers installed on two UH-60 Black Hawk helicopters when oil cooler bearing faults occurred, along with data from helicopters with no bearing faults; and second, CI data that was collected from ten cooler bearings, healthy and faulted, that were removed from fielded helicopters and installed in a test stand. A method using Receiver Operating Characteristic (ROC) curves to compare CI performance was demonstrated. Results indicated the bearing energy CI responded differently for the helicopter and the test stand. Future research is required if test stand data is to be used validate condition indicator performance on a helicopter.
ROOT is an open source framework, freely available on GitHub, at the heart of data acquisition, processing and analysis of HE(N)P experiments, and beyond. It is developed collaboratively: contributions are not authored only by ROOT team members, but also by the user community at large: developers and scientists from universities, labs as well as the private sector. More than 1500 GitHub Pull Requests are merged on average per year. It is in this context that code integration acquires a primary role. The review of code contributions isn’t enough: not only they need to be thoroughly reviewed, they also need to be thoroughly tested through a powerful CI infrastructure on several different platforms to comply with the high code quality standards of the project. Since the end of 2023, ROOT moved its continuous integration system from Jenkins to GitHub Actions. In this contribution, we characterise the transition to the GitHub CI, focussing on our strategy, its implementation and the lessons learned, as well as the advantages the new system offers with respect to the previous one. Particular emphasis will be given to the evaluation of the cost-benefit ratio for Jenkins and GitHub Actions for the ROOT project. We also describe how we manage to run in less than one hour thousands of unit, integration, functional and end-to-end tests on different flavours of Windows, four versions of macOS, as well as about ten of the most used Linux distributions, taking advantage of the CERN computing infrastructure.
For over two decades, the dCache project has provided open-source to satisfy ever-more demanding storage requirements. More than 80 sites worldwide rely on dCache to provide services for LHC experiments, Belle-II, Eu- XFEL, and others. This can be achieved only with a well-established process from a whiteboard, where ideas are created through development, packaging, and testing. The project’s build and test infrastructure is based on Jenkins CI and a set of virtual machines. This infrastructure is maintained by dCache developers. With the introduction of the DESY-central Gitlab server, the developers have started migrating from VM-based testing to container-based deployments in the onsite Kubernetes cluster. As a result, we have packaged dCache containers and Helm charts that can be used by other sites to reproduce our test and build steps quickly or to evaluate new releases on their pre-production systems and, eventually, become a standard model of dCache deployment at the sites. This paper describes the challenges we have faced, the techniques we used to solve them, and the issues that still need to be addressed.
When one or more new values are added to a developing time series, they change its descriptive parameters (mean, variance, trend, coherence). A 'change index (CI)' is developed as a quantitative indicator that the changed parameters remain compatible with the existing 'base' data. CI formulate are derived, in terms of normalized likelihood ratios, for small samples from Poisson, Gaussian, and Chi-Square distributions, and for regression coefficients measuring linear or exponential trends. A substantial parameter change creates a rapid or abrupt CI decrease which persists when the length of the bases is changed. Except for a special Gaussian case, the CI has no simple explicit regions for tests of hypotheses. However, its design ensures that the series sampled need not conform strictly to the distribution form assumed for the parameter estimates. The use of the CI is illustrated with both constructed and observed data samples, processed with a Fortran code 'Sequitor'.
Helicopter health monitoring systems use vibration signatures generated from damaged components to identify transmission faults. For damaged gears, these signatures relate to changes in dynamics due to the meshing of the damaged tooth. These signatures, referred to as condition indicators (CI), can perform differently when measured on different systems, such as a component test rig, or a full-scale transmission test stand, or an aircraft. These differences can result from dissimilarities in systems design and environment under dynamic operating conditions. The static structure can also filter the response between the vibration source and the accelerometer, when the accelerometer is installed on the housing. To assess the utility of static vibration transfer paths for predicting gear CI performance, measurements were taken on the NASA Glenn Spiral Bevel Gear Fatigue Test Rig. The vibration measurements were taken to determine the effect of torque, accelerometer location and gearbox design on accelerometer response. Measurements were taken at the housing and compared while impacting the gear set near mesh. These impacts were made at gear mesh to simulate gear meshing dynamics. Data measured on a helicopter gearbox installed in a static fixture were also compared to the test rig. The behavior of the structure under static conditions was also compared to CI values calculated under dynamic conditions. Results indicate that static vibration transfer path measurements can provide some insight into spiral bevel gear CI performance by identifying structural characteristics unique to each system that can affect specific CI response.
We know that the stones that fall to earth as meteorites are not representative of the full diversity of small solar system bodies, because of the peculiarities of the dynamical processes that send material into Earth-crossing paths [1] which result in severe selection biases. Thus, the bulk of the meteorites that fall are insufficient to understand the full range of early solar system processes. However, the situation is different for pebble- and smaller-sized objects that stream past the giant planets and asteroid belts into the inner solar system in a representative manner. Thus, micrometeorites and interplanetary dust particles have been exploited to permit study of objects that do not provide meteorites to earth. However, there is another population of materials that sample a larger range of small solar system bodies, but which have received little attention - pebble-sized foreign clasts in meteorites (also called xenoliths, dark inclusions, clasts, etc.). Unfortunately, most previous studies of these clasts have been misleading, in that these objects have simply been identified as pieces of CM or CI chondrites. In our work we have found this to be generally erroneous, and that CM and especially CI clasts are actually rather rare. We therefore test the hypothesis that these clasts sample the full range of small solar system bodies. We have located and obtained samples of clasts in 81 different meteorites, and have begun a thorough characterization of the bulk compositions, mineralogies, petrographies, and organic compositions of this unique sample set. In addition to the standard e-beam analyses, recent advances in technology now permit us to measure bulk O isotopic compositions, and major- though trace-element compositions of the sub-mm-sized discrete clasts. Detailed characterization of these clasts permit us to explore the full range of mineralogical and petrologic processes in the early solar system, including the nature of fluids in the Kuiper belt and the outer main asteroid belt, as revealed by the mineralogy of secondary phases.
Our objective is to demonstrate an innovative method combining machine learning with comparative effectiveness research techniques and to investigate a hitherto unstudied question about the effectiveness of common prescribing patterns. For Operation Enduring Freedom/Operation Iraqi Freedom veterans with major depressive disorder, we generate pharmacotherapy pathways (of antidepressants) using process mining and machine learning. We select the medication episodes that were started at subtherapeutic doses by the first assigned primary care physician and observe the paths that those medication episodes follow. Using 2-stage least squares, we test the effectiveness of starting at a low dose and staying low for longer versus ramping up fast while balancing observable and unobservable characteristics of patients and providers through instrumental variables. We leverage predetermined provider practice patterns as instruments. We collected outpatient pharmacy data for selective serotonin reuptake inhibitors and selective norepinephrine reuptake inhibitors, patient and provider characteristics (as control variables), and the instruments for our cohort. All data were extracted for the period between 2006 and 2020. There is a statistically significant positive effect (0.68, 95% CI 0.11–1.25) of “ramping up fast” on engagement in care. When we examine the effect of “ramping up slow”, we see an insignificant negative impact on engagement in care (−0.82, 95% CI −1.89 to 0.25). As expected, the probability of drop-out also seems to have a negative effect on engagement in care (−0.39, 95% CI −0.94 to 0.17). We further validate these results by testing with medication possession ratios calculated periodically as an alternative engagement in care metric. Our findings contradict the “Start low, go slow” adage, indicating that ramping up the dose of an antidepressant faster has a significantly positive effect on engagement in care for our population.
Understanding the global relation between Earthquake-Induced Landslides (EQILs) and the factors contributing to their initiation is still an open topic within the geomorphological community. Accessing EQIL inventories and analyzing them concerning their potential causes is the key to explore such relation. However, each of the existing EQIL inventories has its level of completeness and associated uncertainty which makes any unified relationship which is challenging to obtain. So far, the completeness of EQIL inventories has never been clearly defined. As a result, it has never been accounted for in global EQIL predictive models. In this technical note, we propose a simple definition for the completeness of EQIL inventories. We analyze 30 digital EQIL for 21 earthquakes and develop a semi-quantitative method to estimate the completeness level, which we refer to as Completeness Index (CI). The CI results from a combination of topographic factors, ground shaking parameters, and a measure derived from landslide size statistics. The proposed CI consists of Low, Moderate, and High completeness classes which can be used to evaluate any landslide inventory. We made the whole procedure to compute the CI accessible in an ArcGIS toolbox together with a test dataset.
This is the final of three reports published on the results of this project. In the first report, results were presented on nineteen tests performed in the NASA Glenn Spiral Bevel Gear Fatigue Test Rig on spiral bevel gear sets designed to simulate helicopter fielded failures. In the second report, fielded helicopter HUMS data from forty helicopters were processed with the same techniques that were applied to spiral bevel rig test data. Twenty of the forty helicopters experienced damage to the spiral bevel gears, while the other twenty helicopters had no known anomalies within the time frame of the datasets. In this report, results from the rig and helicopter data analysis will be compared for differences and similarities in condition indicator (CI) response. Observations and findings using sub-scale rig failure progression tests to validate helicopter gear condition indicators will be presented. In the helicopter, gear health monitoring data was measured when damage occurred and after the gear sets were replaced at two helicopter regimes. For the helicopters or tails, data was taken in the flat pitch ground 101 rotor speed (FPG101) regime. For nine tails, data was also taken at 120 knots true airspeed (120KTA) regime. In the test rig, gear sets were tested until damage initiated and progressed while gear health monitoring data and operational parameters were measured and tooth damage progression documented. For the rig tests, the gear speed was maintained at 3500RPM, a one hour run-in was performed at 4000 in-lb gear torque, than the torque was increased to 8000 in-lbs. The HUMS gear condition indicator data evaluated included Figure of Merit 4 (FM4), Root Mean Square (RMS) or Diagnostic Algorithm 1(DA1), + 3 Sideband Index (SI3) and + 1 Sideband Index (SI1). These were selected based on their sensitivity in detecting contact fatigue damage modes from analytical, experimental and historical helicopter data. For this report, the helicopter dataset was reduced to fourteen tails and the test rig data set was reduced to eight tested gear sets. The damage modes compared were separated into three cases. For case one, both the gear and pinion showed signs of contact fatigue or scuffing damage. For case two, only the pinion showed signs of contact fatigue damage or scuffing. Case three was limited to the gear tests when scuffing occurred immediately after the gear run-in. Results of this investigation highlighted the importance of understanding the complete monitored systems, for both the helicopter and test rig, before interpreting health monitoring data. Further work is required to better define these two systems that include better state awareness of the fielded systems, new sensing technologies, new experimental methods or models that quantify the effect of system design on CI response and new methods for setting thresholds that take into consideration the variance of each system.
The diagnostics capability of micro-electro-mechanical systems (MEMS) based rotating accelerometer sensors in detecting gear tooth crack failures in helicopter main-rotor transmissions was evaluated. MEMS sensors were installed on a pre-notched OH-58C spiral-bevel pinion gear. Endurance tests were performed and the gear was run to tooth fracture failure. Results from the MEMS sensor were compared to conventional accelerometers mounted on the transmission housing. Most of the four stationary accelerometers mounted on the gear box housing and most of the CI's used gave indications of failure at the end of the test. The MEMS system performed well and lasted the entire test. All MEMS accelerometers gave an indication of failure at the end of the test. The MEMS systems performed as well, if not better, than the stationary accelerometers mounted on the gear box housing with regards to gear tooth fault detection. For both the MEMS sensors and stationary sensors, the fault detection time was not much sooner than the actual tooth fracture time. The MEMS sensor spectrum data showed large first order shaft frequency sidebands due to the measurement rotating frame of reference. The method of constructing a pseudo tach signal from periodic characteristics of the vibration data was successful in deriving a TSA signal without an actual tach and proved as an effective way to improve fault detection for the MEMS.