SmartIO: An End-to-EndWorkflow for Runtime I/O Prediction and Cross-Layer Optimization of HPC Systems
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We present R2U2, a novel framework for runtime monitoring of security properties and diagnosing of security threats on-board Unmanned Aerial Systems (UAS). R2U2, implemented in FPGA hardware, is a real-time, REALIZABLE, RESPONSIVE, UNOBTRUSIVE Unit for security threat detection. R2U2 is designed to continuously monitor inputs from the GPS and the ground control station, sensor readings, actuator outputs, and flight software status. By simultaneously monitoring and performing statistical reasoning, attack patterns and post-attack discrepancies in the UAS behavior can be detected. R2U2 uses runtime observer pairs for linear and metric temporal logics for property monitoring and Bayesian networks for diagnosis of security threats. We discuss the design and implementation that now enables R2U2 to handle security threats and present simulation results of several attack scenarios on the NASA DragonEye UAS.
In this paper, we propose a learning-based method utilizing the Soft Actor-Critic (SAC) algorithm to train a binary Support Vector Machine (SVM) classifier. This classifier is designed to identify valid input spaces in high-dimensional, highly constrained systems while minimizing the total runtime of offline simulations. The simulations adapt their runtime based on the likelihood that a given training input will be informative to the classifier. Furthermore, we introduce a method for using the trained SAC model to predict whether a desired system input is likely to violate constraints, along with a technique to adjust the input as necessary. Additionally, we explore the potential of this model to detect faults or adversarial attacks within the system. The effectiveness of our approach is demonstrated through various simulations of challenging classification problems and a constrained quadrotor model.
Runtime Verification is a light-weight approach to systems verification, where actual executions of a system are processed and analyzed using rigorous techniques. In this paper we shall narrow the term’s definition to represent the commonly studied variant consisting of verifying that a single system execution conforms to a specification written in a formal specification language. Runtime verification (in this sense) can be used for writing test oracles during testing when the system is too complex for full formal verification, or it can be used during deployment of the system as part of a fault protection strategy, where corrective actions may be taken in case the specification is violated. Specification languages for runtime verification appear to differ from for example temporal logics applied in model checking, in part due to the focus on monitoring of events that carry data, and specifically due to the desire to relate data values existing at different time points, resulting in new challenges in both the complexity of the monitoring approach and the expressiveness of languages. Over the recent years, numerous runtime verification specification languages have emerged, each with its different features and levels of expressiveness and usability. This paper presents an overview and a discussion of this design space.
Runtime verification is the discipline of analyzing program executions using rigorous methods. The discipline covers such topics as specification-based monitoring, where single executions are checked against formal specifications; predictive runtime analysis, where properties about a system are predicted/inferred from single (good) executions; specification mining from execution traces; visualization of execution traces; and to be fully general: computation of any interesting information from execution traces. Finally, runtime verification also includes fault protection, where monitors actively protect a running system against errors. The paper is written as a response to the ‘Test of Time Award’ attributed to the authors for their 2001 paper [45]. The present paper provides a brief overview of what lead to that paper, what has happened since, and some perspectives on the future of the field.
An important issue that must be faced while introducing Ada into the real time world is efficient and prodictable runtime behavior. One of the most effective methods employed during the traditional design of a real time system is the cyclic executive. The role cyclic scheduling might play in an Ada application in terms of currently available implementations and in terms of implementations that might be developed especially to support real time system development is examined. The cyclic executive solves many of the problems faced by real time designers, resulting in a system for which it is relatively easy to achieve approporiate timing behavior. Unfortunately a cyclic executive carries with it a very high maintenance penalty over the lifetime of the software that is schedules. Additionally, these cyclic systems tend to be quite fragil when any aspect of the system changes. The findings are presented of an ongoing SofTech investigation into Ada methods for real time system development. The topics covered include a description of the costs involved in using cyclic schedulers, the sources of these costs, and measures for future systems to avoid these costs without giving up the runtime performance of a cyclic system.
BiblioTech computer program originally developed to capture information about continuing advanced technological developments in fields applicable to spacecraft systems and subsystems in data-base files called "libraries." No element of design restricts BiblioTech from addressing other technological fields beyond specialized area of spacecraft technology. Also contains contact information to enable user to gain more information about technologies in library. Available only as executable code for use with 4th DIMENSION Runtime on Macintosh computers running System 6.0.7 or later.
Runtime assurance is a control framework where a complex controller operates under the observation of a monitor. If the monitor detects the controller exhibiting undesirable behavior, control is passed off to a trusted controller until a desirable state is regained. The runtime assurance architecture provides a layer of assurance to the system being controlled, but special care must be taken that the resulting overall system, consisting of the monitors and controllers, is behaving as intended. This talk aims to formally model and reason about runtime assurance-equipped systems as hybrid programs- which are models that consist of both discrete and continuous components. Using the verification tool Plaidypvs, safety properties of some examples involving RTA architectures is shown.
For unmanned aerial systems (UAS) to be successfully deployed and integrated within the national airspace, it is imperative that they possess the capability to effectively complete their missions without compromising the safety of other aircraft, as well as persons and property on the ground. This necessity creates a natural requirement for UAS that can respond to uncertain environmental conditions and emergent failures in real-time, with robustness and resilience close enough to those of manned systems. We introduce a system that meets this requirement with the design of a real-time onboard system health management (SHM) capability to continuously monitor sensors, software, and hardware components. This system can detect and diagnose failures and violations of safety or performance rules during the flight of a UAS. Our approach to SHM is three-pronged, providing: (1) real-time monitoring of sensor and software signals; (2) signal analysis, preprocessing, and advanced on-the-fly temporal and Bayesian probabilistic fault diagnosis; and (3) an unobtrusive, lightweight, read-only, low-power realization using Field Programmable Gate Arrays (FPGAs) that avoids overburdening limited computing resources or costly re-certification of flight software. We call this approach rt-R2U2, a name derived from its requirements. Our implementation provides a novel approach of combining modular building blocks, integrating responsive runtime monitoring of temporal logic system safety requirements with model-based diagnosis and Bayesian network-based probabilistic analysis. We demonstrate this approach using actual flight data from the NASA Swift UAS.
Unmanned aerial systems (UASs) can only be deployed if they can effectively complete their missions and respond to failures and uncertain environmental conditions while maintaining safety with respect to other aircraft as well as humans and property on the ground. In this paper, we design a real-time, on-board system health management (SHM) capability to continuously monitor sensors, software, and hardware components for detection and diagnosis of failures and violations of safety or performance rules during the flight of a UAS. Our approach to SHM is three-pronged, providing: (1) real-time monitoring of sensor and/or software signals; (2) signal analysis, preprocessing, and advanced on the- fly temporal and Bayesian probabilistic fault diagnosis; (3) an unobtrusive, lightweight, read-only, low-power realization using Field Programmable Gate Arrays (FPGAs) that avoids overburdening limited computing resources or costly re-certification of flight software due to instrumentation. Our implementation provides a novel approach of combining modular building blocks, integrating responsive runtime monitoring of temporal logic system safety requirements with model-based diagnosis and Bayesian network-based probabilistic analysis. We demonstrate this approach using actual data from the NASA Swift UAS, an experimental all-electric aircraft.
The 2023 Performance Monitoring and Construction Completion Report (PM-CCR) presents the findings, observations, and results for Air Sparging (AS) operations and expansion activities, as well as sitewide groundwater monitoring for Launch Complex 39B (LC39B), Solid Waste Management Unit (SWMU) 009, at Kennedy Space Center (KSC), Florida. The reporting period for activities covered under this PM-CCR is from January 1, 2023, to December 31, 2023. At LC39B, AS operations began in 2017 in the area west of the launch pad, in the liquid oxygen (LOX) tank area located northwest of the launch pad, and in an area outside of the perimeter fence to protect nearby Outstanding Florida Waters (OFW). The LC39B AS system was installed with 279 AS wells to depths ranging from 23 to 60 feet below land surface (bls), including the sump, correlating to top of screen depths ranging from 20 to 57 feet bls. In December 2022, a total of 22 AS wells were abandoned to support launch pad crane operations, and in November 2023, the system was expanded with five additional AS wells installed to 13 or 17 feet bls near the LOX tank. The remedial objective of the LC39B AS Interim Measure (IM) is to actively decrease concentrations of contaminants of concern (COCs) in groundwater, specifically trichloroethene (TCE), cis-1,2-Dichloroethene (cDCE), and vinyl chloride (VC), to less than their respective Natural Attenuation Default Concentrations (NADCs), so LC39B can transition into a Long-Term Monitoring (LTM) program. This PM-CCR presents the following information for LC39B: • AS system operations and maintenance (O&M) (Year 7 of operation) from January 2023 to December 2023, to include AS trailer relocation in March 2023 and subsequent replacement and re-start in June 2023. • Construction completion details for AS system expansion, which included installation of five new AS wells and one new monitoring well in November 2023. As part of expansion activities, soil samples were also collected for petroleum analysis; no exceedances were identified, and no further investigation for petroleum is warranted. • Performance monitoring results for groundwater sampling events conducted in May/June 2023 (30 wells) and November 2023 (31 wells) in the AS IM area and in the Low Concentration Plume (LCP) areas located north and east of the launch pad for volatile organic compound (VOC) analysis. • Sampling results for one monitoring well, LOX-IW0012S, which is sampled for aluminum on an annual basis (May/June 2023). This well was resampled in November 2023 for both total and dissolved aluminum. Due to a communication error with the laboratory, the May/June 2023 sample was analyzed for total aluminum only. • Groundwater sampling results for per- and polyfluoroalkyl substances (PFAS) collected from seven monitoring wells during the May/June 2023 event to further investigate the Former Sewage Treatment Plant #6 and Percolation Pond area, west of the launch pad. O&M and performance monitoring results show that the AS system at LC39B is operating as designed and meeting performance criteria. Overall runtime was 45 percent (%) during the reporting period (January to December), but the operational runtime was 78% during the timeframe when the system could run (June to December). The most significant downtime contributor was post-launch crane operations following the Artemis launch on November 16, 2022, which lasted until June 2023. During that timeframe, the AS trailer at LC39B was relocated to another KSC remediation site (Wilson Corners) and was subsequently replaced with the AS trailer from the Paint & Oil Locker (POL) remediation site at KSC to resume AS system operations. Performance monitoring results in the AS IM and LCP areas continue to show reduction in COC concentrations over time when compared to baseline levels. In 2023, only one monitoring well (MW0048) detected a COC exceeding its NADC (VC at 740 micrograms per liter [µg/L]), which marks the baseline result for this new well installed during system expansion. Across the rest of the site, VC concentrations have declined or remained stable during the 2023 sampling events. Excluding MW0048, the highest VC result in 2023 was during the May/June sampling event with a concentration of 63 µg/L at MW0032, which is located near MW0048 and the AS expansion area by the LOX tank. TCE was detected in select monitoring wells in the IM area in 2023, but only two locations exceeded the State of Florida Groundwater Cleanup Target Level (GCTL): MW0032 (21 µg/L in May/June 2023 and 9.1 µg/L in November 2023) and MW0036 (5.0 µg/L in November 2023). MW0036 is also located near the LOX tank, on the north side, where the AS system is still operational (Zone Z4). cDCE and trans-1,2-dichloroethene concentrations were less than laboratory method detection limits or their respective GCTLs in all wells sampled in 2023. Near the OFW located northwest of the launch complex, all COC concentrations were less than laboratory method detection limits from monitoring wells (MW0039, MW0040, and LOXTA0002S) sampled in 2023. Aluminum results from LOX-IW0012S, which has been sampled routinely since 2006, detected a total aluminum concentration of 3,900 µg/L during the May/June 2023 sampling event. Results from the November 2023 event detected 5,700 µg/L for total aluminum and 5,500 µg/L for dissolved aluminum. These concentrations slightly decreased from the previous year but remain relatively consistent with historical detections. Aluminum will continue to be sampled on an annual basis at this well as results still exceed the GCTL of 200 µg/L and the Upper Limit of the KSC Background Concentration of 280 µg/L. PFAS results detected nine different PFAS compounds (out of 32 analyzed) from seven wells sampled. Two PFAS compounds, perfluorooctanesulfonic acid (PFOS) and perfluorooctanoic acid (PFOA), currently have FDEP Provisional GCTLs of 70 nanograms per liter (ng/L). All seven samples collected resulted in concentrations less than the FDEP Provisional GCTLs for both PFAS compounds; no exceedances were observed. PFOS and PFOA also have assigned United States Environmental Protection Agency (USEPA) Maximum Contaminant Levels (MCLs) of 4 nanograms per liter (ng/L). None of the PFOS results exceeded the USEPA MCLs. PFOA was detected in two samples above the USEPA MCL at concentrations of 5.8 ng/L (ECS-IW0009I) and 5.5 ng/L (ECS-IW0009S). Three other PFAS compounds, perfluorohexanesulfonic acid (PFHxS), perfluoro-n-nonanoic acid (PFNA), and hexafluoropropylene oxide dimer acid (GenX), currently have USEPA MCLs of 10 ng/L. PFHxS, PFNA, and GenX were not detected at concentrations greater than their respective USEPA MCLs in any of the seven wells. PFAS compounds without FDEP Provisional GCTLs or USEPA MCLs were screened against USEPA RSLs. No other detections exceeded their respective USEPA RSLs. Additional PFAS sampling will be conducted as part of a future PFAS Site Assessment. Based on O&M activities and performance monitoring, the following is recommended for LC39B: • Continue with Year 8 AS system operation within Zone Z4, which includes the AS expansion area. Zone Z3, which has been off since 2018, should remain off as no rebound has been observed. Zones Z1 and Z2, which were turned off at the end of 2022, will remain shut down as monitoring well results have consistently been below GCTLs or have low-level detections with stable or decreasing trends (Meeting Minute 2402-M10, Decision 2402-D30). • Continue with performance monitoring in 2024 with the same monitoring well network as 2023, except with the addition of MW0048 in both semi-annual events. Baseline concentrations for this well were collected during the November 2023 performance monitoring event. Semi-annual sampling should be planned for the May 2024 and November 2024 timeframes (Meeting Minute 2402-M10, Decision 2402-D31). • Continue sampling monitoring well, LOX-IW0012S, for aluminum (total and dissolved) on an annual basis in May 2024. It is also recommended to re-develop this well prior to the next sampling event (Meeting Minute 2402-M10, Decision 2402-D32). The above recommendations for LC39B were presented at the February 2024 KSCRT Meeting, with Team consensus reached on the path forward. The contents of this PM-CCR were also presented at this meeting.
In this technical report we provide information on the use of the NASA Formal RequirementsElicitation Tool (FRET) to create requirements for a Lift Plus Cruise (LPC) aircraft case study. Furthermore, we provide details on using FRET to translate these requirements into an appropriate format for the Copilot tool, enabling their usage to perform runtime verification on a synthesized LPC system.
This report examines the software supply chain security posture of mobile applications developed for consumer whole-house battery and energy-management products. While these applications are not currently integrated with critical infrastructure, their growing role in connected energy domain spaces underscores the importance of understanding the external dependencies, permission structures, and runtime behaviors that could introduce systemic risk; particularly, if adoption expands into more critical environments.
This presentation goes over some of the tools developed at NASA Ames in the Robust Software Engineering group for the assurance and certification of autonomous systems. The research themes presented include improving safety and risk assessment as early as possible in the lifecycle, elicitation and formalization of requirements to facilitate traceability throughout the lifecycle, especially when formal methods are used, algorithms, tools and techniques for the V&V of ML-enabled systems, advanced testing, use of runtime monitoring to ease use of untrusted components, and contribution to draft regulatory standards and assistance in producing and presenting certification evidences.
Trusted Space Autonomy is challenging in that space systems are complex artifacts deployed in a high stakes environment with complicated operational settings. Thus far these challenges have been met using the full arsenal of tools: formal methods, informal methods, testing, runtime techniques, and operations processes. Using examples from previous deployments of autonomy to the Remote Agent on DS-1, Autonomous Sciencecraft on EO-1, WATCH on MER, IPEX, AEGIS on MER, MSL, and M2020, and the M2020 Onboard planner, we discuss how each of these approaches have been used to enable successful deployment of autonomy. We next focus on relatively limited use of formal methods (both prior to deployment and runtime methods). From the needs perspective, formal methods represent the best chance for reliable autonomy as testing, informal methods, and operations accommodations do not scale well with increasing complexity of the autonomous system. However from the practice perspective, formal methods have been limited in their application due to difficulty in eliciting formal specifications and challenges in representing complex constraints such as metric time and resources. We discuss some of these challenges as well as the opportunity to extend formal and informal methods into runtime validation systems.
Cybersecurity is one of the largest concerns in modern computing, impacting and dictating how governments, private corporations, and individuals interact with and live in an increasingly digital world. The NSA has recently released a memo [ 1] on “Software Memory Safety” where they highlight that both Microsoft and Google have stated around 70% of software vulnerabilities were due to memory safety issues. Although languages such as C and C++ provide freedom and flexibility with memory management, guaran- teeing safety falls mostly on the developer. The NSA recommends using “memory safe” languages whenever possible. In this paper we introduce Lamellar, an asynchronous tasking and PGAS HPC runtime written in Rust, one such "memory safe" language. We describe the entire Lamellar stack, from network interfaces to high- level abstractions such as distributed LamellarArrays and Active Messages. We conclude by showing comparable performance to legacy PGAS runtimes (e.g. OpenSHMEM) on a subset of the BALE kernel suite while maintaining strong memory safety principles.
Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.
Runtime Assurance (RTA) is a design-time architecture for safety-critical systems where an internal monitor acts upon detecting a violation of a property. The simplex architecture is an instance of RTA, where the action taken is to hand control of the overall system to a trusted controller when an untrusted one violates a safety property. Simplex RTA is emerging as a method for allowing AI/ML and other unverified software to be integrated into safety-critical applications like aircraft. To this end, the American Society for Testing and Materials (ASTM) and NASA have each published guidelines on the use of RTA in such systems. In the simplex RTA framework, a system has an advanced controller (AC) and a reversionary controller (RC). The system is allowed to operate with the AC until a runtime monitor detects that some property has been violated and then the RC takes over. Assuming that the sample rate of the monitor will detect improper functioning with enough time for the RC to correct the impending problem, and that the RC is trusted, the system will operate as intended. This use of the simplex RTA framework can allow for the integration of untrusted, but possibly more performant, controllers in a safe way. This paper presents a formalization of a simplex RTA framework in the Prototype Verification System (PVS) theorem prover using an embedding of differential dynamic logic (DDL) called Plaidypvs. A novel feature of this framework is that it can be instantiated at different levels of abstraction. This feature allows for the formal verification of a system with an untrusted black box component, such as an AI/ML controller. This paper does not address the many difficulties in deploying RTA in an industrial-level system. Instead, the focus is on the formal verification of the simplex RTA framework in the language of hybrid programs. Hybrid programs are programs that include both discrete and continuous dynamics and can be used to model complex cyber-physical systems. Plaidypvs is a tool that enables formalization of hybrid programs in the PVS theorem prover. Plaidypvs enables the verification of the general simplex RTA framework and then, by specializing some components of the hybrid program, verifying instances of the framework while treating the untrusted component as a black box. A selection of Unmanned Aircraft Systems (UAS) operations are shown as instances of the general RTA framework in PVS. This offers the benefit of design time verification of relevant safety properties to the system, and it also gives requirements on the sample rate of sensors that determine the time interval in which the ‘switch’ property of the RTA framework is checked.