Search NASA⌕ Search

SEARCH · Search NASA

Results for “continuous integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Overcoming Challenges to Continuous Integration in HPC

Continuous integration (CI) has become a ubiquitous practice in modern software development, with major code hosting services offering free automation on popular platforms. CI offers major benefits, as it enables detecting bugs in code prior to committing changes. While high-performance computing (HPC) research relies heavily on software, HPC machines are not considered “common” platforms. This presents several challenges that hinder the adoption of CI in HPC environments, making it difficult to maintain bug-free HPC projects, and resulting in adverse effects on the research community. Here we explore the challenges that impede HPC CI, such as hardware diversity, security, isolation, administrative policies, and non-standard authentication, environments, and job submission mechanisms. We propose several solutions that could enhance the quality of HPC software and the experience of developers. Implementing these solutions would require significant changes at HPC centers, but if these changes are made, it would ultimately enable faster and better science.

97 MATHEMATICS AND COMPUTING↗

Operational Intelligence in the ATLAS Continuous Integration System

Describes the role of the ATLAS Continuous Integration (CI) System in the ATLAS offline software development infrastructure • Outlines the CI system components and processes • Details Operational Intelligence techniques to accelerate CI jobs and lower operating costs • Explains the Directed Acyclic Graph (DAG) approach in CI pipelines • Reports achieved improvements

97 MATHEMATICS AND COMPUTING↗

BeeSwarm: Enabling Parallel Scaling Performance Measurement in Continuous Integration for HPC Applications

Testing is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using a multi-physics HPC application with Travis CI, GitLab CI and GitHub Actions while using ChameleonCloud and Google Compute Engine as the compute backends. Finally, our results show that BeeSwarm can be used for scalability and performance testing of HPC applications.

97 MATHEMATICS AND COMPUTING↗

Adoption of Test Driven Development and Continuous Integration for the Development of the Trick Simulation Toolkit

This paper describes the adoption of a Test Driven Development approach and a Continuous Integration System in the development of the Trick Simulation Toolkit, a generic simulation development environment for creating high fidelity training and engineering simulations at the NASA/Johnson Space Center and many other NASA facilities. It describes what was learned and the significant benefits seen, such as fast, thorough, and clear test feedback every time code is checked-in to the code repository. It also describes a system that encourages development of code that is much more flexible, maintainable, and reliable. The Trick Simulation Toolkit development environment provides a common architecture for user-defined simulations. Trick builds executable simulations using user-supplied simulation-definition files (S_define) and user supplied "model code". For each Trick-based simulation, Trick automatically provides job scheduling, checkpoint / restore, data-recording, interactive variable manipulation (variable server), and an input-processor. Also included are tools for plotting recorded data and various other supporting tools and libraries. Trick is written in C/C++ and Java and supports both Linux and MacOSX. Prior to adopting this new development approach, Trick testing consisted primarily of running a few large simulations, with the hope that their complexity and scale would exercise most of Trick's code and expose any recently introduced bugs. Unsurprising, this approach yielded inconsistent results. It was obvious that a more systematic, thorough approach was required. After seeing examples of some Java-based projects that used the JUnit test framework, similar test frameworks for C and C++ were sought. Several were found, all clearly inspired by JUnit. Googletest, a freely available Open source testing framework, was selected as the most appropriate and capable. The new approach was implemented while rewriting the Trick memory management component, to eliminate a fundamental design flaw. The benefits became obvious almost immediately, not just in the correctness of the individual functions and classes but also in the correctness and flexibility being added to the overall design. Creating code to be testable, and testing as it was created resulted not only in better working code, but also in better-organized, flexible, and readable (i.e., articulate) code. This was, in essence the Test-driven development (TDD) methodology created by Kent Beck. Seeing the benefits of Test Driven Development, other Trick components were refactored to make them more testable and tests were designed and implemented for them.

Penn, John M.↗

Creating Continuous Integration Infrastructure for Software Development on U.S. Department of Energy High-Performance Computing Systems

The Exascale Computing Project (ECP) software deployment effort developed and advanced DevOps capabilities. One goal was to enable robust continuous integration (CI) workflows that span the protected high performance computing (HPC) environments found within many of the Department of Energy’s (DOE) national laboratories. This article highlights several challenges encountered with enabling automation, such as charging models for CI jobs, and meeting individualized security requirements that revolve around strongly associating running code with a human identity. Here, it also describes how the Jacamar CI tool evolved to meet latter requirements and became a key aspect of the solutions currently offered. Derived from this experience, we offer a conceptual framework for understanding current and future CI challenges at DOE facilities and offer suggestions for long-term solutions.

97 MATHEMATICS AND COMPUTING↗

Integrating continuous atmospheric boundary layer and tower-based flux measurements to advance understanding of land-atmosphere interactions

The atmospheric boundary layer mediates the exchange of energy, matter, and momentum between the land surface and the free troposphere, integrating a range of physical, chemical, and biological processes and is defined as the lowest layer of the atmosphere (ranging from a few meters to 3 km). In this review, we investigate how continuous, automated observations of the atmospheric boundary layer can enhance the scientific value of co-located eddy covariance measurements of land-atmosphere fluxes of carbon, water, and energy, as are being made at FLUXNET sites worldwide. We highlight four key opportunities to integrate tower-based flux measurements with continuous, long-term atmospheric boundary layer measurements: (1) to interpret surface flux and atmospheric boundary layer exchange dynamics and feedbacks at flux tower sites, (2) to support flux footprint modelling, the interpretation of surface fluxes in heterogeneous and mountainous terrain, and quality control of eddy covariance flux measurements, (3) to support regional-scale modeling and upscaling of surface fluxes to continental scales, and (4) to quantify land-atmosphere coupling and validate its representation in Earth system models. Adding a suite of atmospheric boundary layer measurements to eddy covariance flux tower sites, and supporting the sharing of these data to tower networks, would allow the Earth science community to address new emerging research questions, better interpret ongoing flux tower measurements, and would present novel opportunities for collaborations between FLUXNET scientists and atmospheric and remote sensing scientists.

Manuel Helbig↗

Utilizing Testing Frameworks for Launch Control Systems Continuous Integration

Command and control software is an integral part of the launch procedure. The most important part of this type of software is its ability to communicate well with the user and relay information in a correctly formatted way such that the user can understand the data. There is a tool that aides the communication between the different parts of the system, and effectively, the user. This instrument is capable of taking several complex values and ensuring that they are correctly sorted into their distinctive message values and distributed properly among the different facets of the system. This tool will easily translate and publish the data inside of messages in the system to something that is readable and understandable. The tool also allows for transmission of the recorded data to the user, effectively ensuring the communication between different components of the system. As well as keeping track of messages and ensuring that the information contained within each of them reaches the correct location, this tool has the ability to keep track of its own statistics and determine how many messages passed in were erroneous and how many were successfully transmitted. It is able to check and see what the total message failure count is when an invalid message is given, as well as the number of different messages and their respective types passed into the tool. This tool is of great value to the new Space Launch System (SLS). As such, the tool must be thoroughly tested with test cases that, although improbable, are possible, where the tool may not function properly. Testing an interface this complex is necessary to ensure mission safety and create unlikely scenarios where the tool would work as intended, and stretch its limits to test that even under the most uncommon conditions it would still continue to function. This software will be an important part of the control system for the newest spacecraft which will fly deeper into space than humans have ever travelled. It will fly beyond the moon, into deep space to Mars and perhaps set the groundwork for a manned mission even further to create more opportunities for interplanetary and even interstellar travel by humans. This mission relies heavily on software and hardware to ensure the safety of the humans that will be on board and therefore must be checked, exhausting each and every different situation, such that there is not a doubt surrounding the well-being of the humans aboard the rocket. That is why testing is such an important part of the mission. It provides evidence that the systems aboard the rocket and on the launch pad are safe.

Unit Testing↗

Rock Physics-Based Data Assimilation of Integrated Continuous Active-Source Seismic and Pressure Monitoring Data during Geological Carbon Storage

Summary There has been substantial controversy concerning the role of geological carbon storage (GCS) in sequestering anthropogenic carbon emissions to mitigate climate change and global warming. Arguments center on the inability to monitor a geological storage site precisely and continuously, especially highlighting the associated costs and spatiotemporal trade-offs when using conventional subsurface monitoring techniques (well logs, core samples, chemical tracers, and 4D seismics). Active surveillance of GCS sites is essential for managing and mitigating potential leaks but is also required by regulation. With the goal of enhancing the monitoring capability at GCS sites, we present a rock physics-based joint data assimilation model to study a popular GCS site at Cranfield, Mississippi, USA. Synthetic continuous active-source seismic monitoring (CASSM) data (in the form of Vp and Qp measurements) and wellbore pressure monitoring data are assimilated with an ensemble of reservoir realizations to monitor gas saturation and reservoir pressure changes over a period of 100 years. Synthetic seismic attributes are generated using rock physics models (RPMs) and wellbore pressure monitoring data are extracted from the ground truth. Two assimilation methods, ensemble Kalman filter (EnKF) and ensemble Kalman smoother (EnKS), are tested in an observation system simulation experiment (OSSE) environment to assess the prediction accuracy of the individual and composite observation systems. The joint monitoring system achieves more accurate estimates of gas saturation and pressure, across the time span from start of injection to end of forecast, as compared to a single type of monitoring tool and irrespective of data assimilation algorithm choice. These results indicate that jointly assimilated data from two types of sensors (in this case, crosswell seismic and downhole pressure) may lead to a more risk-reducing monitoring design. One would expect that more data, vis-à-vis inclusion of a new sensor type, will improve the accuracy of any GCS monitoring system. However, from a practical standpoint, one important question is whether such a gain in accuracy is worth the additional cost associated with the new sensor. This paper focuses on quantifying the gain in accuracy, such that a practitioner can answer this question.

Engineering↗

Continuous integration data-driven platform of industrial-scale subsurface storage for real-time analytics

This project helped address the growing need for efficient and scalable models to support geological carbon and energy storage, which are crucial for achieving net-zero emissions. Traditionally accurate high-fidelity numerical models have been used to simulate relevant storage processes under a handful of processes, however such models are computationally demanding, making uncertainty quantification impractical. Consequently, we first developed a machine learning framework, based on Graph Neural Operators (GNOs), to improving the accuracy of model predictions for a fixed computational budget. We then developed an Ensemble of Improved Neural Operators (ENO), which uses bagging and Monte Carlo dropout techniques, to further improve prediction accuracy. Lastly, we developed the way to explain progressive transfer learning methods to reduce the amount of training data and computational cost of training (i.e., reduce trainable parameters) when using our models for multiple storage sites. Our numerical investigation, which used real-world case studies, demonstrated that our framework can significantly improve the safety and efficiency of geological storage operations, with potential applications in other domains such as geothermal reservoirs and climate modeling.

54 ENVIRONMENTAL SCIENCES↗

Continuous-Integration Laser Energy Lidar Monitor

This circuit design implements an integrator intended to allow digitization of the energy output of a pulsed laser, or the energy of a received pulse of laser light. It integrates the output of a detector upon which the laser light is incident. The integration is performed constantly, either by means of an active integrator, or by passive components.

Karsh, Jeremy↗

Controls at the Fermilab PIP-II Superconducting Linac

PIP-II is an 800 MeV superconducting RF linac under development at Fermilab. As the new first stage in our accelerator chain, it will deliver high-power beam to multiple experiments simultaneously and thus drive Fermilab’s particle physics program for years to come. In a pivot for Fermilab, controls for PIP-II are based on EPICS instead of ACNET, the legacy control system for accelerators at the lab. This paper discusses the status of the EPICS controls work for PIP-II. We describe the EPICS tools selected for our system and the experience of operators new to EPICS. We introduce our continuous integration / continuous development environment. We also describe some efforts at cooperation between EPICS and ACNET, or efforts to move towards a unified interface that can apply to both control systems.

43 PARTICLE ACCELERATORS↗

Controls at the Fermilab PIP-II Superconducting Linac

PIP-II is an 800 MeV superconducting RF linear accelerator under development at Fermilab. As the new first stage in our accelerator chain, it will deliver high-power beam to multiple experiments simultaneously and thus drive Fermilab's particle physics program for years to come. In a pivot for Fermilab, controls for PIP-II are based on EPICS instead of ACNET, the legacy control system for accelerators at the lab. This paper discusses the status of the EPICS controls work for PIP-II. We describe the EPICS tools selected for our system and the experience of operators new to EPICS. We introduce our continuous integration / continuous development environment. We also describe some efforts at cooperation between EPICS and ACNET, or efforts to move towards a unified interface that can apply to both control systems.

43 PARTICLE ACCELERATORS↗

Containerization of Phase-2 Tracker Data Acquisition and Control Framework

The CMS Experiment has started an extensive upgrade program in the context of the High-Luminosity phase of the LHC (Phase-2). In order to cope with the highly demanding High-Luminosity conditions, CMS will need a completely new inner and outer tracking detectors. On top of R&D development, a Data AcQuisition (DAQ) software is being developed along with different applications to control, monitor and validate the newly produced modules of the future tracker. As more and more developers are getting involved to maintain all those codes, an environment where software and applications can work independently of the host machine operating system is crucial. The Phase-2 Tracker group decided to make use of Docker as containerization solution for its framework. This poster describes the containerization of DAQ software/applications utilized to test and validate the Tracker modules. This includes continuous integration and continuous deployment (so-called CI/CD) and running GUI applications inside containers.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Determining a Method of Enabling and Disabling the Integral Torque in the SDO Science and Inertial Mode Controllers

During design of the SDO Science and Inertial mode PID controllers, the decision was made to disable the integral torque whenever system stability was in question. Three different schemes were developed to determine when to disable or enable the integral torque, and a trade study was performed to determine which scheme to implement. The trade study compared complexity of the control logic, risk of not reenabling the integral gain in time to reject steady-state error, and the amount of integral torque space used. The first scheme calculated a simplified Routh criterion to determine when to disable the integral torque. The second scheme calculates the PD part of the torque and looked to see if that torque would cause actuator saturation. If so, only the PD torque is used. If not, the integral torque is added. Finally, the third scheme compares the attitude and rate errors to limits and disables the integral torque if either of the errors is greater than the limit. Based on the trade study results, the third scheme was selected. Once it was decided when to disable the integral torque, analysis was performed to determine how to disable the integral torque and whether or not to reset the integrator once the integral torque was reenabled. Three ways to disable the integral torque were investigated: zero the input into the integrator, which causes the integral part of the PID control torque to be held constant; zero the integral torque directly but allow the integrator to continue integrating; or zero the integral torque directly and reset the integrator on integral torque reactivation. The analysis looked at complexity of the control logic, slew time plus settling time between each calibration maneuver step, and ability to reject steady-state error. Based on the results of the analysis, the decision was made to zero the input into the integrator without resetting it. Throughout the analysis, a high fidelity simulation was used to test the various implementation methods.

Vess, Melissa F.↗