Search NASA⌕ Search

SEARCH · Search NASA

Results for “continuous integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Overcoming Challenges to Continuous Integration in HPC

Continuous integration (CI) has become a ubiquitous practice in modern software development, with major code hosting services offering free automation on popular platforms. CI offers major benefits, as it enables detecting bugs in code prior to committing changes. While high-performance computing (HPC) research relies heavily on software, HPC machines are not considered “common” platforms. This presents several challenges that hinder the adoption of CI in HPC environments, making it difficult to maintain bug-free HPC projects, and resulting in adverse effects on the research community. Here we explore the challenges that impede HPC CI, such as hardware diversity, security, isolation, administrative policies, and non-standard authentication, environments, and job submission mechanisms. We propose several solutions that could enhance the quality of HPC software and the experience of developers. Implementing these solutions would require significant changes at HPC centers, but if these changes are made, it would ultimately enable faster and better science.

97 MATHEMATICS AND COMPUTING↗

Operational Intelligence in the ATLAS Continuous Integration System

Describes the role of the ATLAS Continuous Integration (CI) System in the ATLAS offline software development infrastructure • Outlines the CI system components and processes • Details Operational Intelligence techniques to accelerate CI jobs and lower operating costs • Explains the Directed Acyclic Graph (DAG) approach in CI pipelines • Reports achieved improvements

97 MATHEMATICS AND COMPUTING↗

BeeSwarm: Enabling Parallel Scaling Performance Measurement in Continuous Integration for HPC Applications

Testing is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using a multi-physics HPC application with Travis CI, GitLab CI and GitHub Actions while using ChameleonCloud and Google Compute Engine as the compute backends. Finally, our results show that BeeSwarm can be used for scalability and performance testing of HPC applications.

97 MATHEMATICS AND COMPUTING↗

Creating Continuous Integration Infrastructure for Software Development on U.S. Department of Energy High-Performance Computing Systems

The Exascale Computing Project (ECP) software deployment effort developed and advanced DevOps capabilities. One goal was to enable robust continuous integration (CI) workflows that span the protected high performance computing (HPC) environments found within many of the Department of Energy’s (DOE) national laboratories. This article highlights several challenges encountered with enabling automation, such as charging models for CI jobs, and meeting individualized security requirements that revolve around strongly associating running code with a human identity. Here, it also describes how the Jacamar CI tool evolved to meet latter requirements and became a key aspect of the solutions currently offered. Derived from this experience, we offer a conceptual framework for understanding current and future CI challenges at DOE facilities and offer suggestions for long-term solutions.

97 MATHEMATICS AND COMPUTING↗

Integrating continuous atmospheric boundary layer and tower-based flux measurements to advance understanding of land-atmosphere interactions

The atmospheric boundary layer mediates the exchange of energy, matter, and momentum between the land surface and the free troposphere, integrating a range of physical, chemical, and biological processes and is defined as the lowest layer of the atmosphere (ranging from a few meters to 3 km). In this review, we investigate how continuous, automated observations of the atmospheric boundary layer can enhance the scientific value of co-located eddy covariance measurements of land-atmosphere fluxes of carbon, water, and energy, as are being made at FLUXNET sites worldwide. We highlight four key opportunities to integrate tower-based flux measurements with continuous, long-term atmospheric boundary layer measurements: (1) to interpret surface flux and atmospheric boundary layer exchange dynamics and feedbacks at flux tower sites, (2) to support flux footprint modelling, the interpretation of surface fluxes in heterogeneous terrain, and quality control of eddy covariance flux measurements, (3) to support regional-scale modeling and upscaling of surface fluxes to continental scales, and (4) to quantify land-atmosphere coupling and validate its representation in Earth system models. Adding a suite of atmospheric boundary layer measurements to eddy covariance flux tower sites, and supporting the sharing of these data to tower networks, would allow the Earth science community to address new emerging research questions, better interpret ongoing flux tower measurements, and would present novel opportunities for collaborations between FLUXNET scientists and atmospheric and remote sensing scientists.

54 ENVIRONMENTAL SCIENCES↗

Rock Physics-Based Data Assimilation of Integrated Continuous Active-Source Seismic and Pressure Monitoring Data during Geological Carbon Storage

Summary There has been substantial controversy concerning the role of geological carbon storage (GCS) in sequestering anthropogenic carbon emissions to mitigate climate change and global warming. Arguments center on the inability to monitor a geological storage site precisely and continuously, especially highlighting the associated costs and spatiotemporal trade-offs when using conventional subsurface monitoring techniques (well logs, core samples, chemical tracers, and 4D seismics). Active surveillance of GCS sites is essential for managing and mitigating potential leaks but is also required by regulation. With the goal of enhancing the monitoring capability at GCS sites, we present a rock physics-based joint data assimilation model to study a popular GCS site at Cranfield, Mississippi, USA. Synthetic continuous active-source seismic monitoring (CASSM) data (in the form of Vp and Qp measurements) and wellbore pressure monitoring data are assimilated with an ensemble of reservoir realizations to monitor gas saturation and reservoir pressure changes over a period of 100 years. Synthetic seismic attributes are generated using rock physics models (RPMs) and wellbore pressure monitoring data are extracted from the ground truth. Two assimilation methods, ensemble Kalman filter (EnKF) and ensemble Kalman smoother (EnKS), are tested in an observation system simulation experiment (OSSE) environment to assess the prediction accuracy of the individual and composite observation systems. The joint monitoring system achieves more accurate estimates of gas saturation and pressure, across the time span from start of injection to end of forecast, as compared to a single type of monitoring tool and irrespective of data assimilation algorithm choice. These results indicate that jointly assimilated data from two types of sensors (in this case, crosswell seismic and downhole pressure) may lead to a more risk-reducing monitoring design. One would expect that more data, vis-à-vis inclusion of a new sensor type, will improve the accuracy of any GCS monitoring system. However, from a practical standpoint, one important question is whether such a gain in accuracy is worth the additional cost associated with the new sensor. This paper focuses on quantifying the gain in accuracy, such that a practitioner can answer this question.

Engineering↗

Continuous integration data-driven platform of industrial-scale subsurface storage for real-time analytics

This project helped address the growing need for efficient and scalable models to support geological carbon and energy storage, which are crucial for achieving net-zero emissions. Traditionally accurate high-fidelity numerical models have been used to simulate relevant storage processes under a handful of processes, however such models are computationally demanding, making uncertainty quantification impractical. Consequently, we first developed a machine learning framework, based on Graph Neural Operators (GNOs), to improving the accuracy of model predictions for a fixed computational budget. We then developed an Ensemble of Improved Neural Operators (ENO), which uses bagging and Monte Carlo dropout techniques, to further improve prediction accuracy. Lastly, we developed the way to explain progressive transfer learning methods to reduce the amount of training data and computational cost of training (i.e., reduce trainable parameters) when using our models for multiple storage sites. Our numerical investigation, which used real-world case studies, demonstrated that our framework can significantly improve the safety and efficiency of geological storage operations, with potential applications in other domains such as geothermal reservoirs and climate modeling.

54 ENVIRONMENTAL SCIENCES↗

Controls at the Fermilab PIP-II Superconducting Linac

PIP-II is an 800 MeV superconducting RF linac under development at Fermilab. As the new first stage in our accelerator chain, it will deliver high-power beam to multiple experiments simultaneously and thus drive Fermilab’s particle physics program for years to come. In a pivot for Fermilab, controls for PIP-II are based on EPICS instead of ACNET, the legacy control system for accelerators at the lab. This paper discusses the status of the EPICS controls work for PIP-II. We describe the EPICS tools selected for our system and the experience of operators new to EPICS. We introduce our continuous integration / continuous development environment. We also describe some efforts at cooperation between EPICS and ACNET, or efforts to move towards a unified interface that can apply to both control systems.

43 PARTICLE ACCELERATORS↗

Controls at the Fermilab PIP-II Superconducting Linac

PIP-II is an 800 MeV superconducting RF linear accelerator under development at Fermilab. As the new first stage in our accelerator chain, it will deliver high-power beam to multiple experiments simultaneously and thus drive Fermilab's particle physics program for years to come. In a pivot for Fermilab, controls for PIP-II are based on EPICS instead of ACNET, the legacy control system for accelerators at the lab. This paper discusses the status of the EPICS controls work for PIP-II. We describe the EPICS tools selected for our system and the experience of operators new to EPICS. We introduce our continuous integration / continuous development environment. We also describe some efforts at cooperation between EPICS and ACNET, or efforts to move towards a unified interface that can apply to both control systems.

43 PARTICLE ACCELERATORS↗

Containerization of Phase-2 Tracker Data Acquisition and Control Framework

The CMS Experiment has started an extensive upgrade program in the context of the High-Luminosity phase of the LHC (Phase-2). In order to cope with the highly demanding High-Luminosity conditions, CMS will need a completely new inner and outer tracking detectors. On top of R&D development, a Data AcQuisition (DAQ) software is being developed along with different applications to control, monitor and validate the newly produced modules of the future tracker. As more and more developers are getting involved to maintain all those codes, an environment where software and applications can work independently of the host machine operating system is crucial. The Phase-2 Tracker group decided to make use of Docker as containerization solution for its framework. This poster describes the containerization of DAQ software/applications utilized to test and validate the Tracker modules. This includes continuous integration and continuous deployment (so-called CI/CD) and running GUI applications inside containers.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Simple, Secure, Internet Delivery of MOOSE-based Applications

Application packaging and distribution are the final steps for delivering software to end-users; both are frequently neglected when creating scientific software. Commercial businesses rely on electronic distribution systems that have rendered disk drives obsolete. Still, national laboratories continue to rely heavily on removable media to distribute and limit access to controlled applications. With increasing concerns of unauthorized copying of sensitive applications, a modern distribution system that utilizes cryptographically secure communication and authentication protocols has been developed. This new distribution system will secure the chain of custody for nuclear software while simultaneously simplifying access to these tools. This report summarizes four primary advancements made toward the secure distribution of Nuclear Energy Advanced Modeling and Simulation (NEAMS)-developed, Multiphysics Object Oriented Simulation Environment (MOOSE)-based applications: application installation, package distribution, automated package building, and distribution of documentation. NEAMS is currently developing more than ten separate applications based on the open-source MOOSE Framework. Distribution of these applications has primarily been accomplished by distributing source code, with end-users compiling the applications themselves. This work created a mechanism where MOOSE applications can be installed in a similar way to any other software. This allows both administrators and end-users simplified access to runnable executables. With this new installation capability, it was then possible to rethink distribution. A new, secure capability for delivering MOOSE-based applications over the internet has been created. This system requires unique cryptographic tokens for authentication, greatly securing the custody chain for software. Once granted access, installation of any NEAMS code can be accomplished with these terminal commands: "conda install ncrc" "ncrc install ncrc-bison." After these two commands (and authenticating) the BISON application will be securely down- loaded from Idaho National Laboratory (INL)’s servers, installed, and ready to use. To enable this new distribution capability to be successful, the open-source Continuous Integration, Verification, Enhancement, and Testing (CIVET) Continuous Integration (CI) capability was augmented to add Continuous Delivery (CD). CD enables the automated building and packaging of MOOSE-based applications as they are modified by development teams, ensuring that our customers can obtain up-to-date versions of the software at any time. The need for instruction on how to use these applications was addressed through modifications to the MOOSE documentation system. The MooseDocs capability, which enables robust documentation of MOOSE-based applications, has been extended to allow both for the installation of documentation and the packaging of documentation with installed applications. Together, these enhancements form the core of a new, secure distribution mechanism for nuclear simulation tools. In concert with the Nuclear Computational Resource Center (NCRC), NEAMS- developed applications will now be straightforward to obtain securely.

97 MATHEMATICS AND COMPUTING↗

CI/CD Efforts for Validation, Verification and Benchmarking OpenMP Implementations

Software developers must adapt to keep up with the changing capabilities of platforms so that they can utilize the power of High-Performance Computers (HPC), including exascale systems. OpenMP, a directive-based parallel programming model, allows developers to include directives to existing C, C++, or Fortran code to allow node level parallelism without compromising performance. This paper describes our CI/CD efforts to provide easy evaluation of the support of OpenMP across different compilers using existing testsuites and benchmark suites on HPC platforms. Our main contributions include (1) the set of a Continuous Integration (CI) and Continuous Development (CD) workflow that captures bugs and provides faster feedback to compiler developers, (2) an evaluation of OpenMP (offloading) implementations supported by AMD, HPE, GNU, LLVM, and Intel, and (3) evaluation of the quality of compilers across different heterogeneous HPC platforms. With the comprehensive testing through the CI/CD workflow, we aim to provide a comprehensive understanding of the current state of OpenMP (offloading) support in different compilers and heterogeneous platforms consisting of CPUs and GPUs from NVIDIA, AMD, and Intel.

Jarmusch, Aaron↗

ACES: Infrastructure As Code. Model Optimization and Performance Capability

Infrastructure as Code (IaC) refers to managing infrastructure (networks, physical/virtual machines, storage, and connection topology) in a descriptive model/language, rather than configuring it manually or using interactive configuration tools. Just like source code can be compiled to generate the same binary code, IaC enables generating the same environment every time it is applied. IaC is a key DevOps practice and is generally used in conjunction with continuous integration (CI) and continuous delivery (CD). In CI, all code changes are merged into a mainline branch and validated multiple times a day as developers check in their changes to the source code. In CD on the other hand, code changes are automatically packaged for a new release-to-production on a regular basis. This typically enables teams to deliver software changes much more quickly and often. Developing and managing the ACES platform using (IaC) is vital for the robust deployment and continued sustainment of this foundational computing capability. IaC and DevOps practices will help us solve many of the common challenges often encountered in developing and maintaining compute infrastructure. First, it will make the provisioning, deployment, and maintenance of the compute infrastructure across multiple environments much more efficient. Second, these processes help make the overall system much more stable by continuously testing new changes as they are introduced to the system. Third, it allows us to be much more confident of the security controls in place since they can be tested as part of the CI process and all new changes can be audited and tracked. Finally, IaC enables the ACES Platform to be adaptable to the emerging technologies due to its ability to spin up different test beds to evaluate and incorporate these technologies. This document addresses common infrastructure-management challenges, describes what happens if they are not addressed, and highlights the value of utilizing IaC to tackle them. Finally, we will provide a high-level overview of the IaC and DevOps practices being utilized by the ACES Platform team.

97 MATHEMATICS AND COMPUTING↗