Search NASASearch

SEARCH · Search NASA

Results for “software development management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Developing and Distributing HEP Software Stacks with Spack

The Computational Science and AI Directorate at Fermilab is using Spack to support the development efforts of a large number of scientific programmers, in many independent projects and experiments. While independent, these projects share many dependencies. They are typically under continuous and fairly rapid development. They have to support deployment on diverse hardware. This is a different context than is typical for the management of HPC software, where Spack was born. To support our community, we have created a model that enables users to develop code with greater efficiency than is possible with Spack’s current development facilities. In this talk we will present: - a brief introduction to the science we support (particle physics) - how the code we work with is naturally organized into several layers of packages - how we are using Spack to manage those layers - how we leverage the layering to provide efficient support for developers, using our Spack extension “MPD”. - some suggestions for changes or additions to Spack to make such work easier.

Knoepfel, Kyle J. [Fermilab]

The Benefits and Weaknesses of Containerizing Software for HPC

Containerization technology has emerged as a transformative tool for software engineers, offering consistent development and deployment environments, simplifying dependency management, and enhancing scalability and portability across diverse systems. However, its application in High-Performance Computing (HPC) presents unique challenges, including the management of virtualization overhead, the need for efficient resource allocation, and the maintenance of optimal performance for compute-intensiv

Ho, Eric Victor [Sandia National Laboratories (SNL

Exascale workflow applications and middleware: An ExaWorks retrospective

Exascale computers offer transformative capabilities to combine data-driven and learning-based approaches with traditional simulation applications to accelerate scientific discovery and insight. However, these software combinations and integrations are difficult to achieve due to the challenges of coordinating and deploying heterogeneous software components on diverse and massive platforms. Here, we present the ExaWorks project, which addresses many of these challenges. We developed a workflow Software Development Toolkit (SDK), a curated collection of workflow technologies that can be composed and interoperated through a common interface, engineered following current best practices, and specifically designed to work on HPC platforms. ExaWorks also developed PSI/J, a job management abstraction API, to simplify the construction of portable software components and applications that can be used over various HPC schedulers. The PSI/J API is a minimal interface for submitting and monitoring jobs and their execution state across multiple and commonly used HPC schedulers. We also describe several leading and innovative workflow examples of ExaWorks tools used on DOE leadership platforms. Furthermore, we discuss how our project is working with the workflow community, large computing facilities, and HPC platform vendors to address the requirements of workflows sustainably at the exascale.

97 MATHEMATICS AND COMPUTING

A Review of Software for Designing and Operating Quantum Networks

Quantum networks development is crucial to realizing a production-grade network that can support distributed sensing, secure communication, and utility-scale quantum computation. However, the transition from laboratory demonstration to deployable networks requires software implementations of architectures and protocols tailored to the unique constraints of quantum systems. This paper reviews the current state of software implementations for quantum networks, organized around a three-plane abstraction of infrastructure, logical, and control/service planes. We cover software for both designing quantum network protocols (e.g., SeQUeNCe, QuISP, and NetSquid) and operating testbeds, with a focus on essential control/service plane functions such as entanglement, topology, and resource management, in a proposed taxonomy. Our review highlights a persistent gap between theoretical architecture and protocol proposals and their realization in simulators or testbeds, particularly in dynamic topology and network management. We conclude by outlining open challenges and proposing a roadmap for developing scalable software architectures to enable hybrid, large-scale quantum networks.

Network Design

Adaptive Solar Energy and Power Storage Platform for Multimetered Properties

Veritel Energy LLC has successfully demonstrated the feasibility of its innovative adaptive solar energy and power storage platform designed for multimetered properties in Phase I of its DOE SBIR grant. The project centered around the development and initial deployment of proprietary hardware and software system that manages and distributes site-generated renewable energy to multiple tenants efficiently. This system integrates seamlessly with existing building infrastructure and utility systems, making it a viable solution for aging multifamily properties often located in at-risk communities. The pilot system was installed in a multifamily building with a single master-meter, and functioned as a pilot to demonstrate a scalable model for renewable energy system adoption across similar properties. This installation not only showed potential for significant reductions in greenhouse gas emissions but also improved the economic viability of renewable installations by allowing property owners to recoup investments through increased energy savings and tenant billing. The Phase I project set the groundwork for further enhancements and testing in Phase II which will aim to apply the technology to projects with dozens of sub-meters, further contributing to the decarbonization of the built environment.

14 SOLAR ENERGY

An Agile/XP software development process for modernizing the accelerator control system at Fermilab

Fermilab is undergoing the most ambitious upgrade to its accelerator control system of the 21st century. As part of the ACORN project, hundreds of legacy control system applications written in C/C++ will be re-imagined and developed from the ground up. In addition, applications to support Fermilab’s new super-conducting linear accelerator are already under construction. To manage the development of modern controls applications, the Controls department has adopted an Agile software development process based on eXtreme Programming. In this paper we will describe our process and detail our experience applying it to the development of two case studies.

Diamond, John [Fermilab]

Optimal Operation of Residential High Performance Water Heater for Reduction of Electricity Cost and Peak Demand Through Field Validation

Water heating accounts for about 18% of a typical US home’s energy use. Modern water heaters have enabled control options through APIs, offering customers the opportunity to reduce their energy cost and peak demand by dynamically adjusting settings. A water heater’s capacity to store energy using its storage tank makes it an asset for peak demand reduction and energy cost savings. For this reason, a mixed-integer linear programming model is proposed to minimize the energy cost of a high-performance water heater while also reducing the peak demand of the residential household under a time-of-use utility rate by dynamically changing the water heater’s running mode. Specifically, a multi-objective optimization model is formulated to determine the mode settings of the water heater considering hot water use, time-of-use rate, and peak demand limit of the residential household. The mode settings are associated with different dead bands of water temperature for triggering on/off action of the heat pump and heating element. A 66-gal hybrid electric high performance water heater was used for numerical simulation and practical experiments. The simulation results were well aligned with measurements of practical experiments, validating the soundness of the thermodynamic model. In addition, reductions of energy cost, enabling affordability, and reducing peak demand are demonstrated. The research team also developed a software framework with dashboards to automatically and continuously monitor and manage devices.

Liu, Guodong [ORNL] (ORCID:0000000213498608)

EVs-at-RISC: A Secure and Resilient Interoperable SCM Control System Architecture for Electric Vehicle’s-at-Scale (Final Technical Report)

The EVs-at-RISC project was a five-year research, development, and demonstration initiative to create foundational tools for utility-scale fleet aggregation and Smart Charge Management (SCM) of Electric Vehicles (EV), Electric Vehicle Charging Infrastructure (EVCI), and related Distributed Energy Resources (DER). Rather than seeking to develop and demonstrate highly perfected SCM algorithms and control strategies, this project instead focused on creating foundational software solutions that enable unprecedented digital interoperability across the communications technologies and vendor platforms used to manage EV , EVCI, and DER, as well as existing energy management infrastructure operated by utilities, grid operators, and aggregators. This project then extends these novel interoperability capabilities to develop and deploy powerful middleware abstractions across grid edge networks and EVCI/DER fleet aggregations incorporating modern software tools and best practices, such as CI/CD, to bring the immense capabilities of infrastructure-as-code and policy-as-code to modern grid edge network environments. This addresses the foremost systemic issues preventing realization of any net operational benefits from scaled deployment of behind-the-meter EV, EVCI, and DER assets in electric power grids and markets today. The results of this approach and project unlock massive potential for new SCM capabilities to be easily prototyped, evaluated, and deployed at-scale within the existing grid edge network infrastructure and EVCI/DER technology ecosystem. The EVs-at-RISC project achieves this by extending Open Field Message Bus (OpenFMB), a conceptual model for digital interoperability and distributed intelligence in traditional front-of-meter utility SCADA networks, validating our hypothesis that OpenFMB could be similarly used to solve systemic digital interoperability issues in behind-the-meter environments and unlock real-world utility-scale SCM capabilities without requiring any new proprietary vendor solutions or significant infrastructure reconfiguration.

24 POWER TRANSMISSION AND DISTRIBUTION

Using containers to speed up development, to run integration tests and to teach about distributed systems

GlideinWMS is a workload manager provisioning resources for many experiments including CMS and DUNE. The software is distributed both as native packages and specialized production containers. Following an approach used in other communities like web development we built our workspaces, system-like containers to ease development and testing. Developers can change the source tree or check out a different branch and quickly reconfigure the services to see the effect of their changes. In this paper, we’ll talk about what differentiates workspaces from other containers. We’ll describe our base system composed of three containers. A one-node cluster including a compute element and a batch system. A GlideinWMS Factory controlling pilot jobs. And a scheduler and Frontend, to submit jobs and provision resources. Additional containers can be used for optional components. This system can easily run on a laptop and we’ll share our evaluation of different container runtimes, with an eye for ease of use and performance. Finally, we’ll talk about our experience as developers and with students. The GlideinWMS workspaces are easily integrated with IDEs like VS Code, simplifying debugging and allowing development and testing of the system also when offline. They simplified the training and onboarding of new team members and Summer interns. And they were useful in workshops where students could have first-hand experience with the mechanisms and components that, in production, run millions of jobs.

Mambelli, Marco

Using Containers to Speed Up Development, to Run Integration Tests and to Teach About Distributed Systems

GlideinWMS is a workload manager provisioning resources for many experiments, including CMS and DUNE. The software is distributed both as native packages and specialized production containers. Following an approach used in other communities like web development, we built our workspaces, system-like containers to ease development and testing. Developers can change the source tree or check out a different branch and quickly reconfigure the services to see the effect of their changes. In this paper, we will talk about what differentiates workspaces from other containers. We will describe our base system, composed of three containers: a one-node cluster including a compute element and a batch system, a GlideinWMS Factory controlling pilot jobs, and a scheduler and Frontend to submit jobs and provision resources. Additional containers can be used for optional components. This system can easily run on a laptop, and we will share our evaluation of different container runtimes, with an eye for ease of use and performance. Finally, we will talk about our experience as developers and with students. The GlideinWMS workspaces are easily integrated with IDEs like VS Code, simplifying debugging and allowing development and testing of the system even when offline. They simplified the training and onboarding of new team members and summer interns. And they were useful in workshops where students could have first-hand experience with the mechanisms and components that, in production, run millions of jobs.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681

FGMS-poster

Idaho National Laboratory (INL) performs post irradiation examination (PIE) of tri-structural isotropic (TRISO)-coated particle fuel to help qualify it for high temperature gas cooled reactors. TRISO fuel compacts are re-irradiated in the Neutron Radiography Reactor (NRAD) to generate the short lived fission products needed for fission product release testing. The Fuel Accident Condition Simulator (FACS) furnace and newly added Screen Neutron Irradiated Fuel for Failure (SNIFF) furnace heat the compacts in helium to temperatures of up to 2,000°C, prompting fission product release—predominantly gaseous xenon and krypton isotopes and condensable products such as cesium—from failed particles. These released isotopes are transported to a fission gas monitoring system (FGMS 1 or FGMS 3), where they accumulate in cryogenic cold traps and are quantified using high-purity germanium (HPGe) detectors. The addition of SNIFF and FGMS 3 increases throughput by enabling simultaneous testing of multiple compacts. Furthermore, automated INL developed software provides continuous, near-real time monitoring of fission product inventories and manages the liquid nitrogen cooling of the traps. These system enhancements improve the efficiency, data quality, and testing capacity of TRISO fuel performance evaluations.

07 - ISOTOPES AND RADIATION SOURCES

Lessons learned from the development and implementation of a workforce training curriculum for advanced controls for high performance HVAC systems

Over the past decade, academic research on advanced controls has slowly transitioned into new software platforms, giving rise to various companies developing and deploying these innovative products, including solutions for light commercial HVAC systems. However, the current workforce remains widely unprepared to install, maintain and operate these systems, particularly complex software-based control platforms, as most workforce training programs still focus on traditional building automation for large commercial buildings. This paper presents the development and piloting of curriculum for three key types of professionals: ● Technicians (trade-level): installing and maintaining modern high-performance HVAC systems and controls ● Programmers (undergrad-level): developing and implementing advanced controls ● Engineers and energy professionals (undergrad/grad-level): managing and evaluating system performance We share details of the material developed including training videos, open-source software, instruction manuals. We also present the results of a pilot implementation of the training materials with real students.

Casillas, Armando

Flexible Pilot Jobs Framework for Distributed High Throughput Computing

Experimental particle physics has been at the forefront of analyzing the world’s largest datasets for decades. The high-energy physics (HEP) community was among the first to develop suitable software and computing tools for this purpose. GlideinWMS is a Glidein-based workload management system whose purpose is to provide experiments like CMS at CERN, DUNE at Fermilab, and others, a way to access and efficiently use vast amounts of computing resources. This system wants to provide a simple way to submit jobs to a set of computing resources, that will be provided to users behind the scenes. Glideins are the pilot jobs executed on the worker nodes at the grid sites, performing operations such as hardware detection, environment setup, and error handling. After all these operations, they will launch the actual user job. Many grid sites are supported, such as shared clusters, Google CE, and AWS. My internship aimed to design and code a flexible pilot jobs framework that will replace the one used by GlideinWMS, developing a modular and flexible skeleton of the Glidein and adding further functionalities. My project also focused on the application of machine learning techniques as support to this management system.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

ECP libraries and tools: An overview

The Exascale Computing Project (ECP) Software Technology and Co-Design teams addressed the growing complexities in high-performance computing (HPC) by developing scalable software libraries and tools that leverage exascale system capabilities. As we enter the exascale era, the need for reusable, optimized software solutions that can handle the unique challenges posed by these systems becomes increasingly important. The primary challenges the ECP teams faced were to create software libraries and tools that are performant on exascale architectures and portable and usable across diverse hardware platforms. Efforts addressed issues related to concurrent execution, memory management, and the integration of heterogeneous computing resources, such as GPUs from multiple vendors. The ECP’s strategy involved a structured development process encompassing the creation, optimization, and deployment of software in collaboration with industry, academia, and national laboratories. The project was organized into several technical areas: co-design of domain-specific suites with target applications, programming models and runtimes, development tools, mathematical libraries, data and visualization tools, and software ecosystem and delivery mechanisms. ECP has successfully developed a large portfolio of software libraries and tools that demonstrate significant improvements in performance and scalability on exascale systems. These products have been integrated into the Department of Energy’s computing facilities, supporting various scientific applications and ensuring robust performance across different hardware setups. ECP advancements in software development for exascale computing highlight the importance of a collaborative and adaptive approach to handling next-generation HPC systems complexities. The lessons learned emphasize the need for continuous engagement with end-users and vendors, and the importance of maintaining a balance between innovation and practical implementation. Future efforts will focus on ensuring scalability, keeping pace with rapid hardware advancements, and further enhancing the interoperability and usability of the software ecosystem. In conclusion, subsequent articles in this special issue provide in-depth discussions and case studies into specific library and tool efforts.

97 MATHEMATICS AND COMPUTING

Multi-package development at Fermilab with Spack

The Spack package manager has been widely adopted in the supercomputing community as a means of providing consistently built on-demand software for the platform of interest. Members of the high-energy and nuclear physics (HENP) community, in turn, have recognized Spack’s strengths, used it for their own projects, and even become active Spack developers to better support HENP needs. Code development in a Spack context, however, can be challenging as the provision of external software via Spack must integrate with the developed packages’ build systems. Spack’s own development features can be used for this task, but they tend to be inefficient and cumbersome. We present a solution pursued at Fermilab called MPD (multi-package development). MPD aims to facilitate the development of multiple Spack-based packages in concert without the overhead of Spack’s own development facilities. In addition, MPD allows physicists to create multiple development projects with an interface that insulates users from the many commands required to use Spack well.

Knoepfel, Kyle James [Fermilab]

Electric Drive Technologies Research: ELT223 Component Modeling, Co-Optimization, and Trade-Space Evaluation Annual Report

This project is intended to support the development of new traction drive systems that meet the targets of 100 kW/L for power electronics and 50 kW/L for electric machines with reliable operation to 300,000 miles. To meet these goals, new designs must be identified that make use of state-of-the-art and next-generation electronic materials and design methods. Designs must exploit synergies between components, for example converters designed for high-frequency switching using wide band gap (WBG) devices and ceramic capacitors. This project included: (1) a survey of available technologies; (2) investigating new technologies, that for example, reduce volume of thermal management or magnetic components; (3) the development of computer aided design tools that consider the converter volume, reliability, and electrical performance; (4) exercising the design software to evaluate performance gaps and predict the impact of certain technologies and design approaches, i.e. GaN semiconductors, ceramic capacitors, ceramic thermal management components, and select topologies; (5) building and testing hardware prototypes to validate models and concepts. The design tools enable co-optimization of the power module and passive elements and provide some design guidance. At the end of the project, new advanced computing methods, such as machine learning approaches, were considered.

33 ADVANCED PROPULSION SYSTEMS

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Multi-package development at Fermilab with Spack

The Spack package manager has been widely adopted in the supercomputing community as a means of providing consistently-built on-demand software for the platform of interest. Members of the high-energy and nuclear physics (HENP) community, in turn, have recognized Spack’s strengths, used it for their own projects, and even become active Spack developers to better support HENP needs. Code development in a Spack context, however, can be challenging as the provision of external software via Spack must integrate with the developed packages’ build systems. Spack’s own development features can be used for this task, but they tend to be inefficient and cumbersome. \medskip We present a solution pursued at Fermilab called MPD (multi-package development). MPD aims to facilitate the development of multiple Spack-based packages in concert without the overhead of Spack’s own development facilities. In addition, MPD allows physicists to create multiple development projects with an interface that insulates users from the many commands required to use Spack well.

Knoepfel, Kyle [Fermilab] (ORCID:000000031433889X)