Search NASA⌕ Search

SEARCH · Search NASA

Results for “AI Integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

Exploration Medical Integrated Product Team Clinical Decision Support Market Survey

NASA’s Exploration Medical Integrated Product Team (XMIPT) has identified Clinical Decision Support (CDS) technology as a critical need for future human space exploration. Such technology will support real time diagnosis, monitoring, and treatment of spaceflight medical conditions. The need for such tools and support systems is critical for in-mission clinical decision-making, especially when Earth-based support is unavailable due to communication delays or blackouts. Supporting technologies may or may not involve Artificial Intelligence (AI), and would support astronaut crew with minimal clinical training, or even those with advanced training if they are, for example, in need of a refresher, experiencing multiple stressors, or temporarily overloaded with tasking. This need traces to the Development of Earth Independent Operations Technologies for NASA’s Mars Campaign Office. The CDS Market Survey purpose, methods, outcomes thus far, and near-term steps will be discussed.

Decision support↗

A knowledge-based system design/information tool

The objective of this effort was to develop a Knowledge Capture System (KCS) for the Integrated Test Facility (ITF) at the Dryden Flight Research Facility (DFRF). The DFRF is a NASA Ames Research Center (ARC) facility. This system was used to capture the design and implementation information for NASA's high angle-of-attack research vehicle (HARV), a modified F/A-18A. In particular, the KCS was used to capture specific characteristics of the design of the HARV fly-by-wire (FBW) flight control system (FCS). The KCS utilizes artificial intelligence (AI) knowledge-based system (KBS) technology. The KCS enables the user to capture the following characteristics of automated systems: the system design; the hardware (H/W) design and implementation; the software (S/W) design and implementation; and the utilities (electrical and hydraulic) design and implementation. A generic version of the KCS was developed which can be used to capture the design information for any automated system. The deliverable items for this project consist of the prototype generic KCS and an application, which captures selected design characteristics of the HARV FCS.

Allen, James G.↗

Scalable Data Center Capacity for DOE's AI Prototype: A Rapidly Available Gigawatt Data Center for DOE

The multilaboratory Gigawatt Data Center working group was commissioned to identify approaches to rapidly establish federal data centers with scalable capacities up to 1,000 MW. These state-of-the-art facilities will serve as hubs for interdisciplinary collaboration, industry partnerships, and transformative applications of artificial intelligence. The proposed strategic shift includes facilitating multilaboratory collaboration, prioritizing operational efficiency, expanding public–private partnerships, optimizing investments, ensuring long-term contractual flexibility, supporting open science and secure data enclaves, and exploiting high-speed national networks. Owing to their extensive experience and best practices, the US Department of Energy national laboratories are uniquely positioned to lead this initiative. We recommend conducting a feasibility analysis to rapidly identify the optimal sites for this initiative, and the effort will likely involve private industry for design, construction, financing, and operational integration. We also propose establishing multiple geographically diverse sites to ensure energy resilience, high operational reliability, and a diverse user base, thereby effectively addressing the nation’s critical needs.

42 ENGINEERING↗

Development of Automated Atom Probe Tomography capability to study the influence of applied voltage and laser power on the final apparent composition of the analyzed specimen

This study presents the development and implementation of an autonomous Bayesian optimization (BO) framework for controlling and optimizing experimental parameters in Atom Probe Tomography (APT). Using commercial silicon needle samples as a benchmark system, we demonstrate that BO can efficiently navigate the complex parameter space of voltage and laser power to achieve target charge state ratios (specifically Si + /(Si + +Si 2+ )) with minimal experimental evaluations. Our implementation integrates Gaussian Process modeling with the CAMECA atom probe control framework, enabling autonomous adjustment of experimental conditions in real-time. Results show that the algorithm successfully converges to target ratios under different scenarios: maintaining a reference ratio, increasing the ratio (favoring Si 1+ ), and decreasing the ratio (favoring Si 2+ ). The system adapts to specimen evolution during analysis, compensating for changes in apex geometry while maintaining optimization targets. This work establishes a proof of concept for AI-driven optimization in APT, addressing the traditional challenges of manual parameter tuning and paving the way for applications to more complex materials where compositional accuracy is critical.

36 MATERIALS SCIENCE↗

Enabling HPC Scientific Workflows for Serverless

The convergence of edge computing, big data analytics, and AI with traditional scientific calculations is increasingly being adopted in HPC workflows. Workflow management systems are crucial for managing and orchestrating these complex computational tasks. However, it is difficult to identify patterns within the growing population of HPC workflows. Serverless has emerged as a novel computing paradigm, offering dynamic resource allocation, quick response time, fine-grained resource management and auto-scaling. In this paper, we propose a framework to enable HPC scientific workflows on serverless. Our approach integrates a widely used traditional HPC workflow generator with an HPC serverless workflow management system to create benchmark suites of scientific workflows with diverse characteristics. These workflows can be executed on different serverless platforms. We comprehensively compare executing workflows on traditional local containers and serverless computing platforms. Our results show that serverless can reduce CPU and memory usage respectively by 78.11% and 73.92% without compromising performance.

Andrei da silva, Anderson↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis – Simulated Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production using a single, simulated wind turbine. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel Hydrogen . While the unit supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. For the simulated wind energy profiles, NLR used OpenFAST to simulate a 3.4-MW International Energy Agency (IEA) reference wind turbine. The hour-long wind energy profiles varied over wind turbulence intensity (Class A or Class C) and average wind speed (5, 7, or 9 m/s). To match the power limits of the 1.25-MW electrolyzer and 3.4-MW IEA wind turbine most effectively and to maximize the efficiency of hydrogen production at a given average wind speed, the profiles were sometimes scaled by two times. This means that, in some cases, the experimental setup assumed two 1.25-MW electrolyzers were coupled with the wind turbine, representing a total maximum electrolysis load of 2.5 MW. Finally, NLR experimented with two settings for the electrolyzer power supply minimum and maximum current ramp rates (gain and slew): 200 and 400 amperes per second. The simulated profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}-{average wind speed}-{turbulence class}_{number of 1.25 MW electrolyzers connected}-{electrolyzer ramp rate in amperes/second} For instance, “windIEA3.4-5ms-C_2-400.zip” represents the hour-long experiment using the IEA 3.4-MW turbine, subjected to an average wind speed of 5 m/s and Class C wind turbulence, and connected to two 1.25-MW electrolyzers with the power supply set to a maximum current ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wind turbine power. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30 minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis .

08 HYDROGEN↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - Simulated Wave

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production using a single, simulated wave energy conversion device. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel Hydrogen. While the unit supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. For the wave energy, NLR used a wave energy converter model from PacWave. These devices can be equipped with accumulators and pressure relief values to smooth the power output by storing and releasing hydraulic energy. Using a peak power output of 10 MW, the model created two 25-minute profiles: one with and one without the accumulators and pressure relief valves. To down select the profile data from the native resolution of 20 Hz to 1 Hz, NLR took the mean of every 20 data points. NLR experimented with two simulated wave energy power plants: one that peaks at 10 MW, and one that peaks at 5 MW. These profiles were scaled for the physical 1.25 MW electrolyzer by multiplying the original profiles by one eighth and one quarter, respectively. The first profile matches the capacity rating of eight of the 1.25 MW electrolyzers, while the second matches four electrolyzers. Finally, NLR experimented with two settings for the electrolyzer power supply minimum and maximum current ramp rates (gain and slew): 200 and 400 amperes per second. The simulated profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wave electrolysis experiment and is formatted as follows: {technology}-{accumulator?}_{number of 1.25 MW electrolyzers connected}-{electrolyzer ramp rate in amperes/second} For instance, “wavePacWave-Noacc_4-400.zip” represents the 25 minute-long experiment using the PacWave’s wave energy converter model, equipped with no accumulator, connected to four 1.25-MW electrolyzers with their power supplies set to a maximum current ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wave power. An experiment, labeled “characterization_200.zip”, demonstrates the MC250 electrolyzer steady-state response with 30 minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all wave profiles combined into one dataset labeled "combined_wave_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis.

08 HYDROGEN↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production by conducting a statistical analysis of historical wind data over a five-year period (2020-2025) from a single 1.5MW turbine manufactured by General Electric (GE) located at NLR’s Flatirons Campus, to generate an experimental test profile that was deployed on a 1.25-MW proton exchange membrane type MC250 electrolyzer system manufactured by Nel Hydrogen . [1] While the electrolyzer balance-of-plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. The historical wind data provided several metrics, however, the analysis particularly focused on the measured power output by the wind turbine. The power output time series of data for each day was categorized by total energy generation and standard deviation, and the day that represented the highest combination of these two metrics was chosen – December 25th, 2022. This process was then repeated for a moving four-hour window within this day to identify the most statistically variable period. Finally, this four-hour period was scaled by 65% to match the 1.25 MW electrolyzer. The electrolysis system controls hydrogen production by varying DC current applied to the stack, from a maximum of 3000 A to a minimum safe operation of 300 A, or 10%. Because the current – voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical wind profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1 Hz frequency. For more details on the statistical analysis process, see the presentation labeled “ Public Reference Data for Megawatt-Scale Hydrogen Electrolysis” provided with each data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}_{scaling factor}-{electrolyzer ramp rate in amperes/second} For instance, “wind-GE1.5MW_0.65-400.zip” represents the hour-long experiment using historical data from the wind-GE1.5MW turbine, scaled to 65%, with the electrolyzer power supply set to a maximum ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and wind power input. A PDF file detailing the historical wind data statistical analysis used to generate the wind profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_historical_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis [1] nelhydrogen.com/product/mc-series-electrolyser .

08 HYDROGEN↗

Oscillatory Hierarchy Controlling Cortical Excitability and Stimulus Integration

Cortical gamma band oscillations have been recorded in sensory cortices of cats and monkeys, and are thought to aid in perceptual binding. Gamma activity has also been recorded in the rat hippocampus and entorhinal cortex, where it has been shown, that field gamma power is modulated at theta frequency. Since the power of gamma activity in the sensory cortices is not constant (gamma-bursts). we decided to examine the relationship between gamma power and the phase of low frequency oscillation in the auditory cortex of the awake macaque. Macaque monkeys were surgically prepared for chronic awake electrophysiological recording. During the time of the experiments. linear array multielectrodes were inserted in area AI to obtain laminar current source density (CSD) and multiunit activity profiles. Instantaneous theta and gamma power and phase was extracted by applying the Morlet wavelet transformation to the CSD. Gamma power was averaged for every 1 degree of low frequency oscillations to calculate power-phase relation. Both gamma and theta-delta power are largest in the supragranular layers. Power modulation of gamma activity is phase locked to spontaneous, as well as stimulus-related local theta and delta field oscillations. Our analysis also revealed that the power of theta oscillations is always largest at a certain phase of delta oscillation. Auditory stimuli produce evoked responses in the theta band (Le., there is pre- to post-stimulus addition of theta power), but there is also indication that stimuli may cause partial phase re-setting of spontaneous delta (and thus also theta and gamma) oscillations. We also show that spontaneous oscillations might play a role in the processing of incoming sensory signals by 'preparing' the cortex.

Shah, A. S.↗

Editorial: Predicting near-earth space environment: new perspective and capabilities in the AI age

Editorial on the Research Topic Predicting near-earth space environment: new perspective and capabilities in the AI age The near-Earth space environment is not only an operational hazard for space missions, but also a scientific laboratory for advancing our understanding and prediction of space plasma populations. This Research Topic is organized around three interconnected themes: observational datasets, machine-learning (ML) model development, and the discovery of new physical insights through those models. Its primary goal is to highlight the emerging capabilities in space environment prediction that are enabled, or will be enabled, by integrating advanced techniques—including AI/ML methods—with long-term curated datasets.

58 GEOSCIENCES↗

Controls and Automation Research in Space Life Support

A highly controlled and automated life support system has long been a NASA goal. It is usually assumed that life support for future long duration missions will use physical/chemical recycling systems that substantially close the oxygen and water circulation loops. Such a tightly coupled life support system has been thought to require an overall supervisory control system to minimize crew operation and maintenance activities. The International Space Station (ISS) Environmental Control and Life Support System (ECLSS) was at first expected to have supervisory control and automation. After this was found infeasible during the design of the ISS ECLSS in the early 1990's, it was then expected that the ISS or future mission systems would be upgraded to meet the original expectations. Since then NASA has extensively researched life support system controls and automation. Automation and Artificial Intelligence (AI) have gone through several cycles of enthusiasm and neglect before their recent great achievements, and NASA life support interest has similarly varied. Since the ISS ECLSS was launched, its on-board operational problems have led NASA to deemphasize system level controls and automation in favor of improving subsystem reliability and maintainability. Recent work has investigated supervisory control for a system similar to the ISS ECLSS. This paper reviews past planning and work on the supervisory control of closed, integrated physical/chemical life support systems similar to the ISS ECLSS and its precursors dating back to the 1960's.

life support↗

Aligning, Bonding, and Testing Mirrors for Lightweight X-ray Telescopes

High-resolution, high throughput optics for x-ray astronomy entails fabrication of well-formed mirror segments and their integration with arc-second precision. In this paper, we address issues of aligning and bonding thin glass mirrors with negligible additional distortion. Stability of the bonded mirrors and the curing of epoxy used in bonding them were tested extensively. We present results from tests of bonding mirrors onto experimental modules, and on the stability of the bonded mirrors tested in x-ray. These results demonstrate the fundamental validity of the methods used in integrating mirrors into telescope module, and reveal the areas for further investigation. The alignment and integration methods are applicable to the astronomical mission concept such as STAR-X, the Survey and Time-domain Astronomical Research Explorer.

mirror alignment↗

A Review and Outlook on Experimental Advances and Innovations in Geological CO 2 Storage: Insights from Depleted Gas Reservoirs and Saline Aquifers

Geological storage of carbon dioxide (CO 2 ) in depleted gas reservoirs and deep saline aquifers is a key part of global decarbonization efforts. As carbon capture and storage advances toward commercial-scale deployment, the credibility and scalability of laboratory experiments are increasingly vital for guiding safe and effective field implementation. This review offers a comprehensive, cross-scale evaluation of experimental methodologies, including core flooding, high-pressure, high-temperature systems, microfluidic visualization, and emerging systems such as multilayer commingled/compartmentalized core flooding, 3D-printed micromodels, and AI-powered digital twins. These innovations are demonstrated to enhance representativeness, reproducibility, and real-time insight, thereby addressing the limitations of conventional workflows. A critical analysis of methodological gaps, such as inconsistent pressure–temperature conditions, oversimplified brine chemistry, and a lack of standardization, reveals experimental sources of scale translation errors and performance uncertainty. By comparing the unique challenges of depleted gas reservoirs (such as low water saturation and legacy well leakage) to those of saline aquifers (including pressure buildup and caprock integrity), this review identifies formation-specific priorities for experimental design. Novel contributions include a synthesis of best practices, integration strategies for model calibration, and recommendations for standardizing core handling, saturation procedures, and reporting protocols. Furthermore, this work serves as a guide for developing robust, field-relevant experimental strategies that can increase the deployment and regulatory acceptance of CO 2 storage technologies at scale.

58 GEOSCIENCES↗

Safety Assessment of a Machine Learning-Based Aircraft Emergency Braking System: A Case Study

Machine Learning (ML) is revolutionizing many technological fields, but its use in aviation remains restricted due to stringent certification requirements. Efforts by the aviation community to establish standards for certifying ML-based systems are progressing, yet challenges persist, particularly with safety assessment methods for ML-based systems. This research addresses these challenges through a case study of an autonomous emergency braking system utilizing a computer vision deep neural network (DNN). We demonstrate a safety assessment process tailored to ML-specific concerns, such as low integrity and performance variability in quantitative safety analysis. This study can serve as an illustrative example to facilitate the discussion and convergence on certification aspects for ML-based systems within the aviation community.

Safety certification↗

Soil Salinity Level Assessment and Prediction Integrating UAV-borne Hyperspectral Imaging and Machine Learning Algorithms to Combat Desertification

In response to the ongoing global food crisis, the United Nations has identified “Zero Hunger” as one of its Sustainable Development Goals. A central contributor to the crisis is the process in which agricultural lands go through desertification. Research has shown a direct correlation between soil salinity and desertification - increased salinity levels indicate a higher risk for desertification. Furthermore, researchers have explored various techniques to map soil salinity, but these methods are oftentimes inefficient and don’t address future salinity predictions. To improve desertification monitoring, soil salinity can be observed via hyperspectral imaging on unmanned aerial vehicles (UAVs) to predict the risk of agricultural desertification using artificial intelligence (AI) and machine learning (ML) techniques. A significant gap exists in past research that applies ML and imaging techniques to soil salinity: convolutional neural networks (CNNs) and regression models are rarely leveraged together, despite the efficiency and accuracy of these models. To compensate for this gap, the proposed system leverages the use of these AI and ML models to improve soil assessment and prediction techniques. This approach involves three steps - data collection, image analysis, and future prediction. Using hyperspectral cameras on UAVs to collect the data from the region, a trained CNN model will output estimated soil salinity levels at a specific time. The estimations will then be analyzed by a regression model to assess the accuracy of future soil salinity predictions. The proposed system will identify regions at risk of desertification to help farmers mitigate agricultural loss, in turn helping alleviate the food crisis.

UAV systems↗

Deploying and Tracking Software with NCCS Software Provisioning

The National Center for Computational Sciences (NCCS) at Oak Ridge National Laboratory has a long history of deploying ground-breaking leadership-class supercomputers for the U.S. Department of Energy. The latest in this line of supercomputers is Frontier, the first supercomputer to break the exascale barrier (1018 floating-point operations per second) on the TOP500 list. Frontier serves a wide array of scientific domains, from traditional simulation-based workloads to newer AI and Machine Learning workloads. To best serve the NCCS user community, NCCS uses Spack to deploy a comprehensive software stack of scientific software packages, providing straightforward access to these packages through Lmod Environment Modules. Maintaining a large software stack while also including multiple new compiler releases each year is a very time-consuming task. Additionally, it is not straightforward to provide a software stack alongside existing vendor-provided software such as the HPE/Cray Programming Environment (CPE), and existing CPE, Spack, and Lmod integration does not allow for multiple versions of GPU libraries such as AMD’s ROCm to be used. To address these challenges and shortcomings, NCCS has developed the NCCS Software Provisioning tool (NSP)1, a tool for deploying and monitoring software stacks on HPC systems. NSP allows NCCS to quickly and effectively provision software stacks from the ground up using template-driven recipes and configuration files. NSP is successfully deployed on Frontier and several other NCCS clusters, enabling the NCCS software team to quickly deploy software stacks for newly-released compilers, expand current software offerings, better support GPU-based software, and monitor Lmod module usage to identify unused software packages that can be removed from the software stack. In this work, we discuss the shortcomings of the previous CPE, Spack, and Lmod usage at NCCS, provide further details on the implementation and structure of NSP, then discuss the benefits that NSP provides.

Rentschler, Asa [ORNL] (ORCID:0009000597694743)↗