Search NASA⌕ Search

SEARCH · Search NASA

Results for “Compilers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

QECC-Synth: A Layout Synthesizer for Quantum Error Correction Codes on Sparse Architectures

Quantum Error Correction (QEC) codes are essential for achieving fault-tolerant quantum computing (FTQC). However, their implementation faces significant challenges due to disparity between required dense qubit connectivity and sparse hardware architectures. Current approaches often either underutilize QEC circuit features or focus on manual designs tailored to specific codes and architectures, limiting their capability and generality. In response, we introduce QECC-Synth, an automated compiler for QEC code implementation that addresses these challenges. We leverage the ancilla bridge technique tailored to the requirements of QEC circuits and introduces a systematic classification of its design space flexibilities. We then formalize this problem using the MaxSAT framework to optimize these flexibilities. Evaluation shows that our method significantly outperforms existing methods while demonstrating broader applicability across diverse QEC codes and hardware architectures.

Yin, Keyi [University of California, San Diego]↗

Agentic AI vs ML-Based Autotuning: A Comparative Study for Loop Reordering Optimization

High Performance Computing (HPC) applications rely heavily on code optimizations to achieve good performance on modern CPU and GPU architectures. Traditional Machine Learning auto-tuning approaches have demonstrated success in exploring high-dimensional spaces, but they often require expensive compile-run evaluations and lack adaptability for large HPC applications. The recent advances in Large Language Models (LLMs) and Agentic AI systems raise intriguing questions about the potential of these approaches to address specific optimization methodologies. This work aims to answer an essential question for the HPC community: “How Agentic AI Systems Compare to Traditional ML Autotuning Techniques?” To address this question, we present a comparative analysis between a traditional ML-based optimization approach and an Agentic AI system, evaluating their respective capabilities and limitations for loop-level optimization. In addition, we introduced a new Agentic AI system named LoopGen-AI using three different Large Language Models: GPT-4.1, Claude 4.0, and Gemini 2.5. A key finding is that LoopGen-AI achieves competitive per-formance with only a few program runs, the reasoning logs from the agents revealed that their decisions rely heavily on the combination of semantic understanding of the target kernel with dynamic feedback from the environment, highlighting a promising new dimension in performance tuning. In contrast, ML-based autotuners focus on statistical exploration, and require orders of magnitude more runs to reach peak performance. Additionally, our analysis shows that prompt engineering, particularly using Persona + Context Manager patterns, significantly impacts the effectiveness of Agentic AI. Our results indicate that while Agentic AI systems are not yet a complete replacement for ML-based autotuners, it can effectively complement traditional methods.

Rosas, Miguel Romero↗

Fusion Neutron Generator

The proposed code, named FROG (Fusion neutron Generator) is built upon the open-source particle transport Monte Carlo toolkit Geant4. Geant4 provides C++ classes that can be leveraged to build application-specific codes dealing with the transport of particles through matter. Geant4-based codes are applied in high-energy particle physics experiments, medical applications, shielding, and space applications for example. The FROG code allows the user to define the geometry of a neutron converter device shaped as a hollow cylinder, where a neutron breeding material such as lithium deuteride (LiD) is cladded by two concentric cylinders. Such neutron converter is then placed inside a regular nuclear fission reactor, where thermal neutrons will react with the neutron breeder material (typically, Lithium 6), and through a series of reactions, will generate high-energy neutrons – neutrons whose kinetic energy are around 14 MeV. The hollowed central portion can hold a specimen that will be bombarded by high-energy neutrons created inside the neutron breeding material. Figuratively speaking, this type of device transforms neutrons from thermal (~0.625 eV) to fusion (~14 MeV) energies and is sometimes termed “fusion-to-thermal neutron converters” in the literature. The code consists of C++ source file compiled and linked to generate an executable. The user can select the dimensions of the converter (radius, length, and thickness of the breeder material), the breeder material type, the cladding material, and the specimen material that will be activated or irradiated. As input, the neutron flux for a specific location inside a reactor, for instance, positions in ATR, is required. As output, the code predicts the number of high-energy neutrons produced, the total neutron flux and fluence as well as its detailed spectrum. The physics involved in such device is very complex, as it requires modeling neutron transport, light-ion (tritons) transport, as well as fusion reactions. The Geant4 toolkit provides the required physical models.

Martin, NicholasP. [Idaho National Laboratory (INL↗

MUPPET: An automated OpenMP mutation testing framework for performance optimization

MUPPET is a tool for OpenMP programs that identifies program modifications, called mutations, aimed at improving program performance. Existing performance optimization techniques, including profiling-based and auto-tuning techniques, fail to indicate program modifications at the source level thus preventing their portability across compilers. MUPPET aims to help HPC developers reason about performance defects and missed opportunities to improve performance at the source code level.

Parasyris, Konstantinos↗

Efficient Xml Interchange (exi) For Python (expy)

EXPy provides a native Python interface into the LF Energy EVerest V2G protocol stack. The protocol stack is implemented in C/C++ and compiled into shared object libraries. EXPy provides the Python Ctypes translation of the C/C++ libraries for use with pure Python software. This project eliminates the need for integrating Python with third-party communications applications and greatly reduces the code base and improves performance. The other major benefit is the ability for EXPy to support new EXI based protocols as additional V2G standards are produced (e.g. upgrade from ISO 15118-2 to ISO 15118-20).

Rohde, Kenneth [Idaho National Laboratory (INL), I↗

Proteus

Proteus enables Just-In-Time (JIT) compilation and optimization of C/C++ code using LLVM.

Georgakoudis, Giorgis↗

pnnl/CARTS

This compiler for ARTS (CARTS) is framework designed to connect a high productive languages with a distributed fine grained runtime that run effectively across clusters. It is built using MLIR and the LLVM infrastructure and it can be used as a platform to test static analysis ideas mapped towards distributed environments with novel technologies

Manzano Franco, Joseph [Pacific Northwest National↗

UnitGuard

Header-only C++17 library for compile-time only dimensional analysis

Tobin, WilliamR [Lawrence Livermore National Labor↗

rummy

A flexible parser for compiled input decks

Dempsey, Adam M. [LANL]↗

mvBayesR

SAND2025-11559O The mvBayesR tool performs multivariate Bayesian analysis on generic data. It includes tools for regression modeling, diagnosis, basis decomposition, sensitivity analysis, and visualization. The tool compiles state-of-the-art methodology into one easy-to-use package. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James [Sandia National Lab. (SNL-CA), Live↗

LAPIS: Linear Algebra Performance for Intermediate Subprograms

SAND2025-11594O LAPIS (Linear Algebra Performance for Intermediate Subprograms) is a compiler infrastructure for linear algebra that targets both high productivity and performance portability. It is based on the open-source MLIR package from the LLVM project. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Kelley, Brian↗

fortfmt-fix

SF-25-126 **fortfmt‑fix** is a small, deterministic codemod that rewrites these descriptors to explicit, standard‑conforming forms (e.g., `I12`, `F25.16`), enabling clean builds on modern compilers without vendor‑specific options. The tool is designed to be idempotent and conservative, preserving semantics while reducing maintenance burden in research codebases.

Bross, David [Argonne National Laboratory (ANL), A↗