Search NASA⌕ Search

SEARCH · Search NASA

Results for “Algorithmic portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Ocean Optics Protocols for Satellite Ocean Color Sensor Validation

The document stipulates protocols for measuring bio-optical and radiometric data for the Sensor Intercomparison and Merger for Biological and Interdisciplinary Oceanic Studies (SIMBIOS) Project activities and algorithm development. This document supersedes the earlier version (Mueller and Austin 1995) published as Volume 25 in the SeaWiFS Technical Report Series. This document marks a significant departure from, and improvement on, theformat and content of Mueller and Austin (1995). The authorship of the protocols has been greatly broadened to include experts specializing in some key areas. New chapters have been added to provide detailed and comprehensive protocols for stability monitoring of radiometers using portable sources, abovewater measurements of remote-sensing reflectance, spectral absorption measurements for discrete water samples, HPLC pigment analysis and fluorometric pigment analysis. Protocols were included in Mueller and Austin (1995) for each of these areas, but the new treatment makes significant advances in each topic area. There are also new chapters prescribing protocols for calibration of sun photometers and sky radiance sensors, sun photometer and sky radiance measurements and analysis, and data archival. These topic areas were barely mentioned in Mueller and Austin (1995).

Fargion, Giulietta S.↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

Parallel variable-band Choleski solvers for computational structural analysis applications on vector multiprocessor supercomputers

A Choleski method used to solve linear systems of equations that arise in large scale structural analyses is described. The method uses a novel variable-band storage scheme and is structured to exploit fast local memory caches while minimizing data access delays between main memory and vector registers. Several parallel implementations of this method are described for the CRAY-2 and CRAY Y-MP computers demonstrating the use of microtasking and autotasking directives. A portable parallel language, FORCE, is also used for two different parallel implementations, demonstrating the use of CRAY macrotasking. Results are presented comparing the matrix factorization times for three representative structural analysis problems from runs made in both dedicated and multi-user modes on both the CRAY-2 and CRAY Y-MP computers. CPU and wall clock timings are given for the various parallel methods and are compared to single processor timings of the same algorithm. Computation rates over 1 GIGAFLOP (1 billion floating point operations per second) on a four processor CRAY-2 and over 2 GIGAFLOPS on an eight processor CRAY Y-MP are demonstrated as measured by wall clock time in a dedicated environment. Reduced wall clock times for the parallel methods relative to the single processor implementation of the same Choleski algorithm are also demonstrated for runs made in multi-user mode.

Poole, E. L.↗

Tolerant (parallel) Programming

In order to be truly portable, a program must be tolerant of a wide range of development and execution environments, and a parallel program is just one which must be tolerant of a very wide range. This paper first defines the term "tolerant programming", then describes many layers of tools to accomplish it. The primary focus is on F-Nets, a formal model for expressing computation as a folded partial-ordering of operations, thereby providing an architecture-independent expression of tolerant parallel algorithms. For implementing F-Nets, Cooperative Data Sharing (CDS) is a subroutine package for implementing communication efficiently in a large number of environments (e.g. shared memory and message passing). Software Cabling (SC), a very-high-level graphical programming language for building large F-Nets, possesses many of the features normally expected from today's computer languages (e.g. data abstraction, array operations). Finally, L2(sup 3) is a CASE tool which facilitates the construction, compilation, execution, and debugging of SC programs.

DiNucci, David C.↗

Automated Subpixel Snow Parameter Mapping with AVIRIS Data

We describe an automated algorithm (MEMSCAG) for mapping subpixel snow covered area (SCA) and snow grain size with AVIRIS data. The algorithm is based on the multiple endmember approach to spectral mixture analysis in which the spectral endmembers and the number of endmembers can vary on a pixel-by-pixel basis. This approach accounts for surface cover heterogeneity within a scene. The mixture analysis runs on endmembers from a spectral library of snow, vegetation, rock, soil, and lake ice spectra. Snow endmembers of varying grain size were produced with a radiative transfer model. All non-snow endmembers were collected with a portable field spectrometer. Mapping is performed through sequential 2-endmember, 3-endmember, and 4- endmember mixture model runs, each subject to constraints on RMS, residuals, fractions and priority. Grain size is determined by the grain size of the snow endmember used in the optimal mixture model. We apply MEMSCAG to AVIRIS data collected over Mammoth Mountain, CA and the northern site of the BOREAS in Manitoba, Canada. MEMSCAG produces appropriate snow covered area estimates in all regions. A preliminary comparison of grain size estimates from MEMSCAG with field measurements demonstrates high accuracy.

Painter, Thomas H.↗

Gesture-Based Robot Control with Variable Autonomy from the JPL Biosleeve

This paper presents a new gesture-based human interface for natural robot control. Detailed activity of the user's hand and arm is acquired via a novel device, called the BioSleeve, which packages dry-contact surface electromyography (EMG) and an inertial measurement unit (IMU) into a sleeve worn on the forearm. The BioSleeve's accompanying algorithms can reliably decode as many as sixteen discrete hand gestures and estimate the continuous orientation of the forearm. These gestures and positions are mapped to robot commands that, to varying degrees, integrate with the robot's perception of its environment and its ability to complete tasks autonomously. This flexible approach enables, for example, supervisory point-to-goal commands, virtual joystick for guarded teleoperation, and high degree of freedom mimicked manipulation, all from a single device. The BioSleeve is meant for portable field use; unlike other gesture recognition systems, use of the BioSleeve for robot control is invariant to lighting conditions, occlusions, and the human-robot spatial relationship and does not encumber the user's hands. The BioSleeve control approach has been implemented on three robot types, and we present proof-of-principle demonstrations with mobile ground robots, manipulation robots, and prosthetic hands.

BioSleeve↗

Challenges and Technology-Driven Opportunities for Safeguarding Microreactors

Nuclear microreactors (MRs) represent a new class of reactors characterized by their compactness, portability, and low power output. These features enable MRs to supply electricity and process heat to remote areas like military bases; inaccessible locations; small grids, such as on islands; or disaster impacted areas. Compared to traditional light water reactors, MRs have a unique set of attributes that need to be considered for the implementation of safeguard strategies. Current safeguard methodologies are reactor technology specific and are employed on large, stationary reactors where there is easy access by safeguards inspectors and where safeguard equipment can be easily installed and retrofitted. While there are numerous benefits to MRs, their compact size, portability, scalability, and operational lifetime create challenges to the traditional safeguard approaches, thus needing novel safeguard strategies. Here, this paper addresses the unique challenges posed by MRs to the international nuclear safeguards regime, including limited human resources, and explores how technology advancements can help mitigate these challenges. Specifically, it examines novel technologies that could contribute to establishing a comprehensive safeguards framework for MRs. These safeguards-enabling technologies encompass safeguards by design, remote sensing and monitoring technologies, applications of artificial intelligence and machine learning algorithms, utilization of digital twins, and system of systems assessments. While each of these safeguards-enabling technologies offers partial solutions to the challenges posed by MRs for the international safeguards regime, none of them alone can entirely address these challenges. Consequently, a combination of the safeguards-enabling technologies outlined in this paper is recommended to establish a robust safeguards regime for MRs.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Enabling Scientific Applications with Performance-Portability and High-Productivity for Multi-GPU Programming with JACC.Multi

This work bridges the gap between multi-GPU computing and high-productivity, performance-portable programming solutions. Our goal is to enhance scientific applications with a productive and portable solution—program once, deploy everywhere—for multi-GPU programming with no cost to programmability. To accomplish this, we implemented JACC.Multi, which is part of the Julia for ACCelerators (JACC) performance-portable framework. JACC. Multi is the only high-level, portable metaprogramming solution that targets multi-GPU environments and is integrated in a readily accessible programming language (e.g., Julia language). With transparent GPU-to-GPU communication, JACC. Multi is optimized for scientific application workloads and is portable for NVIDIA and AMD accelerators. For the evaluation, we use two modern multi-GPU systems: Hudson, which features two NVIDIA H100 Hopper GPUs per node, and Frontier, which features four AMD MI250X GPUs per node, each with two Graphics Compute Dies (GCDs) for a total of eight GCDs per node. Additionally, as part of the evaluation, we use JACC (one GPU), MPI+JACC, and JACC. Multi codes that implement well-known and widely used scientific algorithms/kernels such as the conjugate gradient algorithm and an explicit forward Euler solver that requires GPU-to-GPU communication. Overall, JACC. Multi codes achieve better performance than MPI+JACC codes and significant speedups over JACC (one GPU), with up to 1.9× on Hudson and 6× on Frontier.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Bone mineral computation with a rectilinear scanner

A portable rectilinear transmission scanner and associated computerized data reduction techniques for estimating bone mineral content are described. This unit can be easily disassembled for transport to various measurement sites and has been used to estimate the bone mineral content of the os calcis, radius, and ulna in the Apollo and Skylab astronauts. The scanner is used to obtain multiple rows of data from which a bone profile is derived. Bone edges are determined with the aid of a digital computer program which employs an algorithm that determines the greatest rate of change of the counting rate.

Ullman, J.↗

Sampling Technique for Robust Odorant Detection Based on MIT RealNose Data

This technique enhances the detection capability of the autonomous Real-Nose system from MIT to detect odorants and their concentrations in noisy and transient environments. The lowcost, portable system with low power consumption will operate at high speed and is suited for unmanned and remotely operated long-life applications. A deterministic mathematical model was developed to detect odorants and calculate their concentration in noisy environments. Real data from MIT's NanoNose was examined, from which a signal conditioning technique was proposed to enable robust odorant detection for the RealNose system. Its sensitivity can reach to sub-part-per-billion (sub-ppb). A Space Invariant Independent Component Analysis (SPICA) algorithm was developed to deal with non-linear mixing that is an over-complete case, and it is used as a preprocessing step to recover the original odorant sources for detection. This approach, combined with the Cascade Error Projection (CEP) Neural Network algorithm, was used to perform odorant identification. Signal conditioning is used to identify potential processing windows to enable robust detection for autonomous systems. So far, the software has been developed and evaluated with current data sets provided by the MIT team. However, continuous data streams are made available where even the occurrence of a new odorant is unannounced and needs to be noticed by the system autonomously before its unambiguous detection. The challenge for the software is to be able to separate the potential valid signal from the odorant and from the noisy transition region when the odorant is just introduced.

Duong, Tuan A.↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngaart, Rob F.↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngarrt, Rob F.↗

2025 Advances in NekRS: Supporting improved performance for nuclear applications

This report presents several 2025 advancements in NekRS, a high-fidelity spectral element CFD code developed at Argonne National Laboratory to support the NEAMS thermal-hydraulics program. The forthcoming v25 release consolidates several of these advances, adding new features for portability across heterogeneous GPU architectures, real-time in situ visualization, improved turbulence modeling, and conjugate heat transfer coupling. Over the past year, NekRS has demonstrated strong scalability and performance on DOE’s leading exascale platforms, including Aurora and Frontier, confirming its readiness for some of the largest and most complex simulations attempted to date. These achievements provide a powerful new platform for high-fidelity data generation, which in turn supports the development and validation of advanced closure models critical for reactor safety and design. Significant algorithmic innovations have also been introduced. A new global runtime h-refinement capability simplifies workflows by reducing mesh preparation burdens and enabling coarse-to-fine restarts. Building on this, a novel multigrid strategy was implemented to accelerate pressure and transport solves at scale, addressing long-standing bottlenecks in exascale CFD. Together, these developments improve both the efficiency and accessibility of high-fidelity simulations for reactor-relevant problems. Collectively, these enhancements represent a major step forward in simulation technology, positioning NekRS as a cornerstone of NEAMS efforts to enable accurate, efficient, and scalable high-fidelity analysis of advanced nuclear systems.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

Design and simulation of a muon detector to characterize geological overburden

This study presents the design, construction, and simulation of a mobile muon detector tailored for geological overburden characterization. The detector employs plastic scintillator paddles with silicon photomultipliers (SiPMs) and a QuarkNet data acquisition system, offering a portable solution suitable for remote field deployment. The simulator’s modular aluminum frame allows for adjustable geometry and directional sensitivity, while its battery system supports over a week of autonomous operation. Preliminary experimental tests confirmed that its muon flux measurements were consistent with theoretical expectations. A comprehensive simulation framework using Geant4 and CORSIKA was developed to model detector response and overburden effects. Analytical and Monte Carlo methods were used to assess quadrant resolution and infer muon directionality. This work lays the foundation for future overburden mapping and supports the development of reconstruction algorithms for geological applications.

72 - PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multi-frequency Tomography Radar Observations of Snow Stratigraphy at Fraser during SnowEx

SnowEx is a multi-year airborne snow campaign led by NASA. The purpose of SnowEx is to figure out how much water is stored in Earth’s terrestrial snow-covered regions. As part of the 2017 NASA SnowEx campaign, we deployed a portable triple-frequency (9.6GHz, 13.5GHz and 17.2GHz) and fully polarimetric frequency-modulated continuous-wave (FMCW) radar at Fraser, Colorado. The radar was installed on a 60cmx60cm frame to enable a full reconstruction of the three-dimensional variability per each radar channel. The tomography technique uses the radar echo from the multiple viewing positions and provides a unique access to the vertical structure of the snow layer. With current setup, the range resolution is 30cm. In this paper, we will review the radar design and signal-processing algorithm – time domain back projection. The generated vertical images show the snow stratigraphy, which is consistent with ground snow pit measurement. The continuous operation demonstrates diurnal thawing and refreezing process. The snow density is retrieved by comparing to the snow free image.

Esteban-Fernandez, Daniel↗

Statistical results from the Virginia Tech propagation experiment using the Olympus 12, 20, and 30 GHz satellite beacons

Virginia Tech has performed a comprehensive propagation experiment using the Olympus satellite beacons at 12.5, 19.77, and 29.66 GHz (which we refer to as 12, 20, and 30 GHz). Four receive terminals were designed and constructed, one terminal at each frequency plus a portable one with 20 and 30 GHz receivers for microscale and scintillation studies. Total power radiometers were included in each terminal in order to set the clear air reference level for each beacon and also to predict path attenuation. More details on the equipment and the experiment design are found elsewhere. Statistical results for one year of data collection were analyzed. In addition, the following studies were performed: a microdiversity experiment in which two closely spaced 20 GHz receivers were used; a comparison of total power and Dicke switched radiometer measurements, frequency scaling of scintillations, and adaptive power control algorithm development. Statistical results are reported.

Stutzman, Warren L.↗

Air Traffic Management Technology Demonstration-1 (ATD-1) Avionics Phase 2 Flight Test Training for Interval Management

Prior to the successful flight test validation of a new avionics prototype, participants from Boeing, Honeywell, and United Airlines underwent group training at NASA Langley Research Center. New prototype software for an algorithm which enables greater efficiency in high-density airspace, called Interval Management, was to be incorporated into Electronic Flight Bags and placed in the cockpit for pilot usage. The goals of the training were to teach the flight test pilots how to operate the new software, establish techniques to simultaneously position three aircraft prior to each test scenario, and ensure a common communication protocol among team members when coordinating the position of aircraft for the next scenario. The multi-tiered interactive training regimen consisted of a process that continually built upon previous foundational material. The primary learning elements were 1) a portable computer-based trainer that was provided to the pilots prior to classroom training sessions, 2) classroom learning, 3) full mock-up simulator training, and 4) refresher training just prior to the flight test. Each part of the regimen was designed to repeat and build upon the previous element. The purpose of this Technical Memorandum is to inform the aviation industry how flight training for Interval Management was conducted at Langley Research Center in order to reduce overall development costs of future Interval Management training programs. Secondly, the paper provides insight regarding the decision-making process when attempting to conduct a flight test.

Roper, Roy D.↗