Search NASASearch

SEARCH · Search NASA

Results for “PIM architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

An introduction to the Gilgamesh PIM architecture

The focus of the work conducted is on a class of structures made possible by the merger of logic and memory. Processing In Memory (PIM) extends the design space much farther by closely associating the logic with the memory interface to realize innovative structures never previously possible and thus exposing entirely new opportunities for computer architecture.

Gilgamesh PIM CMOS

Processor-In-Memory (PIM) Based Architectures for PetaFlops Potential Massively Parallel Processing

The report summarizes the work performed at the University of Notre Dame under a NASA grant from July 15, 1995 through July 14, 1996. Researchers involved in the work included the PI, Dr. Peter M. Kogge, and three graduate students under his direction in the Computer Science and Engineering Department: Stephen Dartt, Costin Iancu, and Lakshmi Narayanaswany. The organization of this report is as follows. Section 2 is a summary of the problem addressed by this work. Section 3 is a summary of the project's objectives and approach. Section 4 summarizes PIM technology briefly. Section 5 overviews the main results of the work. Section 6 then discusses the importance of the results and future directions. Also attached to this report are copies of several technical reports and publications whose contents directly reflect results developed during this study.

Kogge, Peter M.

Gilgamesh: A Multithreaded Processor-In-Memory Architecture for Petaflops Computing

Processor-in-Memory (PIM) architectures avoid the von Neumann bottleneck in conventional machines by integrating high-density DRAM and CMOS logic on the same chip. Parallel systems based on this new technology are expected to provide higher scalability, adaptability, robustness, fault tolerance and lower power consumption than current MPPs or commodity clusters. In this paper we describe the design of Gilgamesh, a PIM-based massively parallel architecture, and elements of its execution model. Gilgamesh extends existing PIM capabilities by incorporating advanced mechanisms for virtualizing tasks and data and providing adaptive resource management for load balancing and latency tolerance. The Gilgamesh execution model is based on macroservers, a middleware layer which supports object-based runtime management of data and threads allowing explicit and dynamic control of locality and load balancing. The paper concludes with a discussion of related research activities and an outlook to future work.

management locality load balance

A Survey on the Expanding Scope and Interdisciplinary Opportunities for Processing-in-Memory Techniques

Processing-in-Memory (PIM) is emerging as a practical path to overcome the limitations of traditional von Neumann architectures. At its core, PIM systems implement computing primitives such as logic operations and multiply-accumulate acceleration through compute-in-memory, near-memory processing, or hybrid designs. The role of memory cells varies widely across technologies, acting as inputs, outputs, or analog accumulators through bit-lines and sense amplifiers. This diversity creates trade-offs in precision, bandwidth, latency, and programmability, making it difficult to build a unified understanding on the progress of the field. In this survey, we organize recent advances of PIM into three areas. First, we discuss the progress on the architectural optimizations of PIM and its integration with both DRAM and emerging non-volatile memories. Second, we examine how PIM is being used to accelerate key computing domains, including generative AI workloads and high-performance kernels, along with new approaches. Third, we highlight the growing adoption of PIM in computational sciences, where it is being applied to solve interdisciplinary problems such as genome analysis, mRNA quantification, mass spectrometry, quantum circuit simulation, wave modeling, and secure computation. Finally, we synthesize the major challenges that continue to slow PIM adoption, including manufacturing constraints, power delivery, thermal reliability, data consistency, runtime and memory-management coordination, and the difficulty of building portable software abstractions without sacrificing commercial viability. This work provides an updated, structured perspective on PIM’s potential across computing and computational sciences and the barriers that must be solved for it to reach its full impact.

Asifuzzaman, Kazi [Oak Ridge National Laboratory (

Universal Payload Information Management

As the overall manager and integrator of International Space Station (ISS) science payloads, the Payload Operations Integration Center (POIC) at Marshall Space Flight Center has a critical need to provide an information management system for exchange and control of ISS payload files as well as to coordinate ISS payload related operational changes. The POIC's information management system has a fundamental requirement to provide secure operational access not only to users physically located at the POIC, but also to remote experimenters and International Partners physically located in different parts of the world. The Payload Information Management System (PIMS) is a ground-based electronic document configuration management and collaborative workflow system that was built to service the POIC's information management needs. This paper discusses the application components that comprise the PIMS system, the challenges that influenced its design and architecture, and the selected technologies it employs. This paper will also touch on the advantages of the architecture, details of the user interface, and lessons learned along the way to a successful deployment. With PIMS, a sophisticated software solution has been built that is not only universally accessible for POIC customer s information management needs, but also universally adaptable in implementation and application as a generalized information management system.

Elmore, Ralph B.

Ionic-content-driven restructuring of spirobisindane ionene networks: implications for mechanics, self-healing, and gas transport

Polymers of intrinsic microporosity (PIMs) offer exceptional gas permeability but remain brittle and susceptible to physical aging, limiting their durability in separation applications. Here, we introduce a reconfigurable microporous polymer network that uniquely integrates permanent PIM microporosity with autonomous, intrinsic self-healing driven by imidazolium-based ionic motifs. Spirobisindane units generate the intrinsic free-volume architecture, while an imidazolium-containing polyamide ionene supplies dynamic ionic and hydrogen-bonding interactions that reorganize under mild activation. Incorporation of imidazolium-based ionic liquids further tunes cohesion, mobility, and densification, enabling the network to relax, re-associate, and retain microporosity without structural collapse. Through a comprehensive multiscale approach combining spectroscopy, scattering, thermal and mechanical characterization with all-atom molecular dynamics and density functional theory calculations, we elucidate how ionic content, as a single control parameter that reshapes free-volume distributions, modulates local coordination environments, and governs relaxation and healing kinetics. At intermediate ionic loadings, the networks achieve rapid, repeatable self-healing while maintaining CO$_2$ selectivity, demonstrating an optimal balance between segmental mobility and structural integrity. By establishing how hierarchical ionic interactions couple structure, dynamics, and transport in microporous ionene networks, this work provides generalizable design rules for adaptive soft-matter systems that require simultaneous mechanical resilience, reconfigurability, and selective gas transport.

36 MATERIALS SCIENCE

Cryogenic Pupil Alignment Test Architecture for Aberrated Pupil Images

A document describes cryogenic test architecture for the James Webb Space Telescope (JWST) integrated science instrument module (ISIM). The ISIM element primarily consists of a mechanical metering structure, three science instruments, and a fine guidance sensor. One of the critical optomechanical alignments is the co-registration of the optical telescope element (OTE) exit pupil with the entrance pupils of the ISIM instruments. The test architecture has been developed to verify that the ISIM element will be properly aligned with the nominal OTE exit pupil when the two elements come together. The architecture measures three of the most critical pupil degrees-of-freedom during optical testing of the ISIM element. The pupil measurement scheme makes use of specularly reflective pupil alignment references located inside the JWST instruments, ground support equipment that contains a pupil imaging module, an OTE simulator, and pupil viewing channels in two of the JWST flight instruments. Pupil alignment references (PARs) are introduced into the instrument, and their reflections are checked using the instrument's mirrors. After the pupil imaging module (PIM) captures a reflected PAR image, the image will be analyzed to determine the relative alignment offset. The instrument pupil alignment preferences are specularly reflective mirrors with non-reflective fiducials, which makes the test architecture feasible. The instrument channels have fairly large fields of view, allowing PAR tip/tilt tolerances on the order of 0.5deg.

Bos, Brent

Machine Learning Algorithm Performance on the Lucata Computer

A new parallel computing paradigm (processor in memory, or PIM) has recently become available, one that uses many lightweight threads, and where each thread migrates automatically to the memory used by that thread. Our effort focuses on understanding how suitable this architecture is for our application, and whether the hardware can sustain speedups as high as the system size permits. In particular we explore the kind of code optimizations needed, and how well optimized code scales. This paper describes some of the those optimizations, and the payoff in terms of scaling.

Kogge, Peter

Spread and SpreadRecorder An Architecture for Data Distribution

The Space Acceleration Measurement System (SAMS) project at the NASA Glenn Research Center (GRC) has been measuring the microgravity environment of the space shuttle, the International Space Station, MIR, sounding rockets, drop towers, and aircraft since 1991. The Principle Investigator Microgravity Services (PIMS) project at NASA GRC has been collecting, analyzing, reducing, and disseminating over 3 terabytes of collected SAMS and other microgravity sensor data to scientists so they can understand the disturbances that affect their microgravity science experiments. The years of experience with space flight data generation, telemetry, operations, analysis, and distribution give the SAMS/ PIMS team a unique perspective on space data systems. In 2005, the SAMS/PIMS team was asked to look into generalizing their data system and combining it with the nascent medical instrumentation data systems being proposed for ISS and beyond, specifically the Medical Computer Interface Adapter (MCIA) project. The SpreadRecorder software is a prototype system developed by SAMS/PIMS to explore ways of meeting the needs of both the medical and microgravity measurement communities. It is hoped that the system is general enough to be used for many other purposes.

Wright, Ted

HTMT-class Latency Tolerant Parallel Architecture for Petaflops Scale Computation

Computational Aero Sciences and other numeric intensive computation disciplines demand computing throughputs substantially greater than the Teraflops scale systems only now becoming available. The related fields of fluids, structures, thermal, combustion, and dynamic controls are among the interdisciplinary areas that in combination with sufficient resolution and advanced adaptive techniques may force performance requirements towards Petaflops. This will be especially true for compute intensive models such as Navier-Stokes are or when such system models are only part of a larger design optimization computation involving many design points. Yet recent experience with conventional MPP configurations comprising commodity processing and memory components has shown that larger scale frequently results in higher programming difficulty and lower system efficiency. While important advances in system software and algorithms techniques have had some impact on efficiency and programmability for certain classes of problems, in general it is unlikely that software alone will resolve the challenges to higher scalability. As in the past, future generations of high-end computers may require a combination of hardware architecture and system software advances to enable efficient operation at a Petaflops level. The NASA led HTMT project has engaged the talents of a broad interdisciplinary team to develop a new strategy in high-end system architecture to deliver petaflops scale computing in the 2004/5 timeframe. The Hybrid-Technology, MultiThreaded parallel computer architecture incorporates several advanced technologies in combination with an innovative dynamic adaptive scheduling mechanism to provide unprecedented performance and efficiency within practical constraints of cost, complexity, and power consumption. The emerging superconductor Rapid Single Flux Quantum electronics can operate at 100 GHz (the record is 770 GHz) and one percent of the power required by convention semiconductor logic. Wave Division Multiplexing optical communications can approach a peak per fiber bandwidth of 1 Tbps and the new Data Vortex network topology employing this technology can connect tens of thousands of ports providing a bi-section bandwidth on the order of a Petabyte per second with latencies well below 100 nanoseconds, even under heavy loads. Processor-in-Memory (PIM) technology combines logic and memory on the same chip exposing the internal bandwidth of the memory row buffers at low latency. And holographic storage photorefractive storage technologies provide high-density memory with access a thousand times faster than conventional disk technologies. Together these technologies enable a new class of shared memory system architecture with a peak performance in the range of a Petaflops but size and power requirements comparable to today's largest Teraflops scale systems. To achieve high-sustained performance, HTMT combines an advanced multithreading processor architecture with a memory-driven coarse-grained latency management strategy called "percolation", yielding high efficiency while reducing the much of the parallel programming burden. This paper will present the basic system architecture characteristics made possible through this series of advanced technologies and then give a detailed description of the new percolation approach to runtime latency management.

Sterling, Thomas

Case Study: Using The OMG SWRADIO Profile and SDR Forum Input for NASA's Space Telecommunications Radio System

The Space Telecommunication Radio System (STRS) standard is a Software Defined Radio (SDR) architecture standard developed by NASA. The goal of STRS is to reduce NASA s dependence on custom, proprietary architectures with unique and varying interfaces and hardware and support reuse of waveforms across platforms. The STRS project worked with members of the Object Management Group (OMG), Software Defined Radio Forum, and industry partners to leverage existing standards and knowledge. This collaboration included investigating the use of the OMG s Platform-Independent Model (PIM) SWRadio as the basis for an STRS PIM. This paper details the influence of the OMG technologies on the STRS update effort, findings in the STRS/SWRadio mapping, and provides a summary of the SDR Forum recommendations.

Briones, Janette C.

SALT: The Simulator for the Analysis of LWP Timing

With the emergence of new processor architectures that are highly multithreaded, and support features such as full/empty memory semantics and split-phase memory transactions, the need for a processor simulator to handle these features becomes apparent. This paper describes such a simulator, called SALT.

simulation