Search NASA⌕ Search

SEARCH · Search NASA

Results for “Writing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

SEAS Communication Engine: An Extensible, Flexible Wrapper for Co-Simulation Agents

When modeling and analyzing the power grid and other large scale systems, researchers often express scenarios as optimization problems and feed them into advanced software solvers. In order to allow multiple solvers to communicate with each other and share data from different domains, the National Renewable Energy Laboratory (NREL) and associated Department of Energy (DOE) labs have developed a software framework called the Hierarchical Engine for Large-scale Infrastructure Co-Simulation (HELICS). HELICS allows cosimulation via a collection of client libraries for different languages that can be called from the appropriate optimization software. However, these client libraries do not provide a higher level of abstraction beyond reading and writing data off of the shared HELICS bus. In this paper, we describe a new software library called the SEAS Communication Engine that exposes a higher-level API for running cosimulation problems. The SEAS Engine provides a class-based abstraction on top of the Python HELICS client, in order to allow users to implement their domain-specific cosimulations without needing to interact with core HELICS primitives. This will make adoption of HELICS and cosimulation in general easier, by exposing a simpler API. In the second part of the paper, we validate our library on a collection of different simulation examples, including the canonical IEEE 13 Bus Feeder. Lastly, we demonstrate using the SEAS Engine to directly call domain-specific code written in the Julia programming language. Our hope is that this will serve as a template for easily calling software in different programming languages via the SEAS Engine, thereby avoiding code duplication and complexity.

co-simulation↗

UltraLiM: In-Memory Boolean Logic Architecture Using UltraRAM

Conventional computing architectures encounter ‘von Neumann’ and ‘memory wall’ bottlenecks which arise due to the back-and-forth data movement between the physically separate memory and processing units and the speed mismatch between them, respectively. These bottlenecks hurt both energy efficiency and the throughput of computing systems. To address these challenges, in-memory computing architectures have emerged as a promising alternative. They reduce the need for frequent data movement by executing different computing tasks inside the memory system. Here, we present UltraLiM, a logic-in-memory architecture using the UltraRAM-based memory system. UltraRAM holds the promise of developing a ‘universal memory’, overcoming the limitations of charge-based memories thanks to their non-volatile behavior with lower operating voltage. This work presents an in-memory computing architecture that integrates an UltraRAM-based memory array with a custom-designed peripheral circuitry. With this architecture, we can perform various in-memory Boolean logic operations (such as NOT, NAND, NOR, and XOR) in a single cycle. Leveraging the separate read-write paths in the UltraRAM-based memory array, we optimize read operations without encountering design conflicts. This optimization enhances the sense margin, enabling the use of simpler peripheral circuitry for in-memory logic operations.

Alam, Shamiul [University of Tennessee, Knoxville ↗

An Efficient Checkpointing System for Large Machine Learning Model Training

As machine learning models increase in size and complexity rapidly, the cost of checkpointing in ML training became a bottleneck in storage and performance (time). For example, the latest GPT-4 model has massive parameters at the scale of 1.76 trillion. It is highly time and storage consuming to frequently writes the model to checkpoints with more than 1 trillion floating point values to storage. This work aims to understand and attempt to mitigate this problem. First, we characterize the checkpointing interface in a collection of representative large machine learning/language models with respect to storage consumption and performance overhead. Second, we propose the two optimizations: i) A periodic cleaning strategy that periodically cleans up outdated checkpoints to reduce the storage burden; ii) A data staging optimization that coordinates checkpoints between local and shared file systems for performance improvement.

machine learning, artificial intelligence↗

LuSEE-night power distribution system design

The Lunar Surface Electromagnetic Experiment at Night (LuSEE-Night) is a low-frequency, 0.5 to 50 MHz, radio experiment on the radio-quiet far side of the Moon. The instrument will be launched by NASA Commercial Lunar Payload Services in 2026. The LuSEE-Night instrument core is composed of a radio frequency spectrometer (SPT) processing signals from four antennas, the Data Controller Board (DCB), low electromagnetic interference (EMI) Picket Fence Power Supply (PFPS), and Power Distribution Unit (PDU). The battery powers the instrument during the lunar night and stores energy harvested by the solar panel array during the lunar day. The battery charging is controlled by the Power Conditioning and Distribution Unit (PCDU). The unregulated power is supplied either by the SpaceCraft (S/C) or the battery and gets distributed to the PFPS, communication radio, heaters, and deployables through the PDU. The PFPS generates all regulated low-voltage rails using switching regulation synchronized to the LuSEE-Night clock, which ensures self-generated EMI will be confined to well-defined frequency bins. Here, we discussed the unregulated power distribution system architecture and functionality. The PDU engineering and flight modules are developed and characterized to confirm compliance with LuSEE-Night requirements. At the time of writing, all power subsystem components have been integrated into the payload.

47 OTHER INSTRUMENTATION↗

Unzipping hBN with ultrashort mid-infrared pulses

Manipulating the nanostructure of materials is critical for numerous applications in electronics, magnetics, and photonics. However, conventional methods such as lithography and laser writing require cleanroom facilities or leave residue. We describe an approach to creating atomically sharp line defects in hexagonal boron nitride (hBN) at room temperature by direct optical phonon excitation with a mid-infrared pulsed laser from free space. We term this phenomenon “unzipping” to describe the rapid formation and growth of a crack tens of nanometers wide from a point within the laser-driven region. Formation of these features is attributed to the large atomic displacement and high local bond strain produced by strongly driving the crystal at a natural resonance. This process occurs only via coherent phonon excitation and is highly sensitive to the relative orientation of the crystal axes and the laser polarization. Its cleanliness, directionality, and sharpness enable applications such as polariton cavities, phonon-wave coupling, and in situ flake cleaving.

36 MATERIALS SCIENCE↗

Field-free switching of perpendicular magnetization in a ferrimagnetic insulator with spin reorientation transition

Writing magnetic bits through spin-orbit torque (SOT) switching is promising for fast and efficient magnetic random-access memory devices. While SOT switching of out-of-plane (OOP) magnetized states requires lateral symmetry breaking, in-plane (IP) magnetized states suffer from low storage density. Here, we demonstrate a field-free switching scheme using a 5-nanometer europium iron garnet film grown with a (110) orientation that shows a spin reorientation transition from OOP to IP above room temperature. This scheme combines the benefits of high-density storage in the OOP states at room temperature and the efficient field-free SOT switching in the IP states at elevated temperatures. While conventional switching of OOP bits faces the dilemma that high OOP anisotropy is required to improve bit stability and low OOP anisotropy is required to lower switching current density, this scheme disentangles this interdependence, allowing for low switching currents to be possible without sacrificing the bit stability, offering opportunities for future memory devices.

Science & Technology - Other Topics↗

Orchard: Heterogeneous Parallelism and Fine-grained Fusion for Complex Tree Traversals

Many applications are designed to perform traversals ontree-likedata structures. Fusing and parallelizing these traversals enhance the performance of applications. Fusing multiple traversals improves the locality of the application. The runtime of an application can be significantly reduced by extracting parallelism and utilizing multi-threading. Prior frameworks have tried to fuse and parallelize tree traversals using coarse-grained approaches, leading to missed fine-grained opportunities for improving performance. Other frameworks have successfully supported fine-grained fusion on heterogeneous tree types but fall short regarding parallelization. We introduce a new frameworkOrchardbuilt on top ofGrafter.Orchard’s novelty lies in allowing the programmer to transform tree traversal applications by automatically applyingfine-grainedfusion and extractingheterogeneousparallelism.Orchardallows the programmer to write general tree traversal applications in a simple and elegant embedded Domain-Specific Language (eDSL). We show that the combination of fine-grained fusion and heterogeneous parallelism performs better than each alone when the conditions are met.

Computer Science↗

Using a Large Language Model as a Building Block to Generate Usable Validation and Verification Suite for OpenMP

In the HPC area, both hardware and software move quickly. Often new hardware is developed and deployed, the corresponding software stack, including compilers and other tools, are under active development while leading edge software developers are working to port and tune their applications, all at the same time. While the software ecosystem is in flux, one of the key challenges for users is obtaining insight into the state of implementation of key features in the programming languages and models their applications are using – whether they have been implemented, and whether the implementation conforms to the specification, especially for newly implemented features (less tested by widespread use). OpenMP is one of the most prominent shared memory programming models used for on-node programming in HPC. With the shift towards accelerators (such as GPUs and FPGAs) and heterogeneous programming OpenMP features are getting more complex. It is natural to ask whether generative AI approaches, and large language models (LLMs) in particular, can help in producing validation and verification test suites to allow users better and faster insights into the availability and correctness of OpenMP features of interest. In this work, we explore the use of ChatGPT-4 to generate a suite of tests for OpenMP features. We have chosen a set of directives and clauses, a total of 78 combinations, which first appeared in OpenMP 3.0 (released in May 2008) but are also relevant for accelerators. We prompted ChatGPT to generate tests in the C and Fortran languages, for both host (CPU) and device (accelerator). On the Summit super-computer using the GNU implementation, we found that, of the 78 generated tests 67 C tests and 43 Fortran tests compiled successfully and fewer than those executed to completion. On further analysis we show that not all generated tests are valid. We document the process, results, and provide detailed analysis regarding the quality of tests generated. With the aim of providing input to a production quality validation and verification suite, we manually implement the corrections required to make the tests valid according to the current OpenMP specification. We quantify this effort as small, medium, or large, and record the lines of code changed to correct the invalid tests. With the corrected tests we validate recent implementations from HPE, AMD, and GNU on the Frontier supercomputer. Our experiment and subsequent analysis show that although LLMs are capable of producing HPC specific codes, they are limited by their understanding of the deeper semantics and restrictions of programming models such as OpenMP. Unsurprisingly more commonly used features have better support, while some OpenMP 3.0 directives such as sections and tasking are not universally supported on accelerators. We demonstrate that successful compilation and execution to completion are inadequate metrics for evaluating generated code and that, at this time, commodity LLMs require expert intervention for code verification. This points to gaps in the training data that is currently available for HPC. We demonstrate that with "small" effort 37% of generated invalid C tests and 63% of generated invalid Fortran tests could be corrected. This improves productivity of test generation as we circumvent writing from scratch and the common programming errors associated with it.

Pophale, Swaroop [ORNL] (ORCID:0000000185446367)↗

Stability-preserving Lossy Compression for Large-scale Partial Differential Equations

Checkpoint/Restart (C/R) strategies are vital for fault tolerance in PDE-based scientific simulations, yet traditional checkpointing incurs significant I/O overhead. Lossy compression offers a scalable solution by reducing checkpoint data size, but conventional methods often lack control over physical invariants (e.g., energy), leading to instability such as oscillations or divergence in Partial Differential Equations (PDE) systems. This paper introduces a stability-preserving compression approach tailored for PDE simulations by explicitly controlling kinetic and potential energy perturbations to ensure stable restarts. Extensive experiments conducted across diverse PDE configurations demonstrate that our method maintains numerical stability with minimal error magnification—even across multiple checkpoint-restart cycles—outperforming state-of-the-art lossy compressors. Parallel evaluations on the Frontier supercomputer show up to 8.4× improvement in checkpoint write performance and 6.3× in read performance, while maintaining relative L2 errors ∼ 2e-6 throughout continued simulation. These results provide practical guidance for balancing compression accuracy, stability, and computational efficiency in large-scale PDE applications.

Gong, Qian [ORNL] (ORCID:0000000235704142)↗

Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability

Resolving the most fundamental questions in cosmology requires simulations that match the scale, fidelity, and physical complexity demanded by next-generation sky surveys. To achieve the realism needed for this critical scientific partnership, detailed gas dynamics must be treated self-consistently with gravity for end-to-end modeling of structure formation. Exascale computing enables simulations that span survey-scale volumes while incorporating key astrophysical processes that shape complex cosmic structures. We present results from CRK-HACC, a cosmological hydrodynamics code built for extreme scalability. Using separation-of-scale techniques, GPU-resident tree solvers, in situ analysis pipelines, and multi-tiered I/O, CRK-HACCexecuted Frontier-E: a four trillion particle full-sky simulation, over an order of magnitude larger than previous efforts. The run achieved 513.1 PFLOPs peak performance, processing 46.6 billion particles per second and writing more than 100 PB of data in just over one week of runtime. Frontier-E marks a significant advance in predictive modeling for next-generation cosmological science.

Frontiere, Nicholas [Argonne National Laboratory (↗

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗

The Right Time at the Right Places

In my undergraduate studies at the Massachusetts Institute of Technology (MIT), the short-lived chemical physics major allowed me to evade a number of courses required for chemistry majors. Thus, it was possible to take many physics courses and most of the advanced PhD-level courses in physical chemistry. I also took the introductory electrical engineering course in computer programming. The latter allowed me to write lots of computer code as a part of my (passing, but largely unsuccessful) senior thesis directed graciously by Professor Walter Thorson. As recommended by MIT Professor John C. Slater, I moved to Stanford University with Professor Frank Harris as my PhD supervisor. Frank was the perfect advisor for me, providing very close direction during the first year, and then allowing me to develop more and more independently. Within a few days of my twenty-fifth birthday, I became an assistant professor of chemistry at the University of California, Berkeley. Eighteen years later, I moved to the University of Georgia as director of a new research institute.

autobiography↗

Chemically Amplified, Dry-Develop Poly(aldehyde) Photoresist

The catalytic decomposition of poly(phthalaldehyde) with a photoacid generator can be used as dry-develop photoresist, where the exposed film depolymerizes into small molecules to allow the development of features via controlled vaporization. Higher temperatures enabled shorter dry-development times, but also promoted faster photoacid diffusion that compromised pattern fidelity. Trihexylamine was used as a base quencher to counteract acid diffusion in a phthalaldehyde-propanal co-polymer photoresist. The propanal co-monomer in the polymer improves the vaporization rate because it has a higher vapor pressure than phthalaldehyde. Addition of the base quencher was found to improve the contrast, pattern fidelity, and ease-of-handling of the dry-develop resist in a direct-write UV lithography tool. The dry-development of 4 μ m features was achieved with no appreciable residue. For large area features, a spatially variable exposure method was used to direct the residue away from the exposed area. The gradient exposure method was used to produce 100 μ m features. Plasma etching after dry-development was also used to achieve residue-free dry-developed patterns. These results show the benefits of incorporating base additives into a dry-develop depolymerizable resist system and highlight the need for addressing residue formation.

Materials Science↗

Data-Driven Protection Software to classify fault locations by protective zone in distribution systems with high PV penetration

The software contains (a) the source codes to generate Point-on-Wave (PoW) transient data for any feeder model in Alternative Transient Program (ATP) format. Codes provide options to change different steady state settings, including the loading condition and PV capacity and transient state setting like faults type, location and initiation time (b) data post-processing source code to converted data from native format to COMTRADE, csv, HDF5 (c) Docker container to train CNN to classify fault locations by protective zone. The container takes dataset and other training parameters (sampling rate, training epochs, batch size etc) as input to train CNN. The container writes back the trained CNN model, training and testing metrics and plots to the local workstation

Ramesh, Meghana↗

ISO_Fortran_binding_m v0.1.0

The Fortran programming language standard defines a broad feature set supporting the interoperability of Fortran programs with program written according to the C programming language standard. Among Fortran's C-interoperability features is a a C header file "ISO_Fortran_binding.h" This header file defines the interface to various C data structures and functions that C programs may use to access Fortran data entities. The ISO_Fortran_binding_m software defines a native Fortran module that presents an interface to these same data structures and functions. ISO_Fortran_bind_m thus enables Fortran programs to access and manipulate Fortran entities in ways that precisely mirror what C programs can do using ISO_Fortran_binding.h. ISO_Fortran_binding_m facilitates writing portable standard-conforming Fortran programs that emulate non-interoperable features, e.g., dynamic polymorphism, in a standard-conforming interoperable way similar but broader than what is demonstrated in Berkeley Lab's Caffeine software [1]. ISO_Fortran_binding_m also enables a Fortran programmer to extend Fortran's capabilities to emulate certain C functionality such as memory address arithmetic or computing C's "sizeof" function. [1] https://github.com/BerkeleyLab/caffeine/blob/213e3df1c319f0663306354414f352acda42a24f/src/caffeine/collective_subroutines/co_reduce_s.f90#L88 [2] https://github.com/BerkeleyLab/ISO_Fortran_binding_m/blob/0c585362bb4f2c72cf9049c800a7115b529ec533/src/iso_fortran_binding_m.F90#L196

Rouson, Damian↗

SOURCES4D

SOURCES is a code for computing neutron source rates and spectra from spontaneous fission (including delayed neutrons) and (alpha,n) reactions in homogeneous materials and (alpha,n) reactions in single-interface and two-interface geometries. SOURCES is a Los Alamos National Laboratory (LANL) code that is written in FORTRAN and distributed through the Radiation Safety Information Computation Center (RSICC). LANL’s last release of SOURCES to RSICC was SOURCES4C in 2002. This disclosure covers the latest version of SOURCES, SOURCES4D. This version adds sensitivity capabilities for (alpha,n) sources in homogeneous materials. Specifically, SOURCES4D writes new output that can be used to calculate, in post-processing, first and second derivatives of the (alpha,n) source rate density and spectrum with respect to nuclide densities in a homogeneous material and first derivatives of the (alpha,n) source rate density and spectrum with respect to nuclide stopping powers and (alpha,n) cross sections (nuclear data). These derivatives are useful for uncertainty quantification, predictive modeling, and other applications in neutron transport problems.

Favorite, Jeffrey A.↗

ALchemist (Active Learning Toolkit for Chemical and Materials Research) [SWR-25-102]

ALchemist is a modular Python toolkit that brings active learning and Bayesian optimization to experimental design in chemical and materials research. It is designed for scientists and engineers who want to efficiently explore or optimize high-dimensional variable spaces—without writing code—using an intuitive graphical interface.

Coatney, Caleb [National Renewable Energy Laborato↗

becquerel (bq) v0.7.0

Becquerel is a Python package for analyzing nuclear spectroscopic measurements. The core functionalities are reading and writing different spectrum file types, fitting spectral features, rebinning spectrum counts to different bin edges, performing detector calibrations and interpreting measurement results. It also includes tools for visualizing radiation spectra and fits of different spectral features, as well as convenient access to tabulated nuclear data both from remote servers and local caches. It relies heavily on the standard scientific Python stack of numpy, scipy, matplotlib, pandas, and numba. It is intended to be general-purpose enough that it can be useful to anyone from an undergraduate taking a laboratory course to the advanced researcher.

Bandstra, Mark [Lawrence Berkeley National Laborat↗