Search NASASearch

SEARCH · Search NASA

Results for “chunking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Chemistry imaging and distribution analysis of rare earth elements in coal using LIBS and LA-ICP-MS instruments

Currently, demand for rare earth elements (REEs) increased significantly. Coal is actively evaluated as potential economic sources for extraction of REEs. Here, in this work, laser-induced breakdown spectroscopy (LIBS) was evaluated for rapid estimation of REEs content and their distribution in the natural coal samples. The results were compared with similar laser ablation–inductively coupled plasma–mass spectrometry (LA-ICP-MS) measurements. Thirteen coal samples (nine standard samples and five natural samples) were used in this study. Powder samples were pressed into pellets while coal chunks were directly ablated for data recording. Pellets of the powder standard samples were used to optimize the data acquisition system and then data recorded with this optimized system was used to identify the proper data acquisition and analysis models. After establishing the proper data acquisition system and analysis model using the standard samples, natural coal samples in powder form and their chunks were utilized to record LIBS and LA-ICP-MS spectra. Multivariate calibration models were developed using four of the natural samples, which were evaluated by predicting the REE content in the fifth sample. Principal component analysis was performed on the LIBS data obtained from the natural samples and it classified all the samples with high accuracy. Two-dimensional (2D) elemental mapping on coal chunk samples was also performed using both LIBS and LA-ICP-MS to study the distribution of REEs in the samples. The resulting elemental images and their correlations can be used to infer mineral distributions.

01 COAL, LIGNITE, AND PEAT

An Open-source Llm Enhanced-tool Specialized In Helping Moose Related Problems And Tasks

MOOSEenger is an open-source, terminal-first chat application for the MOOSE ecosystem that couples specialized parsing of MOOSE documentation and “.i” input files with retrieval-augmented generation to deliver grounded answers about multiphysics modeling and workflows. It includes dedicated readers for MOOSE-style HTML and a pyhit-based parser that uses the MOOSE syntax tree to preserve block structure and attach retrieval metadata. A data-ingestion pipeline performs semantic chunking into atomic facts and stores them hierarchically in a local Chroma vector database that maintains parent–child relationships across documents; the system can ingest directories, individual files, and single-page web content, and it provides CRUD operations (insert, update, delete) to manage the corpus. At query time, relevant chunks are embedded, retrieved, and fused into the model context, with interactive features such as token streaming, persistent chat history, and dynamic RAG (retrieval triggered by user input or intermediate model output). Deployment is flexible: MOOSEenger runs with local Ollama models or remote Hugging Face/OpenAI backends—typically coordinating generation, lightweight tagging/summarization, and embeddings across three models—and it also supports a server mode and integration with the VS Code Continue interface.

Li, Mengnan [Idaho National Laboratory (INL), Idah

Pythia8 Quark and Gluon Jets (float8 e4m3FN)

A float8 (e4m3FN) quantized version of the quark and gluon jet dataset originally published by Komiske, Metodiev, and Thaler (Zenodo record 3164691). Only the 20-file subset without charm and bottom quark jets is included here. All simulation parameters and jet selection criteria are identical to the original: Pythia 8.226, √s = 14 TeV Quarks from WeakBosonAndParton:qg2gmZq, gluons from WeakBosonAndParton:qqbar2gmZg with the Z decaying to neutrinos FastJet 3.3.0, anti-k_t jets with R = 0.4 p_T^jet ∈ [500, 550] GeV, |y^jet| < 1.7 There are 20 files, each in compressed NumPy format (QG_jets_fp8e4m3fn_0.npz through QG_jets_fp8e4m3fn_19.npz). Each file contains two arrays: X: (100000, M, 4) — 50k quark and 50k gluon jets, randomly sorted, padded to max multiplicity M, with particle features (pt, rapidity, azimuthal angle, pdgid) y: (100000,) — jet labels, gluon = 0, quark = 1 Since NumPy has no native fp8 dtype, X is stored as float32, but the values have been quantized through TensorFlow's float8_e4m3fn type and carry only fp8 precision. The quantization procedure is as follows: a global per-channel scale factor is computed from the absolute maximum value across all 20 chunks (with FP8_MAX = 448.0, the maximum representable value of e4m3FN). Each chunk is then scaled into the fp8 dynamic range, round-tripped through tf.experimental.float8_e4m3fn, and scaled back. This global scaling ensures a consistent quantization grid across the full dataset. The y labels are unchanged. Users should be aware that e4m3FN has limited dynamic range and precision. We recommend verifying this format is appropriate for your application; for a less aggressive reduction see the float16 and float32 versions linked below. If you use this dataset, please cite the original Zenodo record and its associated paper: Komiske, Metodiev, Thaler, Energy Flow Networks: Deep Sets for Particle Jets, JHEP 01 (2019) 121, arXiv:1810.05165

DiLullo, Nicholas [Brown University] (ORCID:000000

BSEC flux towers: CSAT3B and TRH

The data were collected as part of the BSEC project, during the period from June 2025 to May 2026. Directory "broadway" contains data collected on a multi-level flux tower (US-BWf) in the Broadway East neighborhood (1808 North Patterson Park Ave., Baltimore City, MD 21213; LAT: 39o18'40.31'' N; LONG: 76o35'12.43'' W). At each of the four measurement heights (8.5 m, 11.1 m, 13.4 m, 15.9 m), a Campbell Scientific CSAT3B sonic anemometer was operated at 50 Hz to measure virtual temperature (tc) and three velocity components (u: 270 degrees; v: 180 degrees; w: vertical), and a RM Young temperature sensor (model 41382VC) was operated at 1 Hz inside a compact aspirated radiation shield (model 43502) to measure absolute temperature (T) and relative humidity (RH). Inside directory "broadway", directory "netcdf" contains data collected each day in 5-minute chunks that have been converted to NetCDF format (before quality checking), while "4hr" contains data arranged into 4-hour chunks (also in NetCDF format) that have been through basic quality checking steps (treating data points with nonzero diagnostic codes as missing data; fixing six or fewer consecutive missing data points using linear interpolation). Users are recommended to start with data in directory "4hr", while data in directory "netcdf" can be used for reference purposes.

Baltimore

Ptychographic reconstructions performed in real time and offline have equivalent quality

Abstract Ptychography is a burgeoning imaging technique that enables high-resolution, lensless reconstruction of complex samples by analysing overlapping diffraction patterns, making it invaluable in fields like materials science, biology, and nanotechnology. Real-time ptychographic reconstructions are gaining interest in the scientific community as they provide immediate feedback. Yet their potential to replace offline reconstructions remains uncertain, in part due to questions about the quality of the resulting images. This study quantitatively compares real-time and offline reconstructions at different overlap conditions. Offline reconstructions, using all diffraction patterns at once, and real-time reconstructions, where new frames are added to the reconstructions in small chunks as the diffraction patterns are recorded, were indistinguishable and identical in reconstruction quality. These results hold consistently across all tested overlap ratios. This study represents the first quantitative analysis of real-time ptychographic reconstruction using a growing dataset, demonstrating the potential for real-time reconstructions to replace or at least complement offline reconstructions.

Science & Technology - Other Topics

Quantum error mitigation by layerwise Richardson extrapolation

A widely used method for mitigating errors in noisy quantum computers is Richardson extrapolation, a technique in which the overall effect of noise on the estimation of quantum expectation values is captured by a single parameter that, after being scaled to larger values, is eventually extrapolated to the zero-noise limit. We generalize this approach by introducing layerwise Richardson extrapolation (LRE), an error mitigation protocol in which the noise of different individual layers (or larger chunks of the circuit) is amplified and the associated expectation values are linearly combined to estimate the zero-noise limit. The coefficients of the linear combination are analytically obtained from the theory of multivariate Lagrange interpolation. LRE leverages the flexible configurational space of layerwise unitary folding, allowing for a more nuanced mitigation of errors by treating the noise level of each layer of the quantum circuit as an independent variable. Furthermore, we provide numerical simulations demonstrating scenarios where LRE achieves superior performance compared to traditional (single-variable) Richardson extrapolation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]

Physical properties, internal structure, and the three‐dimensional petrography of CI chondrites

physical properties and the nature of their breccation, we investigated nine samples of the Ivuna and Orgueil CI chondrites ranging in size from 1 mm to 4 cm in approximate diameter. The combined mass of unique material investigated in this work is 113 g. For our investigations, we use ideal gas pycnometry, 3-D laser scanning, x-ray computed microtomography (μCT), and accompanying digital data extraction techniques. We found that the bulk density of the samples ranged from 1.61 to 2.10 g cm −3 . Larger samples tend to have a lower bulk density. Grain density (ranging from 2.44 to 2.55 g cm −3 ) is significantly less variable than the bulk density in our samples and the quantity of porosity (ranging from 14.6% to 33.8%) is the dominant factor in determining the bulk density of CI chondrite material. Our μCT results show that the visible porosity across all sizes of our CI chondrite samples is in the form of cracks, but these cracks can account for less than two-thirds of the porosity in the CI chondrites. Other porosity is not visible, even at μCT resolutions of 2.7 μm voxel edge −1 and we conclude that it is sub-micron in nature. It is not clear if the cracks seen in our samples are indigenous to the chondrites or are a result of terrestrial processes. We also find that the CI chondrites are excellent examples of the fractal-like nature of brecciation, where clasts can be observed at all scales we imaged. The breccias are composed of sub-equant-shaped and sub-rounded-textured clasts like melt-free impact breccias on other solar system bodies. From our μCT volume and digital data extraction, we determine that the Ivuna CI chondrite breccia is organized: the mostly sub-equant clasts within our ~2 cm chunk of Ivuna have a mean diameter of 1.33 mm and their aligned longest axes define a lineation structure. We speculate that the lineation was imparted after fragmentation of the clasts by slight shear on the parent asteroid which could be the result of seismic-related granular flow or mild non-axial impact-related compaction. These data will help to place returned asteroidal material from asteroids 162173 Ryugu and 101955 Bennu and the CI chondrites into a mutual geological context.

CI chondrite

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat

Globus service enhancements for exascale applications and facilities

Many extreme-scale applications require the movement of large quantities of data to, from, and among leadership computing facilities, as well as other scientific facilities and the home institutions of facility users. These applications, particularly when leadership computing facilities are involved, can touch upon edge cases (e.g., terabyte files) that had not been a focus of previous Globus optimization work, which had emphasized rather the movement of many smaller (megabyte to gigabyte) files. We report here on how automated client-driven chunking can be used to accelerate both the movement of large files and the integrity checking operations that have proven to be essential for large data transfers. In conclusion, we present detailed performance studies that provide insights into the benefits of these modifications in a range of file transfer scenarios.

97 MATHEMATICS AND COMPUTING

Communications Reliability for Vehicle Grid Integration

Electric Vehicles (EVs) adoption rate has been steadily increasing in the US leading to a growing number of charging stations including faster DC (Direct Current) chargers and slower Level 1 and Level 2 AC (Alternating Current) chargers. This increase in demand for electricity is further exacerbated by recent developments in Artificial Intelligence (AI) technology, advanced manufacturing, and digitization. These factors will require electric utilities to upgrade their infrastructure to keep up with the increasing electrical demand (especially during peak hours). An easy way to counteract the need for these upgrades is to shift a major chunk of active charge sessions (durations where there is energy transfer from charger to EV's propulsion battery) to off-peak hours thereby flattening the load curve and making the infrastructure more resilient. This concept is known as Smart Charge Management (SCM). EV owners also benefit from SCM since it lowers their charging costs and consequently their transportation costs by prioritizing charging during off-peak hours. SCM takes advantage of EV's capability to act as a controllable load or DER (Distributed Energy Resource). This report summarizes the reliability analysis performed on the communication required for two of these SCM use-cases. This analysis only focuses on SCM strategies for unidirectional charging (energy transfer from EVSE to EV or V1G) and not bidirectional charging.

24 POWER TRANSMISSION AND DISTRIBUTION

Knowledge Graph for End-to-End Traceability of an Integrated Human-Earth System Model

Integrated human-Earth system models inform energy-water-land system dynamics and policies, yet their results are difficult to trace through input-data, model structure, scenario configurations, and solved outputs. Because this information is siloed across disconnected artifacts, process-based IAMs have historically lacked a unified, queryable representation. Such lack of traceability prevents researchers from systematically isolating the multi-sector drivers of complex outcomes (such as tracing water-scarcity results back to distant energy-system dynamics) or conducting holistic uncertainty attribution across hundreds of interacting parameters. To address this concern, our work documents the software engineering process of a knowledge graph that unifies these four layers for the Global Change Analysis Model (GCAM-USA_Reference scenario, GCAM v9.1). The graph was built as a relational property graph in DuckDB from the run’s own artifacts: the input-preparation dependency map (gcamdata chunk map), the model’s XML input files, the run configuration, and the results database (BaseX), successfully mapping the model’s declared structure. The resulting graph comprises 204,321 nodes and 1,687,814 edges across 16 node types and 15 edge types, with approximately 16.3 million time-series values stored separately to maintain structural efficiency. To ensure representation fidelity, every edge carries an epistemic-status annotation recording the warrant for the relationship (structural, provenance, dependency, or model-derived), and a machine-readable provenance ledger classifying the origin of every schema element. Evaluation against a fixed five-benchmark suite with locked baselines reports zero structural orphans, zero dangling edge endpoints, and 100% of output-producing technologies traceable to raw input files. Two interactive interfaces present the graph, including a serverless browser application built on DuckDB-Wasm. By establishing the first end-to-end provenance framework for an IAM, this work enables researchers and scientists to systematically audit complex policy scenarios, debug model structures, and trace policy-relevant outputs to their data origins in real time.

Artifical Intelligence