Search NASA⌕ Search

SEARCH · Search NASA

Results for “raw data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

An R shiny graphical user interface for highprecision mass spectrometric data analysis

• There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2, IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

LABONE, ELIZABETH↗

SUBTASK 1.6 – BASIN ELECTRIC CARBON STORAGE RESEARCH PROJECT: NOVEL MONITORING TECHNIQUES

The Energy & Environmental Research Center (EERC) conducted baseline activities associated with an applied research project at Basin Electric Power Cooperative’s (Basin’s) carbon capture and storage (CCS) site in Beulah, North Dakota, to establish novel carbon storage-monitoring techniques as commercial methods under Cooperative Agreement No. DE-FE0024233, Subtask 1.6. The following report summarizes the baseline activities performed and briefly describes the subsequent (operational monitoring) activities that have been proposed to the U.S. Department of Energy (DOE) as part of the overall project to develop and demonstrate novel monitoring techniques at North America’s largest permitted CCS operation. Dakota Gasification Company (DGC), a wholly owned subsidiary of Basin, owns and operates the Great Plains Synfuels Plant (GPSP) approximately 5 miles northwest of the town of Beulah, North Dakota (Figure 1). In 2023, DGC received approval from the North Dakota Industrial Commission (NDIC) to develop a storage facility on-site for injecting a stream of carbon dioxide (CO2) captured from GPSP. DGC will transport the captured CO2 stream with approximately 6.8 miles of transmission lines that extend north of GPSP and inject >1 million tonnes (MMt) of CO2 annually (>1 MMt/yr) over a 12-year period with up to six underground injection control (UIC) Class VI-compliant injection wells completed in the Broom Creek Formation, a predominantly sandstone reservoir and saline aquifer underlying GPSP. The Broom Creek Formation lies approximately 5900 feet (ft) below ground surface (bgs) at GPSP. The commercial scale (i.e., >1 MMt/yr) of DGC’s permitted carbon storage project is ideal for developing and testing the novel monitoring techniques included within Subtask 1.6. The goals of this project are to demonstrate 1) the cost-effectiveness of novel monitoring technologies included as part of this research, 2) technology capability for tracking the CO2 plume and/or associated pressure response in the subsurface and monitoring out-of-zone migration, and 3) compliance with UIC Class VI program requirements. The research activities proposed for the overall project include 1) design of an automated, integrated, modular (AIM) monitoring station; 2) time-lapse electromagnetic (EM) field surveys; 3) drone-based surveillance studies; 4) time-lapse monitoring with seismic methods; 5) advanced wellbore-monitoring methods; 6) deployment of an AIM monitoring network; 7) EM monitoring of CO2 with real-time data processing; 8) continued seasonal drone-based surveillance studies; 9) seismic monitoring with passive and active surveys; and 10) wellbore monitoring with nuclear magnetic resonance (NMR) for near-surface characterization. Completion of Activities 1.0–5.0 (baseline activities) are described in this report. Upon authorization of funding by DOE, the EERC will initiate Activities 6.0– 10.0 (operational monitoring activities). Current state-of-the-art (SOA) carbon storage-monitoring techniques require countless labor hours dedicated to the acquisition of data. Once data are gathered, these SOA techniques often rely on commercial facilities to process raw data from the field. However, it is anticipated that next-generation monitoring techniques, such as those being demonstrated, will lower acquisition footprints, be less operationally intensive, and improve data acquisition efficiencies. These new techniques are more conducive to the application of machine learning, artificial intelligence, and automation, thus providing a pathway for integration into active control systems, informing site operability, and improving the integration of data for future CCS projects across the United States. Additionally, reclaimed and active mining lands are present within the project site, creating a unique opportunity to demonstrate the effectiveness of remote sensing and surface-based geophysics monitoring techniques at similar project sites that may include disturbed, unconsolidated, or actively excavated near-surface environments. The efforts included in the overall project will produce necessary designs, learnings, and data acquired during the baseline and operational monitoring periods that are necessary for time-lapse demonstration and validation of the described monitoring techniques. In addition, it is anticipated that the monitoring technologies included in this study will be compliant with UIC Class VI requirements to enable the potential for implementation at other CCS sites across the United States.

42 ENGINEERING↗

TSDC: Transportation Secure Data Center: Real-World Data for Planning, Modeling, and Analysis

The Transportation Secure Data Center is a centralized repository for high-resolution transportation data from hundreds of travel and transit surveys and studies. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. It houses surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Meanwhile, the Livewire Data Platform empowers research, industry, and academic partners to easily and securely preserve, maintain, share, discover, and gain access to transportation and mobility data. Livewire accommodates a range of datasets, including behavioral, experimental, model, analytical, and raw data at the vehicle, traveler, and system levels. Datasets support mobility research and planning spanning urban science, connected and automated vehicles, fueling and charging infrastructure, mobility decision science, multimodal transportation, vehicle efficiency, and more.

33 ADVANCED PROPULSION SYSTEMS↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Data Selection Improvement For MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal, and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. We also need to choose the right reconstruction method, a tool to take raw data from the detector and convert it into physics results. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions, and compared them to the truth information directly from the MicroBooNE simulation pipeline. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Fermilab]↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Zero-Power Analog Optical Processing

The motivation behind this research is the growing challenge of handling the massive amounts of data generated by modern imaging systems. Conventional digital image processing techniques are struggling to keep pace with the demands of high-resolution and high-speed imaging systems for remote sensing due to their high-power consumption and data storage requirements. We present a novel approach based on analog photonics to address this challenge. The proposed system utilizes a silicon-photonics-based image encoder positioned after image formation and initial optical-to-electrical conversion. The photonic encoder compresses image data using a passive disordered photonic structure to perform kernel-type random projections of the raw data. The compressed data is then processed by a back-end neural network, which reconstructs the original image with high fidelity (structural similarity exceeding 90%). Our proposed approach has the potential to compress images with ~ 1000X lower power consumption compared to digital approaches with data rates exceeding 1 terapixel/second.

97 MATHEMATICS AND COMPUTING↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

Data for reproducing the figures of the paper Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas

This deposit contains the raw data for reproducing research results of the paper Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas. The main contribution of this work is to utilize machine learning techniques to reconstruct and enhance the resolution of a diagnostic measurement from other available diagnostics in a system. The proposed techniques is called Diag2Diag.

diag2diag↗

FL‐ADS: Federated learning anomaly detection system for distributed energy resource networks

Abstract With the ongoing development of Distributed Energy Resources (DER) communication networks, the imperative for strong cybersecurity and data privacy safeguards is increasingly evident. DER networks, which rely on protocols such as Distributed Network Protocol 3 and Modbus, are susceptible to cyberattacks such as data integrity breaches and denial of service due to their inherent security vulnerabilities. This paper introduces an innovative Federated Learning (FL)‐based anomaly detection system designed to enhance the security of DER networks while preserving data privacy. Our models leverage Vertical and Horizontal Federated Learning to enable collaborative learning while preserving data privacy, exchanging only non‐sensitive information, such as model parameters, and maintaining the privacy of DER clients' raw data. The effectiveness of the models is demonstrated through its evaluation on datasets representative of real‐world DER scenarios, showcasing significant improvements in accuracy and F1‐score across all clients compared to the traditional baseline model. Additionally, this work demonstrates a consistent reduction in loss function over multiple FL rounds, further validating its efficacy and offering a robust solution that balances effective anomaly detection with stringent data privacy needs.

Purohit, Shaurya [Iowa State University Ames Iowa ↗

Dataset for manuscript "Equipartition and the temperature of maximum density of TIP4P/2005 water"

We simulate TIP4P/2005 water in the temperature range of 257 K to 318 K with time-steps 0.25, 0.50, 1.00, 2.00, and 4.00 fs. The density-temperature behavior obtained using 0.25 or 0.50 fs are in excellent agreement with each other but differ from those obtained using time-steps that have been shown earlier to lead to a breakdown of equipartition. The temperature of maximum density (TMD) is 277.15 K with time-step 0.25 or 0.50 fs, but is shifted to progressively lower values for longer time-steps, a trend that holds for different thermostat/barostat combinations. Enhancing the water-water dispersion interaction, as has been recommended for simulating disordered proteins in TIP4P/2005, degrades the description of the liquid-vapor phase envelope. We present a simple physically transparent reasoning to highlight the separation of the time-scales between translational and rotational motion. We also develop a metric, Chi, that we term the equipartition anomaly, to detect equipartition violations in simulations that include molecules that are treated as rigid objects. Calculating Chi is shown to be straightforward and sensitive to equipartition violations. A key takeaway from this study is that using sufficiently short time-steps (less than or equal to 0.5 fs) to preserve equipartition is essential for obtaining meaningful liquid water properties and for producing reliable simulation data, as correct-ensemble sampling is fundamental to ensure reproducibility across codes and simulation alogrithms. The included dataset provides the raw data used in the preparation of the graphs noted in the manuscript.

36 MATERIALS SCIENCE↗

Vanderbilt CMS Heavy-Ion Tier-2 Facility (Final Report)

The CMS experiment at the LHC at CERN in Geneva, Switzerland spends a portion of its running time colliding heavy ions. These heavy ion collisions are a key probe of a state of matter known as a Quark Gluon Plasma, which is a state of matter found in the very early universe 10s of microseconds after the big bang. These collision data are recorded by the CMS experiment and then undergo several stages of refinement to extract physics properties of the collisions from the raw data recorded by the various detector subsystems in CMS. In addition, copies of these datasets are stored on tape archival storage for data preservation. This award funds the deployment, maintenance, and operation of a computing facility at Vanderbilt University which processes, stores, and analyzes these collision data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data for "Which plant traits increase soil carbon sequestration? Empirical evidence from a long-term poplar genetic diversity trial"

This archive contains all data and code used by the following publication: Field, J. L., Sloan, B. P., Craig, M. E., Calloway, P., Ottinger, S. L., Mead, T., Abramoff, R. Z., Venegas, M. P., Chhetri, H. B., Haiby, K., Kalluri, U. C., Muchero, W., Schadt, C. W., & Mayes, M. A. (2025). Which plant traits increase soil carbon sequestration? Empirical evidence from a long-term poplar genetic diversity trial (p. 2025.02.17.638464). bioRxiv. https://doi.org/10.1101/2025.02.17.638464 Our analysis combined several soil and root data sets collected by Oak Ridge National Laboratory (ORNL) researchers/collaborators from the Clatskanie Poplar Common Garden in Clatskanie, OR by from 2009-2024. The raw data data files are located */02-data/01-raw/* which we harmonized using the codes in */01-codes/01-harmonize-clatskanie-data-pub.qmd*. The final processed data set used in the paper is found at */02-data/02-processed/clatskanie-c-fit-data.csv* and its columns are described in the table below.

Sloan, Brandon [ORNL] (ORCID:0000000316304271)↗

Solovay-Kitaev Algorithm and Randomized Compilation Data Availability

This zipped folder contains simulation notebooks, simulated data, and experimental data from the QSCOUT trapped-ion device that were used in the publication "Solovay-Kitaev Algorithm and Randomized Compilation" (https://doi.org/10.1103/ll6m-dbl7). The raw data is in the form of measurement outcomes of simple tomographic quantum circuits that were executed on the QSCOUT device and simulated using JAQALPAQ. These data are used to create plots within the jupyter notebooks that were included in the publication.

Quantum benchmarking↗

Real-time data processing for serial crystallography experiments

We report the use of streaming data interfaces to perform fully online data processing for serial crystallography experiments, without storing intermediate data on disk. The system produces Bragg reflection intensity measurements suitable for scaling and merging, with a latency of less than 1 s per frame. Our system uses the CrystFEL software in combination with the ASAP::O data framework. In a series of user experiments at PETRA III, frames from a 16 megapixel Dectris EIGER2 X detector were searched for peaks, indexed and integrated at the maximum full-frame readout speed of 133 frames per second. The computational resources required depend on various factors, most significantly the fraction of non-blank frames ('hits'). The average single-thread processing time per frame was 242 ms for blank frames and 455 ms for hits, meaning that a single 96-core computing node was sufficient to keep up with the data, with ample headroom for unexpected throughput reductions. Further significant improvements are expected, for example by binning pixel intensities together to reduce the pixel count. We discuss the implications of real-time data processing on the `data deluge' problem from recent and future photon-science experiments, in particular on calibration requirements, computing access patterns and the need for the preservation of raw data.

47 OTHER INSTRUMENTATION↗

GAT (Grid Analysis Toolkit) [SWR-25-41]

Grid Analysis Toolkit (GAT) is a unified Python API and plotting for power system PCM and CEM results (Sienna, PLEXOS, ReEDS™). It's a toolkit for wrangling data for Bulk Grid Dispatch and Transmission Analysis. GAT aims to provide simplified access to PCM and CEM results in a standard format while also allowing raw data access to underlying datasets specific to the model. This software can also be found on PyPI at For plotting, GAT defaults to standard National Lab of the Rockies (NLR) color schemes and standard styles while allowing customization.

Webb, Micah [National Laboratory of the Rockies (N↗