Search NASA⌕ Search

SEARCH · Search NASA

Results for “raw data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Online and Offline Data Quality Monitoring for the Mu2e Calorimeter

This thesis presents the design, implementation, and validation of a calorimeter Data Quality Monitoring (DQM) toolchain for the Mu2e experiment at Fermilab. Mu2e searches for charged lepton flavor violation via coherent muon-to-electron conversion in the field of an aluminum nucleus, $\mu^- Al \rightarrow e^-Al$, a process whose observation would constitute clear evidence of physics beyond the Standard Model. Achieving target sensitivity requires stringent control of detector performance and data integrity during acquisition, as subtle issues in readout configuration, data formatting, or electronics behavior can compromise reconstruction and bias downstream analyzes. To address these challenges, this work develops a multi-layer DQM approach spanning both raw data validation and reconstructed digi-level diagnostics. At the low level, a fragment analysis component performs word- and bit-field decoding of calorimeter readout blocks, enabling sanity checks of the expected structure and producing detailed error and integrity statistics useful for commissioning and troubleshooting. At the digi level, the CaloDigiDQM analyzer is implemented within the art framework and transforms each CaloDigiCollection into a structured hierarchy of ROOT histograms designed for fast drill-down diagnostics. The module generates coherent monitoring views at global, disk, board, and channel granularity, including occupancy, waveform-derived features (baseline, RMS, peak amplitude and position), and left-right sensor consistency metrics. Detector-aware channel-to-electronics mapping is performed through the conditions system (CaloDAQMap), ensuring that diagnostics remain aligned with hardware identifiers used in operations. For end-to-end testing without reliance on live DAQ data, a synthetic CaloDigi producer is developed to generate realistic waveforms with controlled noise and pulse shapes. The resulting system supports both offline ROOT-file production and online operation, including optional histogram streaming through otsdaq via ots::HistoSender. This toolchain provides a practical and scalable foundation for calorimeter commissioning and stable data collection, enabling early detection of anomalies and reducing operational risk for Mu2e.

Vakulenko, Mark [Drew U.] (ORCID:0009000276197818)↗

WFIP3 - SHIP site - NREL Profiling Lidar (Windcube v2.1) / Reviewed data

This dataset contains reviewed data from the profiling lidar (Windcube v2.1) deployed on WFIP3's SHIP. The reviewed files herein are based on the lidar's RTD files (i.e., the real-time raw data files at near 1 Hz resolution). The data have been corrected for the motion of the ship.

17 WIND ENERGY↗

WFIP3 - BARG site - NREL Profiling Lidar (Windcube v2.1) / Reviewed Data

This dataset contains reviewed data from the profiling lidar (Windcube v2.1) deployed on WFIP3's barge. The reviewed files herein are based on the lidar's RTD files (i.e., the real-time raw data files at near 1 Hz resolution). The data have been corrected for the motion of the ship.

17 WIND ENERGY↗

Supporting Data for "On the magnetic contribution of itinerant electrons to neutron diffraction in the topological antiferromagnet CeAlGe"

Contents of this DOI are data for the research paper "On the magnetic contribution of itinerant electrons to neutron diffraction in the topological antiferromagnet CeAlGe" by the authors: V. Pomjakushin et al. The dataset contains raw data from neutron scattering experiments of CeAlGe powder performed at the CNCS spectrometer at Spallation Neutron Source at ORNL.

36 MATERIALS SCIENCE↗

Legacy Survey of Space and Time Data Preview 1: raw dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the raw dataset type. These are unprocessed images from the LSST Commissioning Camera. This release contains 16,125 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Federated learning for 2D synchrotron x-ray diffractometry: a cross-institutional approach for phase quantification of Ti–6Al–4V alloy

High-energy Two dimensional (2D) synchrotron x-ray diffractometry provides important insights into the atomistic structure and phase evolution of materials, yet traditional analysis methods remain complex, knowledge-intensive, and computationally demanding. Deep-learning models offer a powerful alternative for automating their analysis. Institutions that hold these datasets may be unwilling to share their data due to privacy and security policies, as well as the challenges associated with large-scale data transfer. As a result, models trained on local datasets often perform well only on their own data but exhibit bias and poor generalization across different instruments or facilities. To overcome these limitations, we explore federated learning (FL) for 2D synchrotron diffractograms, enabling collaborative model training without exchanging raw data. In this study, 2D synchrotron diffractograms of Ti–6Al–4V alloy collected from two independent facilities are used to train convolutional neural networks for predicting the β-phase volume fraction. Experimental results show that federated global models significantly outperform locally trained models in terms of generalization and achieve accuracy comparable to centralized trained models. These findings demonstrate the potential of FL to enable secure, cross-institutional collaboration and enhance the scalability of deep-learning-based materials characterization.

36 MATERIALS SCIENCE↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

QProR: An Efficient Framework for Quantity-of-Interest Based Progressive Retrieval with Guaranteed Error Control

Scientific applications generate an unprecedented volume of data, overwhelming the network and file systems’ bandwidth and posing challenges for efficient and scalable data retrieval and analysis. Progressive data compression offers a promising solution by enabling on-demand retrieval at reduced size. However, existing progressive methods either fail to bound the errors in essential quantities of interest (QoIs) derived from raw data or suffer from suboptimal retrieval efficiency. In this work, we propose QProR, an efficient QoI-based progressive framework that optimizes progressive retrieval for target QoIs. Our key contributions include: (1) a systematic framework that integrates error-controlled lossy compressors with bitplane encoding while decoupling the two processes for high flexibility and adaptability; (2) a novel weighted bitplane encoding method which incorperates QoI knowledge into data refactoring to enhance retrieval efficiency; (3) an optimized retrieval strategy that accounts for the varying impacts of different variables on multivariate QoIs; (4) comprehensive evaluations using six real-world datasets from multiple scientific applications and thorough comparisons against state of the arts. Experimental results demonstrate that QProR achieves up to 80.38% reduction in the retrieval size under the same requested QoI error tolerance, when compared with the best-performing existing methods. When transferring 384 GB of scientific data to remote sites, QProR delivers up to 1.68 × speedup in the end-to-end data transfer performance.

Li, Wenbo [University of Kentucky]↗

Arctic shrub and Eriophorum leaf and root decomposition, northern Alaska, 2017-2018

This data package contains litter decomposition data collected from 170 plots of rapidly expanding shrub genera (Alnus, Betula, and Salix) and a widespread sedge (Eriophorum vaginatum) along a latitudinal and temperature gradient in northern Alaska. These data were produced from a litter bag experiment that took place from July 2017 to July 2018 and include mass loss and nitrogen loss decomposition metrics for both leaf and root litters. These raw data support a submitted manuscript that examines the variability in decomposition between shrub and graminoid leaf and root litters across a 1-year experiment across the graminoid-dominated Arctic tundra and reveals how deciduous shrub expansion affects litter decomposition in tundra ecosystems. Data are presented by site (n=5) and patch (shrub or sedge plot) in csv files. The site and plot location data and environmental measurement data are provided in Fraterrigo and Chen (2020). Additional methods regarding plot distribution and environmental measurements are in Chen et al. (2020) and Fraterrigo et al. (2024).

54 ENVIRONMENTAL SCIENCES↗

Intelligent experiments through real-time AI: Fast Data Processing and Autonomous Detector Control for sPHENIX and future EIC detectors (Phase-I)

With an ever increasing demand for high precision data from modern detectors for discovery science and precision measurements, all major high energy nuclear and particle experiments, current and future, are facing the challenge on how to deal with the large volume of raw data generated from sophisticated state-of-the-art detectors in high rate collisions. These goals need to be balanced with available hardware and cost limits on DAQ (Data AcQuisition system) bandwidth and offline computing resources to capture, store and process the signal events. Two prototypical examples are the upcoming sPHENIX experiment, the DOE next generation heavy ion physics experiment at the Relativistic Heavy Ion Collider at BNL, and the future EIC experiments that are planned to be online circa 2030.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

LCLS RF Station Phase Anomaly Candidate Dataset

A public anomaly detection dataset constructed from RF station faults for phase at SLAC's LCLS (Linac Coherent Light Source). We have compiled a dataset of the RF station diagnostic phase data and the beam-position monitor (BPM) signals, alongside the hand labels, for a labeled study period. The dataset consists of two HDF5 files (one for train and one for test) containing the raw data, two CSV files containing information about the candidates. The CSV file for the test dataset also contains the label.

Liang, Jia [Stanford Univ., CA (United States). In↗

Investigation of fast and efficient lossless compression algorithms for macromolecular crystallography experiments

Structural biology experiments benefit significantly from state-of-the-art synchrotron data collection. One can acquire macromolecular crystallography (MX) diffraction data on large-area photon-counting pixel-array detectors at framing rates exceeding 1000 frames per second, using 200 Gbps network connectivity, or higher when available. In extreme cases this represents a raw data throughput of about 25 GB s −1 , which is nearly impossible to deliver at reasonable cost without compression. Our field has used lossless compression for decades to make such data collection manageable. Many MX beamlines are now fitted with DECTRIS Eiger detectors, all of which are delivered with optimized compression algorithms by default, and they perform well with current framing rates and typical diffraction data. However, better lossless compression algorithms have been developed and are now available to the research community. Here one of the latest and most promising lossless compression algorithms is investigated on a variety of diffraction data like those routinely acquired at state-of-the-art MX beamlines.

36 MATERIALS SCIENCE↗

Differentially Private Map Matching (DPMM) v1.0

Human mobility trajectories provide valuable information for developing mobility applications, as they contain diverse and rich information about the users. User mobility data is valuable for various applications such as intelligent transportation systems (ITS), commercial business models, and disease-spread models. However, such spatio-temporal traces may pose a threat to user privacy. GPS trajectories in their raw form are not suitable for transportation studies, as they require matching locations with nearest road links — a process called map-matching. This software implements a differential privacy (DP)-based map-matching algorithm, called DPMM, that generates link-level location trajectories in a privacy-preserving manner to protect users' origin destinations (OD) and travel paths. OD privacy is achieved by injecting Planar Laplace noise to the user OD GPS points. Travel-path privacy is provided with randomized travel path construction using exponential DP mechanism. The injected noise level is selected adaptively, by considering the link density of the location and the functional category of the localized links. For path privacy, our mechanism samples waypoints and selects candidate paths between waypoints. DPMM provides privacy effectively with respect to link density instead of other trajectory samples in the database compared to other privacy mechanisms. Compared to the different baseline models our DP-based privacy model offers closer query responses to the raw data in terms of individual and aggregate trajectory-level statistics with an average at absolute deviation from the baseline for individual statistics on ϵ = 1.0. Beyond individual trajectory statistics, the DPMM outperforms the other benchmark DP-based mechanisms on different aggregate statistics with up to 8x improvement in utility.

Peisert, Sean [Lawrence Berkeley National Laborato↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

fluxfinder: An R Package for Reproducible Calculation and Initial Processing of Greenhouse Gas Fluxes From Static Chamber Measurements

Fluxes of greenhouse gases are a critical component of the earth's natural climate, but anthropogenic emissions have created an imbalance and resulted in global climate change. Quantifying the emission of these gases is vital to our understanding of their sources and sinks, both natural and anthropogenic. The static chamber method, in which a system of interest is enclosed, and gas concentrations are measured over time, is widely used to estimate fluxes of greenhouse gases. With the development of instruments such as infrared gas analyzers (IRGAs) supporting high-frequency concentration data, there is a growing need for open-source workflows to calculate fluxes. Here we present fluxfinder, an R package designed to support reproducible calculations and processing of greenhouse gas fluxes measured with the static chamber method. The package includes raw data file parsing from widely used IRGAs, metadata matching, unit conversion, flux estimations, and initial quality assurance/quality control (QA/QC). Diagnostic graphical plots provide a transparent way to differentiate between measurement issues and nonlinear behavior. The package is also designed to be easily integrated with the gasfluxes package for further fitting of nonlinear concentration-time models, allowing alternative or additional flux QA/QC. The fluxfinder package offers a flexible workflow that is easily adaptable to promote open and reproducible greenhouse gas flux estimations.

Wilson, Stephanie J.↗

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM↗

1H-NMR characterization of soil dissolved organic matter from soil samples in control and warming plots in Blodgett Forest, CA (2014 and 2018)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory Terrestrial Ecosystem Science Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM (soil organic matter) decomposition and stabilization. This package contains metabolite data obtained through 1H nuclear magnetic resonance (NMR) spectroscopy on water-extracted soils. Soil samples were collected in 2014/06/03 and 2018/06/04 from 3 replicated paired plots that had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. The following files are included: (1) nmr_h2o_data_raw.csv: raw data, (2) nmr_h2o_data_processed.csv: computed compound concentrations and metadata, (3) nmr_h2o_compound_metadata.csv: compound metadata, (4) nmr_h2o_sample_metadata.csv: sample metadata

1H-NMR (nucleic magnetic resonance) spectroscopy↗