Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data processing methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

WHONDRS Surface Water Geochemistry and Organic Matter Characterization Data from Streams Distributed across Latin America

This dataset supports a broader study examining global transferability of stream biogeochemistry and was generated in collaboration with the MicroSudAqua (µSudAqua) network (https://microsudaqua.netlify.app/en/). The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen, cations) and organic matter characterization (FTICR-MS) from streams in Argentina, Brazil, Chile, and Colombia. Samples were collected across stream orders (1st to 6th order) within five basins. Related data were collected and will be published separately in collaboration with the µSudAqua network. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos; (2) a folder of surface water sample data, (3) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data; (4) file-level metadata; (5) data dictionary; (6) field metadata; (7) readme; (8) international generic sample number (IGSN) mapping file; and (9) field protocol. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total dissolved nitrogen data and averages; (3) anions and averages; (4) methods codes; (5) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4.

Anions↗

Integrated Off-Gas System Tests on the DM1200 Melter with RPP-WTP LAW Sub-Envelope C1 Simulants (Final Report)

The report documents melter and off-gas performance results obtained on the DM1200 during processing of LAW Sub-Envelope Cl feed. The principal objectives of the DM1200 melter testing were to demonstrate suitable processing and production rates for LAW Cl feed; characterize the glass product for elemental composition, the presence of organic compounds, and TCLP response for metals; and measure destruction and removal efficiencies (DREs) for organic compounds across the melter and off-gas treatment components. The feed was spiked with organic compounds for part of the testing duration. Data were collected for both process engineering needs and to support regulatory activities. The data needed for regulatory support were collected using EPA methods and independent data analyses and data checking were performed.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

SUBTASK 1.6 – BASIN ELECTRIC CARBON STORAGE RESEARCH PROJECT: NOVEL MONITORING TECHNIQUES

The Energy & Environmental Research Center (EERC) conducted baseline activities associated with an applied research project at Basin Electric Power Cooperative’s (Basin’s) carbon capture and storage (CCS) site in Beulah, North Dakota, to establish novel carbon storage-monitoring techniques as commercial methods under Cooperative Agreement No. DE-FE0024233, Subtask 1.6. The following report summarizes the baseline activities performed and briefly describes the subsequent (operational monitoring) activities that have been proposed to the U.S. Department of Energy (DOE) as part of the overall project to develop and demonstrate novel monitoring techniques at North America’s largest permitted CCS operation. Dakota Gasification Company (DGC), a wholly owned subsidiary of Basin, owns and operates the Great Plains Synfuels Plant (GPSP) approximately 5 miles northwest of the town of Beulah, North Dakota (Figure 1). In 2023, DGC received approval from the North Dakota Industrial Commission (NDIC) to develop a storage facility on-site for injecting a stream of carbon dioxide (CO2) captured from GPSP. DGC will transport the captured CO2 stream with approximately 6.8 miles of transmission lines that extend north of GPSP and inject >1 million tonnes (MMt) of CO2 annually (>1 MMt/yr) over a 12-year period with up to six underground injection control (UIC) Class VI-compliant injection wells completed in the Broom Creek Formation, a predominantly sandstone reservoir and saline aquifer underlying GPSP. The Broom Creek Formation lies approximately 5900 feet (ft) below ground surface (bgs) at GPSP. The commercial scale (i.e., >1 MMt/yr) of DGC’s permitted carbon storage project is ideal for developing and testing the novel monitoring techniques included within Subtask 1.6. The goals of this project are to demonstrate 1) the cost-effectiveness of novel monitoring technologies included as part of this research, 2) technology capability for tracking the CO2 plume and/or associated pressure response in the subsurface and monitoring out-of-zone migration, and 3) compliance with UIC Class VI program requirements. The research activities proposed for the overall project include 1) design of an automated, integrated, modular (AIM) monitoring station; 2) time-lapse electromagnetic (EM) field surveys; 3) drone-based surveillance studies; 4) time-lapse monitoring with seismic methods; 5) advanced wellbore-monitoring methods; 6) deployment of an AIM monitoring network; 7) EM monitoring of CO2 with real-time data processing; 8) continued seasonal drone-based surveillance studies; 9) seismic monitoring with passive and active surveys; and 10) wellbore monitoring with nuclear magnetic resonance (NMR) for near-surface characterization. Completion of Activities 1.0–5.0 (baseline activities) are described in this report. Upon authorization of funding by DOE, the EERC will initiate Activities 6.0– 10.0 (operational monitoring activities). Current state-of-the-art (SOA) carbon storage-monitoring techniques require countless labor hours dedicated to the acquisition of data. Once data are gathered, these SOA techniques often rely on commercial facilities to process raw data from the field. However, it is anticipated that next-generation monitoring techniques, such as those being demonstrated, will lower acquisition footprints, be less operationally intensive, and improve data acquisition efficiencies. These new techniques are more conducive to the application of machine learning, artificial intelligence, and automation, thus providing a pathway for integration into active control systems, informing site operability, and improving the integration of data for future CCS projects across the United States. Additionally, reclaimed and active mining lands are present within the project site, creating a unique opportunity to demonstrate the effectiveness of remote sensing and surface-based geophysics monitoring techniques at similar project sites that may include disturbed, unconsolidated, or actively excavated near-surface environments. The efforts included in the overall project will produce necessary designs, learnings, and data acquired during the baseline and operational monitoring periods that are necessary for time-lapse demonstration and validation of the described monitoring techniques. In addition, it is anticipated that the monitoring technologies included in this study will be compliant with UIC Class VI requirements to enable the potential for implementation at other CCS sites across the United States.

42 ENGINEERING↗

Carbon fiber classification using raman spectroscopy

Carbon fiber characterization processes are described that include multi-condition Raman spectroscopy-based examination combined with multivariate data analyses. Methods are a nondestructive material characterization approach that can provide predictions as to carbon fiber bulk physical properties, as well as identification of unknown carbon fiber materials for quality control purposes. The framework of the multivariate analysis methods includes a principal component-based identification protocol including comparison of Raman spectral data from an unknown carbon fiber with a data library of multiple principal component spaces.

Houk, Amanda L.↗

WHONDRS Surface Water and Sediment Geochemistry and Organic Matter Characterization Data from Streams across HJ Andrews Experimental Forest, Oregon (v2)

This dataset supports a broader study developing conceptual models for river corridor critical zone processes across spatial scales and was generated in collaboration with the HJ Andrews River Corridor Critical Zone Workshop in 2025. The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen) from 48 sites across the HJ Andrews Experimental Forest, Oregon (https://andrewsforest.oregonstate.edu). Some of the sites have been impacted by the Holiday Farm Fire and the Lookout Fire in 2020 and 2023, respectively. Related data were collected as part of the workshop and will be published separately in collaboration with other workshop attendees and available at http://www.hydroshare.org/resource/b274c4a234bf4b12b7cb8a54a696c629. Related genomic data can be found on the National Center for Biotechnology Information (NCBI) under BioProject PRJNA1503030 (see critical details section below for more information). Additional related data collected in 2016 from a similar effort can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3377027 and http://www.hydroshare.org/resource/ea6c0832885a46c3939e7bb22e48e754 and are described within https://doi.org/10.5194/essd-11-1-2019 (Ward et al., 2019). This data package was originally published in March 2026. It was updated in August 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos, (2) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, (3) a data checks report, (4) a folder of sample data, (5) file-level metadata, (6) data dictionary, (7) field metadata, (8) readme, (9) international generic sample number (IGSN) mapping file; and (10) field protocol. The sample data subfolder contains surface water and sediment (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages, (2) total dissolved nitrogen data and averages, (3) methods codes, (4) FTICR-MS methods; and (5) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the CoreMS processed data and seven subfolders, thee containing .xml files for each sample type (sediment, surface water and blank samples), three containing the sediment CoreMS output files for each sample type (sediment, surface water and blank samples), and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .Rmd, .py, .cal, .json, .jpg, or .jpeg.

Biogeochemistry↗

Survey-wide asteroid discovery with a high-performance computing enabled non-linear digital tracking framework

Modern astronomical surveys detect asteroids by linking together their appearances across multiple images taken over time. This approach faces limitations in detecting faint asteroids and handling the computational complexity of trajectory linking. Here, we present a novel method that adapts “digital tracking” – traditionally used for short-term linear asteroid motion across images – to work with large-scale synoptic surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time (Rubin/LSST). Our approach combines hundreds of sparse observations of individual asteroids across their non-linear orbital paths to enhance detection sensitivity by several magnitudes. To address the computational challenges of processing massive data sets and dense orbital phase spaces, we developed a specialized high-performance computing architecture. We demonstrate the effectiveness of our method through experiments that take advantage of the extensive computational resources at Lawrence Livermore National Laboratory. This work enables the detection of significantly fainter asteroids in existing and future survey data, potentially increasing the observable asteroid population by orders of magnitude across different orbital families, from near-Earth objects (NEOs) to Kuiper belt objects (KBOs).

Asteroid discovery↗

Multi-Scale Integrated Monitoring System for Enhancing Methane Emission Detection, Quantification & Prediction

This report details the progress and findings of a comprehensive study on reviewing existing solutions, identifying technology gaps, and formulating an “all-in-one” integrated strategy for developing the next-generation multiscale methane monitoring and modeling platform, conducted under grant number DE-FE0032292. Co-led by Dr. David Ebert, Dr. Binbin Weng, and Dr. Chenghao Wang at the University of Oklahoma, the project’s goal was to develop an integrated approach for building this engineering platform to detect, quantify, and mitigate methane emissions across various temporal scale, spatial scales, and sectors. The planning grant study began with an extensive review of various methane sensing and monitoring technologies and systems, surveying over 100 technology providers globally. This review revealed the prevalence of optical methods over chemical methods in commercially available sensors, with Non-Dispersive Infrared (NDIR), Tunable Diode Laser Absorption Spectroscopy (TDLAS), and Optical Gas Imaging (OGI) cameras being the most prevalent options. A trend towards more advanced optical techniques was observed, driven by increased regulatory focus and technological advancements. The technical evaluation of these sensing technologies provided crucial insights into their capabilities and limitations. The study examined emerging technologies such as Differential Absorption LiDAR (DIAL), which show promise for high-precision and long-range detection. The team then investigated the features and application bandwidth of various sensing platforms, including handheld, fixed/stationary, mobile, aerials, and spaceborne monitors. Pilot field studies were conducted to assess the capabilities of solutions for different emission scenarios. Field work with sensor deployments was conducted at three distinct site types: an oil & gas industry site, a cattle ranching operation, and a waste processing facility. The team also conducted a thorough review of methane flux inverse modeling approaches, focused on physically based methods. These approaches were categorized into simple, intermediate, and advanced methods. A realtime WRF-GHG (Weather Research and Forecasting-Greenhouse Gas) modeling system was developed and applied, incorporating multiple data sources to guide field experiments and inform methane plume detection. The project identified and analyzed numerous categories of methane data sources, including satellite measurements, ground-based sensors, and inventory databases. Key platforms examined include EDGAR, EPA GHGI, NASA TROPOMI, Carbon Mapper, and Climate TRACE, among others. The team proposed an architecture for a comprehensive methane monitoring platform. This system incorporates multi-source data acquisition, advanced data processing and assimilation, interactive visualization tools, and analytical capabilities for emissions forecasting and scenario analysis. The proposed platform aims to provide a user-friendly interface catering to various stakeholders, from researchers to policymakers. The architecture includes sophisticated data ingestion methods, a centralized data warehouse, and advanced analytical tools for data fusion and interpretation. To ensure the relevance and effectiveness of the proposed system, a comprehensive survey was conducted to gather stakeholder input on system requirements. Key findings include a strong need for integrating various data types and formats, a preference for real-time data updates and advanced visualization tools, and a demand for user-friendly interfaces catering to different expertise levels.

03 NATURAL GAS↗

Bayesian stability and force modeling for uncertain machining processes

Accurately simulating machining operations requires knowledge of the cutting force model and system frequency response. However, this data is collected using specialized instruments in an ex-situ manner. Bayesian statistical methods instead learn the system parameters using cutting test data, but to date, these approaches have only considered milling stability. This paper presents a physics-based Bayesian framework which incorporates both spindle power and milling stability. Initial probabilistic descriptions of the system parameters are propagated through a set of physics functions to form probabilistic predictions about the milling process. The system parameters are then updated using automatically selected cutting tests to reduce parameter uncertainty and identify more productive cutting conditions, where spindle power measurements are used to learn the cutting force model. The framework is demonstrated through both numerical and experimental case studies. Results show that the approach accurately identifies both the system natural frequency and cutting force model.

42 ENGINEERING↗

Kernel methods for evolution of generalized parton distributions

Generalized parton distributions (GPDs) characterize the 3-dimensional structure of hadrons, combining information about their internal quark and gluon longitudinal momentum distributions and transverse position within the hadron. The dependence of GPDs on the factorization scale Q 2 allows one to connect hard exclusive processes involving GPDs at disparate energy and momentum scales, which is needed in global analyses of experimental data. Here, in this work, we explore how finite element methods can be used to construct fast and differentiable Q 2 evolution codes for GPDs in momentum space, which can be used in a machine learning framework. We show numerical benchmarks of the methods' accuracy, including a comparison to an existing evolution code from PARTONS/APFEL++, and provide a repository where the code can be accessed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Detector alignment for X-ray crystallography using Millepede-II

I describe a method for accurately refining the geometrical parameters of segmented X-ray area detectors on the basis of serial crystallography data, using 'Millepede' – an algorithm created for a very similar problem in high-energy physics. The Millepede method for serial crystallography builds on the approach of Brewster et al. [Acta Cryst. (2018), D74, 877–894], in which the detector parameters are refined simultaneously with the parameters for each individual crystal. This accounts for the mutual dependency between the parameters and thereby avoids the bias and slow convergence problems that have afflicted older approaches in which the deviations between observed and calculated Bragg peak positions were taken directly as the updates for the detector panel positions. The Millepede method uses the special structure of the least-squares normal equations to reduce them to a much smaller form that can be solved very quickly, even compared with the sparse matrix methods used previously. This makes it practical to refine the detector geometry frequently and thereby maintain accurate calibration without specialized alignment campaigns. Tilts of detector panels out of the plane can be reliably refined, as can the overall distance of the detector in the beam direction. With a simulated test case, the new method produced panel shifts within 7% of the correct values with only one iteration, and produced almost exactly correct shifts after a second iteration. A simulated out-of-plane panel rotation was correctly determined to within 0.001°. Applied to experimental data from an X-ray free-electron laser, the method increased the indexable fraction of frames from 30% to 91% in a single iteration, and to 96% after two further iterations. Computing the geometry updates on the basis of 2060 crystals took only 0.819 s on desktop computing hardware, including the time taken to read the required data from disk. The scaling was found to be very close to linear for up to 100 980 sets of crystal parameters, which took only 78.2 s to process under the same conditions. The method has been applied as part of a real-time feedback system at a synchrotron radiation beamline, in which an out-of-plane detector tilt of 0.04° was detected and corrected. Possible further applications are also described here.

Millepede-II↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

A Framework for Compressing Unstructured Scientific Data via Serialization

We present a general framework for compressing unstructured scientific data with known local connectivity. A common application is simulation data defined on arbitrary finite element meshes. The framework employs a greedy topology preserving reordering of original nodes which allows for seamless integration into existing data processing pipelines. This reordering process depends solely on mesh connectivity and can be performed offline for optimal efficiency. However, the algorithm’s greedy nature also supports on-the-fly implementation. The proposed method is compatible with any compression algorithm that leverages spatial correlations within the data. The effectiveness of this approach is demonstrated on a large-scale real dataset using several compression methods, including MGARD, SZ, and ZFP.

Reshniak, Viktor [ORNL] (ORCID:0000000315454462)↗

Field testing and validation of a low-cost MPC for demand flexibility for grid-interactive K-12 schools

K-12 school buildings account for the highest energy consumption within the public sector. Implementing advanced HVAC controls in grid-interactive K-12 schools could bring substantial economic advantages and grid flexibility. Our previous study demonstrated that a low-cost model predictive control (MPC) solution, which coordinates multiple packaged units, can enable demand flexibility without major hardware upgrades. However, a significant gap remains between academic pilots and market-ready scalable solutions. This paper extends the previous single-site pilot to a multi-site demonstration involving three school campuses (95 total units) through a commercial technology transfer process. Addressing the challenge of verifying performance with sparse field data, we present a new statistical approach using Bayesian methods to estimate the MPC’s effect on peak demand. Unlike traditional methods, this approach robustly quantifies uncertainty in non-normal, limited datasets. The results confirm the solution’s replicability, achieving a 21.6–38.9% reduction in HVAC peak demand (10.8–22.1% at the site-level) with > 98% probability across diverse locations. Finally, we document critical barriers to scaling software-as-a-service (SaaS) solutions–such as API instability and diverse legacy systems–and offer practical strategies to accelerate the commercial adoption of grid-interactive efficient buildings.

Ham, Sang Woo↗

Progress in end-to-end optimization of fundamental physics experimental apparata with differentiable programming

In this article we examine recent developments in the research area concerning the creation of end-to-end models for the complete optimization of measuring instruments. The models we consider rely on differentiable programming methods and on the specification of a software pipeline including all factors impacting performance — from the data-generating processes to their reconstruction and the inference on the parameters of interest — along with the careful specification of a utility function well aligned with the end goals of the experiment. Building on previous studies originated within the MODE Collaboration, we focus specifically on applications involving instruments for particle physics experimentation, as well as industrial and medical applications that share the detection of radiation as their data-generating mechanism. This report illustrates the most recent advancements in the area, and outlines, for each of the discussed applications as well as for automatic differentiation itself, ongoing and future work.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Conditional Latent Diffusion for High-Resolution Prediction of Electrochemical Surface Morphology

A conditionally guided generative latent diffusion process that is trained on a set of experimental processing parameters and their associated resulting electron microscope images of the electrodeposition process is able to interpolate between processing parameters in a physically consistent way. Electrodeposition of rhenium with pulse and pulse-reverse waveforms is used as a model system, and the process is adaptable to other electrodeposition, electropolishing, or corrosion processes. The method is able to extrapolate, predicting estimates of material morphologies for experimental setups unseen in the training data. The results are demonstrated with experimental data.

36 MATERIALS SCIENCE↗

Evaluation of Seismic Artificial Intelligence with Uncertainty

Artificial intelligence has transformed the seismic community with deep learning models (DLMs) that are trained to complete specific tasks within workflows. However, there is still a lack of robust evaluation frameworks for evaluating and comparing DLMs. Here, we address this gap by designing an evaluation framework that jointly incorporates two crucial aspects: performance uncertainty and learning efficiency. To target these aspects, we meticulously construct the training, validation, and test splits using a clustering method tailored to seismic data and enact an expansive training design to segregate performance uncertainty arising from stochastic training processes and random data sampling. The framework’s ability to guard against misleading declarations of model superiority is demonstrated through the evaluation of PhaseNet (Zhu and Beroza, 2018), a popular seismic phase picking DLM, under three training approaches. Our framework helps practitioners choose the best model for their problem and set performance expectations by explicitly analyzing model performance with uncertainty at varying budgets of training data.

58 GEOSCIENCES↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

Danovo Energy Solution's presented its paper named: Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events at the 2026 Georgia Tech Fault & Disturbance Analysis Conference. The full paper can be found at OSTI ID# 3169150 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danova Energy Solutions]↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗