Search NASA⌕ Search

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

Simulator Data Analysis to Inform Digitalized Environment Impacts on Human Reliability

The U.S. Nuclear Regulatory Commission (NRC) has developed a human reliability analysis (HRA) method, termed the Integrated Human Event Analysis System for Event and Condition Assessment (IDHEAS-ECA), in order to estimate human error probabilities (HEPs) in risk-informed regulatory applications. To update the quantification part of IDHEAS-ECA, the NRC required human performance and error data from fully digitalized main control rooms (MCRs); therefore, it requested that Idaho National Laboratory (INL) revisit previous data collection studies and investigate how the following three factors impact human reliability: self-checking, peer-checking, and automation. The HRA data collection studies revisited were the Human Reliability Data Extraction (HuREX) project, developed by the Korea Atomic Energy Research Institute (KAERI), and the Simplified Human Error Experimental Program (SHEEP), developed by INL. HuREX is a representative HRA data collection study that collects human reliability data from full-scope simulators staffed by licensed operators. SHEEP, on the other hand, has been proposed to complement such full-scope studies by collecting data via simplified simulators staffed by non-licensed student operators. In the HuREX study, KAERI collected HRA data from fully digitalized MCRs for the Advanced Power Reactor (APR)–1400. The SHEEP data were obtained from simplified simulators that partially mimicked the features of digitalized MCRs. The present report mainly discusses how the impacts of the aforementioned three factors on human errors were derived from these two data collection studies.

99 GENERAL AND MISCELLANEOUS↗

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige↗

Meteorological Services Annual Data Report for 2024

This document presents the meteorological data collected at Brookhaven National Laboratory (BNL) by Meteorological Services (Met Services) for the calendar year 2024. The purpose is to publicize the data sets available to emergency personnel, researchers and facility operations. Met services has been collecting data at BNL since 1949. Data from 1994 to the present is available in digital format. Data is presented in monthly plots of one-minute data. This allows the reader the ability to peruse the data for trends or anomalies that may be of interest to them. Full data sets are available to BNL personnel and to a limited degree outside researchers. The full data sets allow plotting the data on expanded time scales to obtain greater details (e.g., daily solar variability, inversions, etc.).

54 ENVIRONMENTAL SCIENCES↗

Automated qualification data tool for high temperature metallic materials

This report describes a framework for storing, processing, and displaying qualification data for high temperature mechanical properties. The framework automates the process of generating design data from mechanical test results, for example for a data qualification report for the ASME Boiler \& Pressure Vessel Code. The framework has three parts: a data storage model with common formats for several types of typical mechanical property tests, a backend based on the \pycreep Python library for correlating and extrapolating the data to generate design material properties and allowable stresses, and a demonstration user interface for displaying, sorting, and filtering the data and exploring different options for modeling the design mechanical properties. The report discusses the options available for data processing, with illustrations from real test data on Alloy 617, Alloy 709, Alloy 740H, and Laser-Powder Bed Fusion 316H. The framework is complete for ASME type data analysis and will be used to store test data generated by the Department of Energy, Office of Nuclear Energy, Advanced Materials and Manufacturing Technologies sponsored qualification programs. Future work could extend the tool to other types of material properties and/or expand the demo user interface to make it accessible across the AMMT program.

36 MATERIALS SCIENCE↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Self-potential tomography preconditioned by particle swarm optimization—Self-potential monitoring and streamflow data acquired March 26–September 14, 2023 at East Fork Poplar Creek near Oak Ridge Tennessee

This data release contains self-potential (SP) monitoring data measured on the flood plain of East Fork Poplar Creek (East Fork) in Oak Ridge, Tennessee and streamflow data measured at streamgage EFK5.4 about 310 meters upstream from the SP monitoring site. Additionally, forward and inverse numerical modeling scripts used to model the electrical-potential field on the East Fork flood plain are provided. SP monitoring data included in this data release were measured at 39 different data-collection points on the east flood plain; 30 points were spaced 3-m apart along an 87-m profile parallel to the edge of the streambank, and 9 points were spaced 5-m apart along a 40-m profile approximately perpendicular to the streambank. The two profiles of SP data-collection points intersected at the approximate midpoint of the profile parallel to the streambank. Transient voltages were measured at each data-collection point every 60 seconds between 16:13 Eastern Standard Time (EST) on March 26, 2023, and 11:41 EST on September 14, 2023. Streamflow data included in this data release overlap the time-period of self-potential monitoring and were measured every 900 seconds between 16:23 on March 26, 2023, and 23:53 on September 14, 2023.

54 ENVIRONMENTAL SCIENCES↗

Data for: Climatic Imprint on Interfacially-Controlled Platinum-Palladium Resources

Data package for manuscript "Climatic Imprint on Interfacially-Controlled Platinum-Palladium Resources" by Emily G. Wright, Ivey Wang, Yihang Fang, Elaine D. Flynn, and Jeffrey G. Catalano. This dataset contains adsorption results from experiments designed to investigate the effect of chloride on Pd(II) adsorption to goethite and Pt(II) adsorption to hematite and goethite, including lab experiments, X-ray absorption fine structure spectroscopy, and models of retention within a laterite. See the associated manuscript for full methods information. The file "Wright2025_PtAds_data.csv" contains the target starting Pt concentration (uM), final aqueous Pt and associated error (in uM), calculated adsorbed Pt and associated error (in umol/m2), target and measured aqueous chloride (mM), target aqueous nitrate (mM), final pH, and mineral concentration/loading (g/L). Associated mineral-free controls (mineral loading = 0 g/L) are included; the aqueous Pd error was not calculated and chloride was not measured in every sample. These data appear in Figures 1, S3, S4, S5, S20, and S22 in the associated manuscript. The file "Wright2025_PdAds_data.csv" contains the target starting Pd concentration (uM), final aqueous Pd and associated error (in uM), calculated adsorbed Pd and associated error (in umol/m2), target and measured aqueous chloride (mM), and mineral concentration/loading (g/L). Associated mineral-free controls (mineral loading = 0 g/L) are included; the aqueous Pd error was not calculated and chloride was not measured in every sample. These data appear in Figures 1, S3, S4, S5, and S20 in the associated manuscript. The file "Wright2025_MineralBatches_data.csv" contains the mineral identity and BET specific surface area (m2/g) for every mineral batch synthesized and used in experiments. The annealing time used is listed for hydrothermally annealed goethite. These data appear in Table S2 in the associated manuscript. The file "Wright2025_XRD_data.csv" contains the XRD patterns for every mineral batch synthesized as the counts as a function of two theta (in degrees). See "Wright2025_MineralBatches_data.csv" for more details on specific mineral batches. These data appear in Figure S2 in the associated manuscript. The file "Wright 2025_ZetaPotential_data.csv" contains the measured zeta potentials for samples of goethite (batch G2) at pH 4 the presence of varying amounts of sodium chloride. These data appear in Table S3 in the associated manuscript. The file "Wright2025_XAFSSamples_data.csv" contains the specific mineral batch, measured final aqueous Pd or Pt (uM), measured final aqueous chloride (mM), and estimated adsorbed Pd or Pt (umol/m2) of all XAFS samples. These data appear in Tables S4, S7, S8, and S10 in the associated manuscript. The files "Wright2025_PdXAFS_data.csv" and "Wright2025_PtXAFS_data.csv" contain the normalized spectra of Pd and Pt, respectively, adsorbed to minerals at varying chloride concentrations. See "Wright2025_XAFSSamples_data.csv" for a guide to sample names. Note that "05" in a sample name is equivalent to "0.5". These data appear in Figures 2, S6, S7, S8, S12, S13, and S14 in the associated manuscript. The file "Wright2025_LateriteProfileProfileModelParameters_data.csv" include the ratio of hematite to hematite and goethite in two synthetic, modeled profiles, as well as the modeled surface areas of goethite and hematite as a function of relative depth within the modeled weathering zone. These data were used, in conjunction with equations presented in the paper, to calculate the theoretical concentrations of Pd and Pt (and the resulting Pt/Pd ratio) within the profiles. These data appear in Figure 3 in the associated manuscript. The file "Wright2025_Imagery_data.zip" is a zipped folder containing the TEM and STEM images appear in Figures S18 and S19. Individual files are labeled as either STEM (Fig. S18) or TEM (Fig. S19) with a letter representing the part of the multipart figure.

58 GEOSCIENCES↗

Replication Data for: Measurement of the mean number of muons with energies above 500 GeV in air showers detected with the IceCube Neutrino Observatory

<b>Measurement of the mean number of muons with energies above 500 GeV in air showers detected with the IceCube Neutrino Observatory</b> <br><br> This data release accompanies results submitted to Physical Review D describing the measurement of the average multiplicity of TeV muons with IceCube. It contains the data necessary to reproduce the main plots from the paper (Figs. 7 and 9), i.e. the numerical results for the average number of muons with energies above 500 GeV as a function of primary cosmic ray energy. <br><br> For any questions about this data release, please write to analysis@icecube.wisc.edu. <br><br> Files included in this release: <ul> <li>A README file <li>Files including data to reproduce the results plots from the paper (see below for details) <li>An example python script showing how to read and plot the data </ul> <br> <u>What is in the files icecube_Nmu500_X_Y.txt:</u> <br> Y indicates wether the file contains values obtained from experimental data (Y="data") or air-shower simulations (Y="MC"). <br> X indicates the hadronic interaction model for which the plot is made. If Y="data", this means that the experimental data was interpreted using this model. If Y="MC", it means that the simulations were performed with this model. The three models included are Sibyll 2.1, QGSJet-II.04, and EPOS-LHC (see paper for references). The file with X="modelaverage" gives the average over the three individual results with the deviations from the average included in the systematic uncertainties. <br><br> Please see the README file for details on how the data is structured in the files.

Astroparticle Physics↗

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗

Data Selection Improvement For MicroBooNE

Data selection is an extremely important part of data analysis for any experiment. Finding a physics result is often the result of sifting through a massive amount of data, keeping data that we believe to be signal, and throwing out data we do not. This process is called data selection. Creating a selection algorithm is an intensive process that must balance keeping enough data to have statistics and maximizing the signal purity of that data. We also need to choose the right reconstruction method, a tool to take raw data from the detector and convert it into physics results. In this study, we used three different reconstruction tools, Pandora, WireCell, and LANTERN, for the MicroBooNE experiment in conjunction to improve the selection algorithm for analysis. For the case of this study, we look into the charged current N proton 0 pions (CCNp0$\pi$) interaction channel. This is the dominant channel for the Short Baseline Neutrino (SBN) program and is expected to be a large contributor to the Deep Underground Neutrino Experiment (DUNE). We first investigated each of the three tools to find out more about their strengths and weaknesses as reconstructions, and compared them to the truth information directly from the MicroBooNE simulation pipeline. We then put together a direct comparison of the three methods to find which method or combination of methods would return the best result for us. While the study is ongoing, we have learned a lot about data selection for the experiment and the differences between the reconstruction tools.

Dillon, Brayden [Fermilab]↗

Open Energy Data Initiative (OEDI) FY22-24 (Final Technical Report)

Final technical report for the Open Energy Data Initiative (OEDI) project covering fiscal years FY22 through FY24. The DOE Open Energy Data Initiative (OEDI) is a partnership between the National Renewable Energy Laboratory (NREL), the U.S. Department of Energy (DOE), and major cloud providers including Amazon, Microsoft, and Google to provide universal access to big data in the cloud. At the heart of OEDI is a centralized repository of high-value energy research datasets aggregated from the U.S. Department of Energy's Program Offices, National Laboratories and other collaborators. It aggregates smaller, domain-specific repositories, allows direct data submissions, and includes support for big data through its energy data lakes. OEDI's data lakes make high-value data universally accessible and help researchers, collaborators and the general public overcome many of the obstacles to accessing and using big data.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING↗

Development of physics-consistent conditional diffusion model to overcome data scarcity in critical heat flux

Deep generative modeling provides a powerful pathway to overcome data scarcity in energy-related applications where experimental data are often limited. By learning the underlying probability distribution of the training dataset, deep generative models, such as the diffusion model, can generate high-fidelity synthetic samples that statistically resemble the training data. Such synthetic data generation can significantly enrich the size and diversity of the available training data, and more importantly, improve the robustness of downstream machine learning models in predictive tasks. The objective of this paper is to investigate the effectiveness of diffusion models for overcoming data scarcity in nuclear energy applications. By leveraging a public dataset on critical heat flux which covers a wide range of commercial nuclear reactor operational conditions, we developed a diffusion model that can generate an arbitrary amount of synthetic samples. Since a vanilla diffusion model can only generate samples randomly, we also developed a conditional diffusion model capable of generating targeted critical heat flux data under user-specified thermal-hydraulic conditions. The performance of the diffusion model was evaluated based on its ability to capture empirical feature distributions and pair-wise correlations, as well as to maintain physical consistency. The results showed that both the diffusion model and conditional diffusion model can successfully generate realistic and physics-consistent critical heat flux data. Furthermore, uncertainty quantification results demonstrate that the conditional diffusion model is highly effective in augmenting critical heat flux data while maintaining acceptable levels of uncertainty.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Shedding light on U.S. small and midsize data centers: Exploring insights from the CBECS survey

As demand for digital services accelerates, the energy and environmental footprint of data centers faces increasing scrutiny. While hyperscale cloud facilities have driven efficiency gains, small and midsize U.S. data centers remain a critical yet underexamined segment with significant untapped potential for energy savings. This study leverages data from the Commercial Buildings Energy Consumption Survey (CBECS) to analyze trends in server stocks, computing customers, cooling system adoption and efficiency, and geospatial distribution from 2012 to 2018. Findings reveal a sharp decline in small and midsize data centers, from 1.764 million to 1.398 million, with server counts dropping from 5.177 million to 4.262 million—aligning with the broader shift toward cloud computing. More than 40 % of servers in small data centers and 55 % in midsize data centers are housed in office buildings, and over half of all servers are concentrated in climate zones 5A (cold), 3A (mixed-humid), and 4A (mixed-humid), with the highest densities in metropolitan hubs. While direct expansion units remain the dominant cooling system, a clear transition toward more energy-efficient solutions, particularly air economizers, is evident. By integrating server and cooling system distributions, we estimate Power Usage Effectiveness (PUE) and Water Usage Effectiveness (WUE) for U.S. data centers by size and year. Results show that midsize data centers are more energy-efficient but more water-intensive due to the widespread use of water-cooled chillers. These findings highlight the trade-offs in cooling system selection and provide a critical foundation for policies aimed at enhancing efficiency in an evolving data center landscape.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Implementation of high-speed data acquisition at DIII-D

Research at the DIII-D National Fusion Facility in San Diego focuses on short pulse plasma discharges that specialize on various shaping profiles. High-speed data collection is a critical component for the operation of many of DIII-D’s diagnostics and is fundamental for capturing high-resolution data used in experimental data analysis. Differing techniques enable the plasma control system (PCS) to perform complex real-time feedback control on microsecond time scales. This work presents a comprehensive overview of data acquisition, focusing on the hardware and software used in reliable data acquisition at DIII-D. The robust nature of the data acquisition system allows for various techniques to coexist seamlessly. However, as modern systems capable of nanosecond resolution become more common, existing architectures will need to be modified. Here, by addressing the key challenges of high-speed data acquisition, DIII-D is able to provide real-time data used in plasma operation and has the ability to acquire high fidelity data needed for future experimental fusion reactors, such as ITER.

Control↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗