Search NASASearch

SEARCH · Search NASA

Results for “data cleaning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory

Magnetic edge fields in UTe 2 near zero background fields

Chiral superconductors are theorized to exhibit spontaneous edge currents. Here, in this study, we found magnetic fields at the edges of UTe 2 , a candidate odd-parity chiral superconductor, that seem to agree with predictions for a chiral order parameter. However, we did not detect the chiral domains that would be expected, and recent polar Kerr and muon spin relaxation data in nominally clean samples argue against chiral superconductivity. Our results show that hidden sources of magnetism must be carefully ruled out when using spontaneous edge currents to identify chiral superconductivity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Queued Up: 2025 Edition – Characteristics of Power Plants Seeking Transmission Interconnection As of the End of 2024 [Slides]

Electric transmission system operators (ISOs, RTOs, or utilities) require proposed power plants seeking to connect to the transmission grid to undergo a series of impact studies before they can be built. This process establishes what new transmission equipment or upgrades may be needed before a project can connect to the system and assigns the costs of that equipment. The lists of projects in this process are known as “interconnection queues”. In collaboration with interconnection.fyi, Berkeley Lab compiled, aggregated, and cleaned interconnection queue data from >50 transmission grid operators (7 ISO/RTOs and 49 non-ISO balancing areas), which collectively represent ~97% of currently installed U.S. electric generating capacity. The dataset includes requests submitted to queues through the end of 2024, and only includes requests seeking to connect to the transmission grid (not distribution-connected or behind-the-meter projects). The files below include both a PDF report and an Excel data file. The PDF report analyzes interconnection data and metrics through the end of 2024. The Excel data file includes (a) the full project-level interconnection queue dataset through 2024, (b) a codebook (data dictionary) describing each data field, and (c) 35 additional tabs featuring tables summarizing a range of interconnection metrics. Key highlights from the Queued Up: 2025 Edition (featuring data through 2024) include: • As of the end of 2024, there were ~10,300 projects actively seeking grid interconnection in the U.S., representing 1,400 GW of generation and approximately 890 GW of storage. • Historic withdrawal rates alongside relatively fewer new requests resulted in a 12% decrease in total active queue volume compared to the prior year. • Active natural gas capacity (136 GW, +72% year-over-year) increased in 2024, while solar (956 GW, -12%), storage (890 GW, -13%), and wind (271 GW, -26%) capacity decreased. • 408 GW of capacity already has a draft or executed interconnection agreement (IA) but has not yet reached commercial operations. • The time projects spend in queues before reaching COD is increasing. For the regions with available data, the median duration from IR to COD has doubled from <2 years for projects built in 2000-2007 to over 4 years for those built in 2018-2024. • Ultimately, most of this proposed capacity will not be built. Only 13% of capacity that submitted interconnection requests from 2000-2019 had reached commercial operations by the end of 2024; 77% of that capacity had been withdrawn and 10% was still active. • FERC Order 2023 and various other reforms are being implemented. These are important measures to reduce interconnection bottlenecks and enhance grid system reliability, but it is too early to measure and assess their full impact. • New additions for the 2025 edition include: (a) additional detail on data processing and gaps; (b) updates on interconnection reforms; (c) new analysis on interconnection agreements, and more.

24 POWER TRANSMISSION AND DISTRIBUTION

Generator Interconnection Costs to the Transmission System in non-ISO Balancing Authorities [Slides]

Electric transmission system operators—including Independent System Operators (ISOs), Regional Transmission Organizations (RTOs), and utilities—require proposed power plants to undergo a series of interconnection studies before connecting to the grid. These studies assess what transmission upgrades or new infrastructure may be necessary and assign the associated costs to the project. Lawrence Berkeley National Laboratory has compiled, aggregated, and cleaned interconnection cost data, originally for ISOs/RTOs, and now for five non-ISO Balancing Authorities: PacifiCorp, Bonneville Power Authority, Duke Energy Progress, Duke Energy Carolinas and Duke Energy Florida. Insufficient transparency in interconnection cost data may contribute to rapidly expanding interconnection queues, with active queue capacities tripling between 2020 and 2024 in the studied BAs. Most projects withdraw after receiving high interconnection cost estimates. Interconnection costs have increased since the early 2000s, with average costs for "complete" projects reaching $194/kW between 2018 and 2024. Active queue projects and withdrawn projects incur substantially higher costs, primarily due to rising network upgrade costs. Recent interconnection costs in non-ISO balancing authorities are higher than in ISO regions, potentially due to a greater willingness to pay among developers. Utility-scale solar, wind, and storage projects have interconnection costs that exceed those for natural gas. However, when focusing on projects that do not withdraw from the queue, the interconnection costs for these technologies are more similar to natural gas projects. Other key findings include: (1) Larger generation projects benefit from lower proportional interconnection costs, (2) capacity transmission service (NRIS) often requires additional network investments, and (3) projects with high network upgrade costs are often clustered geographically. The dataset includes results from 2,104 interconnection studies conducted between 2000 and 2024, covering projects that are operational, withdrawn, or still progressing through the study process. The Excel file contains (a) the complete project-level interconnection cost dataset, and (b) seven additional tabs summarizing cost metrics across dimensions such as time, market structure, cost category (point of interconnection vs. broader network upgrades), fuel type, service type (ERIS vs. NRIS), generator size, and geography.

24 POWER TRANSMISSION AND DISTRIBUTION

Queued Up: 2026 Edition, Characteristics of Power Plants Seeking Transmission Interconnection As of the End of 2025 [Slides]

Electric transmission system operators (ISOs, RTOs, or utilities) require proposed power plants seeking to connect to the transmission grid to undergo a series of impact studies before they can be built. This process establishes what new transmission equipment or upgrades may be needed before a project can connect to the system and assigns the costs of that equipment. The lists of projects in this process are known as “interconnection queues”. In collaboration with https://www.interconnection.fyi. Berkeley Lab compiled, aggregated, and cleaned interconnection queue data from >50 transmission grid operators (7 ISO/RTOs and 50 non-ISO balancing areas), which collectively represent ~98% of currently installed U.S. electric generating capacity.

24 POWER TRANSMISSION AND DISTRIBUTION

Data from: Lowland Tropical Forests Remain a Methane Sink Under Warming and Long-Term Hurricane Disturbance Recovery

The repository folder contains spreadsheets and script for soil greenhouse gas (GHG) fluxes, soil moisture, soil temperature, air temperature, and precipitation measurements collected from the Tropical Responses to Altered Climate Experiment (TRACE) at the Sabana Research Field Station, El Yunque National Forest (USDA Forest Service; 18°19′28.74″ N, 65°43′50.09″ W) — an open-air field warming experiment located in a lowland tropical forest in Puerto Rico within the Luquillo Experimental Forest (LEF) — six to seven years after Hurricanes Irma and Maria (2017). All spreadsheets for soil and air microclimate data, as well as soil greenhouse gas data, are included as csv files. Air temperature data are also included as Excel spreadsheets (.xlsx). The script is built in R Studio, which is the only software required to run data analysis. This dataset is associated with the manuscript “Larocca Conte G ; Zuvela L ; Cruz-Pérez R ; Barreto-Vélez T ; Becerra-Santillan N ; Campbell S ; Chu H ; Dam T ; Grullón-Penkova I ; Kleit M ; Ortiz-Iglesias D ; Rubio-Lebrón L ; Cavaleri M ; Reed S ; Sihi D ; Wood T ; O'Connell C., 2026. Lowland Tropical Forests Remain a Methane Sink Under Warming and Long-Term Hurricane Disturbance Recovery. Agricultural and Forest Meteorology. In review". The dataset was used to test the effect of warming on soil CH4 dynamics following long-term legacy effects of hurricane disturbance. The dataset includes: - An overall README file in word and pdf format describing methodology and spreadsheets’ structure. - Continuous measurements of soil temperature and moisture from January 2023 to July 2024 measured with Campbell CS655 probes (“TRACE_soil_temperature_and_moisture_2023_cleaned(in).csv” and “TRACE_soil_temperature_and_moisture_2024_cleaned. csv”). - Air temperature data measured with a HOBO MX23O1A data logger (“Hobo air temperature 2023 Sep 2024” and “Hobo air temperature 2023 Sep 2024” – “CSV FILES folders”). - Precipitation data from a nearby weather tower downloaded from González et al. (2025; “sabana_2020-2025.csv”). - Soil CH4 and CO2 effluxes measured intermittently in two summer campaigns (June – August 2023 and June – July 2024) with a LI-COR 8200-01S Portable Smart Chamber coupled with a LI-COR LI-7810 CH4/ CO2/H2O Trace Gas Analyzer (“23_24COMBO2.0.csv”). - R markdown script for data analysis (“Trace new_PLOTS.Rmd”).

54 ENVIRONMENTAL SCIENCES

Characterization of Electronic Stress-Induced Changes in Multilayer MoS 2

Transition metal dichalcogenides like molybdenum disulfide (MoS 2 ) are compelling for next-generation electronic devices. In this work, we investigate the impact of electronic stress on MoS 2 to illustrate that observational and phenomenological information on multiple devices can be useful to describe changes in the device, and caution against the rationalization of paltry results as representative or correlative to device behavior. Here, we stress MoS 2 by applying a sustained 20 V DC bias to study the material’s response. Post-stress electronic characterization revealed nonuniform shifts in current–voltage (I–V) behavior alongside microscale changes. Complementary mechanical, spectroscopic, and scanning microwave impedance measurements showed that stress-induced features locally modulate stiffness, surface potential, Raman intensity, and charge carrier density. We correlated I–V behavior with morphological features (wrinkles, tears, folds, height) and device-level geometry (MoS 2 overlap with electrodes, channel area, contact length) on 50 test structures across five chips to move beyond anecdotal conclusions. We found no universal correlations before DC stress. However, device-level geometry was correlated with I–V behavior after DC stress, suggesting that electrode contacts play a more dominant role than morphology in determining performance. Delamination and thinning induced by DC stress led to localized reductions in charge carrier density within the affected regions. Further, delamination and thinning appear to map to I–V device performance in a few samples, but the correlation is lost when a larger sample size is considered. This suggests significant sample-to-sample variability in surface electronic states of the test structures. We also discuss how environmental factors introduced during fabrication may contribute to the observed heterogeneous device response. Progress will require high-resolution, multimodal analysis across many samples constructed under controlled, clean conditions. By building data sets that capture variability, we can better identify the true drivers of performance.

36 MATERIALS SCIENCE

FervoFlex: Long duration in-reservoir energy storage and load-following, dispatchable geothermal generation

Fervo recently achieved a groundbreaking milestone with the successful completion of its pioneering commercial geothermal project in Northern Nevada. This project is partially supported by the ARPA-E OPEN grant, FervoFlex™ technology—a long-duration in-reservoir energy storage and load-following, dispatchable geothermal generation system. Notably, this project marks a significant stride in commercial viability, delivering uninterrupted, 24/7 clean energy to Google data centers in Nevada (Terrell, Michael 2023)

15 GEOTHERMAL ENERGY

Flexible Operation of Natural Gas Power Plants in Texas: Startup and Shutdown Durations and Nitrogen Dioxide Emissions

This dataset provides insights into historical flexible operation of natural gas power plants in Texas, with a focus on startup and shutdown events. The dataset includes tables summarizing startup and shutdown durations as well as nitrogen oxide (NOx) emission factors during these events, and compares these emission factors with those observed during all other operating phases (referred to here as “steady-state operation”). The dataset is derived using the U.S. Environmental Protection Agency (EPA)’s Clean Air Markets Program Data (CAMPD). Historical hourly data from 2015–2024, including electricity generation, heat input, and NOx emission factors, are used for natural gas combined cycle units, combustion turbine units, and steam turbine units in Texas. This work was authored by the National Laboratory of the Rockies, operated by the Alliance for Energy Innovation, LLC, for the U.S Department of Energy (DOE) under contract no. DE-AC36-08GO28308. Funding was provided by the U.S. Department of Energy as part of its Grid Modernization Laboratory Consortium, a strategic partnership between DOE and the national laboratories to bring together leading experts, technologies, and resources to collaborate on the goal of modernizing the nation’s grid. The views expressed in the dataset do not necessarily represent the views of the DOE or the U.S. Government.

24 POWER TRANSMISSION AND DISTRIBUTION

Development and Preliminary Analysis of a U.S. Geothermal Heat Pump Installation Database

This paper seeks to addresses the significant gap in the literature regarding the installation and adoption of geothermal heat pump (GHP) systems in the United States. While the "2021 U.S. Geothermal Power Production and District Heating Market Report" published by the National Renewable Energy Laboratory (NREL) focused on direct-use geothermal district heating systems, it did not include an analysis of GHP installations (Robins et al. 2021). To bridge this gap, NREL has compiled a novel database currently containing 70,470 records of GHP installations, primarily sourced from state well permits and small-scale studies. Our methodology emphasizes the collection, cleaning, and standardization of data, addressing challenges such as inconsistent reporting formats and privacy concerns. Despite limitations in data on capacity, costs, and performance, our preliminary geospatial analysis reveals insights into the distribution of GHP systems across urban and rural areas and climate zones. The paper highlights the importance of publicly accessible data for advancing GHP technology adoption with a discussion of existing data sources and their limitations, advocating for improved collaboration between NREL and industry stakeholders.

data collection

Dwarf galaxy halo masses from spectroscopic and photometric lensing in DESI and DES

We present the most precise and lowest-mass weak lensing measurements of dwarf galaxies to date, enabled by spectroscopic lenses from the Dark Energy Spectroscopic Instrument (DESI) and photometric lenses from the Dark Energy Survey (DES) calibrated with DESI redshifts. Using DESI spectroscopy from the first data release, we construct clean samples of galaxies with median stellar masses $\log_{10}(M_*/M_{\odot})=8.3-10.1$ and measure their weak lensing signals with sources from DES, KiDS, and SDSS, achieving detections with $S/N$ up to 14 for dwarf galaxies ($\log_{10}(M_*/M_{\odot})<$9.25) -- opening up a new regime for lensing measurements of low-mass systems. Leveraging DES photometry calibrated with DESI, we extend to a photometric dwarf sample of over 700,000 galaxies, enabling robust lensing detections of dwarf galaxies with combined $S/N=38$ and a significant measurement down to $\log_{10}(M_*/M_{\odot})=8.0$. We show that the one-halo regime (scales $\lesssim 0.15h^{-1}\rm Mpc$) is insensitive to various systematic and sample selection effects, providing robust halo mass estimates, while the signal in the two-halo regime depends on galaxy color and environment. These results demonstrate that DESI already enables precise dwarf lensing measurements, and that calibrated photometric samples extend this capability. Together, they pave the way for novel constraints on dwarf galaxy formation and dark matter physics with upcoming surveys like the Vera C. Rubin Observatory's LSST.

Treiber, Helena [Princeton U., Astrophys. Sci. Dep

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato

Data Qualification Report: SRNL Glass Composition-Properties (ComPro) Database

The Savannah River National Laboratory Glass Composition-Properties (ComPro) database is an extensive database containing pertinent composition and durability data to support the accelerated clean-up mission at the Defense Waste Processing Facility. The activities described in this data qualification report were performed to support the information contained in the database. There were two objectives of the original data qualification process. The first objective was to review supporting documentation to determine if DOE/RW-0333P Quality Assurance Requirements and Description had been implemented during the original work. If the DOE/RW-0333P Quality Assurance Requirements and Description had not been directly implemented during the original work, the second objective was to determine if the controls that were used were adequate to meet the intent of the DOE/RW-0333P Quality Assurance Requirements and Description. The results of these two objectives and the activities performed to support these decisions are described in this document. An assessment of each dataset was made to determine if the data were RW-0333P Compliant, RW-0333P Equivalent or Non-RW-0333P Compliant. The original data qualification was performed in accordance with E7, Conduct of Engineering Manual, Procedure 3.70, Revision 4, Qualification of Data. The specific method that was used was Equivalent Controls as described in E7, 3.70. Revision 2 of this document adds supporting information for the RW-0333P Compliant datasets added to Revision 3 of the database.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING

EPCAPE-Dalhousie Field Campaign Report

This project supported the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE), which aimed to characterize the extent, radiative properties, aerosol interactions, and precipitation of stratocumulus clouds in the Eastern Pacific coastal region of La Jolla, California. Given the frequent anthropogenic aerosols emitted into this region, characterization of this important coastal cloud region and the aerosol impact on the cloud characteristics will improve the understanding of aerosol indirect effects and its representation in global climate models. This specific project deployed the fog droplet monitor (FM-120, Droplet Measurement Technologies) at the Scripps Institution of Oceanography’s Mt. Soledad site from 16 February 2023 to 15 February 2024. This instrument measured droplet number size spectra between 2 and 50 microns, from which droplet number concentration, liquid water content, and effective diameter were calculated. These measurements will be used to characterize the seasonal variability of clouds at Mt. Soledad and investigate differences in cloud properties under regional polluted and clean marine conditions. The collected data will soon be posted as part of the digital collection for EPCAPE data hosted by the University of California, San Diego (Russell et al. 2023).

54 ENVIRONMENTAL SCIENCES

Heating effects on jack pine pyrogenic organic matter properties from a pyrocosm study in 2022

This dataset contains data associated with the preprint “Fire removes preexisting pyrogenic organic matter from the ecosystem through the mechanisms of both direct combustion and increasing mineralizability” (Luo et al., 2025b), which is the complementary study to the published paper “Reburning pyrogenic organic matter: a laboratory method for dosing dynamic heat fluxes from above” (Luo et al., 2025a). We designed a full-factorial experiment with different burial depths of jack pine (Pinus banksiana Lamb) pyrogenic organic matter (PyOM) (Surface, 1 cm, and 5 cm) and different heat-flux profiles (High, Low, and Control) to examine how subsequent fires affect the properties of preexisting PyOM. We measured total carbon (C), pH, dissolved organic carbon (DOC), dissolved inorganic carbon (DIC), and mineralized C (as CO₂-C, from a 12-week incubation).We found that high heat flux and/or surface placement resulted in substantial direct C losses through combustion. Intermediate heat exposure produced both combustion losses and increases in DOC and mineralizability, which may have complex long-term implications: an increased dissolved fraction of PyOM may promote downward transport into mineral soils and potentially contribute to deeper, longer-term C storage, but it may also make PyOM more susceptible to microbial decomposition. Under the lowest heat flux and deepest burial, most PyOM was retained, and changes in DOC and C mineralization were minimal. Finally, PyOM pH, an important chemical property, decreased under low-temperature heating but increased under higher temperatures.We uploaded pH data for all samples (“pH_of_all_samples.csv”); pH and temperature-related data (peak temperature and degree hours) for samples in High and Low heat-flux treatments (“pH_vs_peakT_and_degree_hours_only_for_heated_samples.csv”); total C data (“CN_pct_C_stock_C_loss_in_samples.csv”); DOC and DIC data (“doc_dic.csv”); and mineralized C (CO₂-C) data (“CO2-C_all_original.csv”). Additional details can be found in the Methods & Sampling section.All datasets uploaded to ESS-DIVE are clearly labeled, cleaned, and include both raw and derived data, ready for reuse in other analyses. All analysis code and raw datasets are also available on GitHub: https://github.com/MengmengLuo/Fire-removes-preexisting-pyrogenic-organic-matter-from-the-ecosystem.

54 ENVIRONMENTAL SCIENCES

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)