Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗

Developing a Database of Bio-based Materials for Building Envelope Applications

Oak Ridge National Laboratory (ORNL) has been funded by the Department of Energy (DOE) to help accelerate the introduction of building envelope materials that would reduce the carbon footprint of the buildings sector. The DOE’s Building Technologies Office has historically sought to resolve the knowledge gaps regarding the energy efficiency and moisture durability of building envelope systems and to develop the data, guidance, and tools needed to facilitate rapid industry adoption of high-performance, moisture-managed envelope systems. This project will help accelerate the widespread acceptance of a new generation of building materials developed specifically with the intent of reducing the carbon footprint of buildings. We have produced a database of hygrothermal transport properties on low embodied carbon building materials that can be added to energy and durability simulation tools. Properties that were measured include density, heat capacity, thermal conductivity as a function of temperature and relative humidity, moisture dependent permeance, and sorption isotherms as a function of relative humidity. These data sets were measured following consensus national standards using state-of-the-art facilities. The data has been compiled and is being made available to building designers who require these data to assess these new materials in their designs. We will publish the data and seek its addition to reference databases such as the ASHRAE Handbook of Fundamentals.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Monitored Natural Attenuation (MNA) Assessment for the Chemicals, Metals, and Pesticides (CMP) Pits Operable Unit (OU) and the Pen Branch Wetland

In October 2024, Savannah River National Laboratory (SRNL) was tasked to conduct an independent review of groundwater data and assess the monitored natural attenuation (MNA) performance related to the Chemicals, Metals, and Pesticides (CMP) Pits Operable Unit (OU) located in the central portion of the Savannah River Site (SRS). The SRNL assessment of MNA entailed an independent analysis of groundwater concentration data, groundwater elevation data, available surface water concentration data, soil concentration data, and relevant historical geological characterization logs for the CMP Pits OU. Historical data tables and records were obtained from both the Savannah River Nuclear Solutions - Area Completion Projects (SRNSACP) team and South Carolina State University (SCSU) and condensed into new data sets by the SRNL project team for more targeted analysis of MNA performance characteristics. All data pertaining to groundwater concentration, surface water concentration, and groundwater elevation were restricted to collection dates after any known active remediation for the CMP OU.

54 ENVIRONMENTAL SCIENCES↗

NPFTURBULENCE: Best Estimate Aerosol Size Distribution by airborne measurements

The original data were collected during the field campaign of “Turbulent layers promoting New Particle Formation” experiment (NPFTURBULENCE; https://www.arm.gov/research/campaigns/aaf2024npfturbulence) over the Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) Atmospheric Observatory (https://www.arm.gov/capabilities/observatories/sgp ) in north-central Oklahoma. The ARM Aerial Facility ArcticShark uncrewed aerial system (UAS) was based at Blackwell–Tonkawa Municipal Airport (IATA: BWL, ICAO: KBKN, FAA LID: BKN, 36.74475° N, 97.34918° W, 313.9 m MSL), for the field campaign from May 5 through May 29, 2024. The ArcticShark UAS performed 11 flights, including 10 research flights over the Central Facility of the ARM SGP to measure atmospheric state, turbulence, surface IT temperature and imagery, aerosol number concentration and size distribution. The current data set presents Best Estimate Aerosol Size Distribution: a merged aerosol size distribution composed of the data from 2 sensors: miniaturized Scanning Electrical Mobility Sizer (mSEMS) and Portable Optical Particle Spectrometer (POPS). The mSEMS data were interpolated to 1 second from “native” time resolution of about 15 second to match the other probe. The POPS data were converted from equivalent optical size into geometric size using value of aerosol refractive index of 1.477 from the HISCALE field campaign (same geographical area, altitudes, and time of year; http://www.arm.gov/campaigns/aaf2016hiscale ).

54 ENVIRONMENTAL SCIENCES↗

Instrumentation for the Investigation of Pitch Bearing Design and Reliability

Recently, there has been an increasing level of industry interest in pitch system and pitch bearing reliability. Pitch bearings are used in wind turbines to connect the blade root to the hub. Some populations of pitch bearings have demonstrated a 12% failure rate in 20 years. As rotor diameters continue to increase for tall land-based and offshore wind turbines, pitch bearings are becoming even larger in diameter, which can make them vulnerable to deflections and consequent stress concentrations. There is an increased need to more accurately study pitch bearing deformations, misalignment, load distributions, and contact stresses. A significant body of work has investigated fatigue lives and wear characteristics of pitch bearings on ground-based test rigs. NREL has also recently begun a research program related to pitch bearing reliability, recognizing its growing importance for wind turbines. The purpose of this paper is to describe a set of instrumentation that was recently installed on a 1.5 MW wind turbine at the NREL Flatirons Campus and provide an example data set. To the authors' knowledge, this will be the first publicly available pitch bearing data collection campaign on an operational wind turbine.

17 WIND ENERGY↗

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON↗

Cauchy problems for Einstein equations in three-dimensional spacetimes

Abstract We analyze existence and properties of solutions of two-dimensional general relativistic initial data sets with a negative cosmological constant, both on spacelike and characteristic surfaces. A new family of such vacuum spacelike data parameterised by poles at the conformal boundary at infinity is constructed. We review the notions of global Hamiltonian charges, emphasizing the difficulties arising in this dimension, both in a spacelike and characteristic setting. One or two, depending upon the topology, lower bounds for energy in terms of angular momentum, linear momentum, and center of mass are established.

Chruściel, Piotr T. (ORCID:0000000183627340)↗

Dark Energy Survey Year 6 Results: Weak Lensing and Galaxy Clustering Cosmological Analysis Framework

We present the methodology for the weak lensing and galaxy clustering analyses of the Dark Energy Survey (DES) Year 6 data set. In this work, we design and validate the analysis pipeline for the cosmic shear, galaxy clustering plus galaxy$-$galaxy lensing ($2 \times 2$pt), and the joint analysis in the $3 \times 2$pt. Our framework accounts for key theoretical uncertainties, such as baryonic feedback and galaxy bias, incorporating both linear and non-linear models. We apply scale cuts in regimes where theoretical modeling becomes unreliable. The robustness of the pipeline is validated using mock data and simulations, confirming unbiased cosmological constraints and highlighting the importance of posterior projection effects in the validation process. As a result, we deliver robust and validated analysis pipelines for cosmic shear, $2 \times 2$pt, and $3 \times 2$pt in $Λ$CDM and $w$CDM scenarios, including a well-defined set of scales suitable for real data analysis, a robust prescription for theoretical systematics, and the theoretical covariance of the signal. This comprehensive methodology also lays the groundwork for future galaxy surveys such as the Vera C. Rubin Observatory Legacy Survey of Space and Time.

Sanchez-Cid, D. [Zurich U.; Madrid, CIEMAT; Madrid↗

Addendum to: Combined analysis of neutrino decoherence at reactor experiments

We update our analyses to constrain neutrino decoherence induced by wavepacket separation with RENO and Daya Bay data, now including the final data sets of the two experiments. We find that while the individual bounds from Daya Bay and RENO data improve relative to our original estimates, the combined fits are still dominated by KamLAND data and are only minimally improved.

de Gouvêa, André↗

The Role of Snowmelt and Subsurface Heterogeneity in Headwater Hydrology of a Mountainous Catchment in Colorado: A Model‐Data Integration Approach

Mountainous headwater streams are sustained by both snowmelt‐driven streamflow and groundwater discharge in the Upper Colorado River Basin. However, predicting headwater stream discharge magnitude and peak flow timing is challenging in mountainous terrains, where snowmelt rates vary with vegetation type and elevation, and heterogeneous subsurface physical properties influence groundwater storage and its release. We used a model‐data integration approach to investigate the roles of snowmelt and subsurface structure in stream discharge and groundwater level. We ran an ensemble of 100 integrated surface‐subsurface hydrologic models for a mountainous headwater catchment near Crested Butte, Colorado, USA. We also evaluated and calibrated these models against observed data sets, including snow depth measurements using distributed temperature probes, stream discharge, and groundwater levels. Calibration with multiple data sources using neural density estimators has further constrained uncertainty in subsurface properties and snowmelt rates. Results indicated that observed slower snowmelt rates in evergreen forests delayed the peak flow and baseflow onset. In upstream areas with lower subsurface permeability, water was stored within the subsurface but was not released as interflow or shallow groundwater flow, and thereby not contributing to downstream streamflow during recession limb periods. Double peaks in groundwater occurred in areas with spatial subsurface heterogeneity, in our case due to the contrast between granodiorite and Mancos shale. These process‐based insights into groundwater and snowmelt dynamics in mountainous headwaters will help improve predictions of headwater hydrology.

Wang, Lijing [University of Connecticut, Storrs, C↗

Data for "Depth of nutrient uptake by deep-rooted plants is regulated by water availability"

The data set consists of strontium (Sr) isotope ratios (87Sr/86Sr), water isotopes, soil cation concentrations, soil water potential sensor data, and results of 87Sr/86Sr mixing model. The plant canopy size files include the dataset of canopy dimension of sagebrush, lupine, and sunflower. The soil and plant ICPMS (Inductively Coupled Plasma Mass Spectrometry) data file includes both of 87Sr/86Sr, and cation concentration dataset from soil exchangeable pool, apatite pool, silicate extract, atmospheric rain deposition, and plant leaf and stem tissues. The plant dendrochronology file includes the dendrochronogical ring width of several sagebrush, and dendrochemical sample data includes the 87Sr/86Sr for each separated growth ring. The modeling result gives the proportion of nutrient sources of each plants (based on their 87Sr/86Sr in leaf tissues and growth rings) from atmospheric deposition and mineral weathering. Soil water potential data includes continuous collection of soil water potential dataset at 2 depths (30 cm and 60 cm, from Nov 24 - Jun 25) of the sampling site. All the samples were collected from 2 sampling campaign June and July 2023, and rain water is a separate sampling from Aug - Sept 2023, at north-facing hillslope near pumphouse site. The data showed that the depth of cation nutrient acquisition is thus tightly coupled with, and likely determined by, water availability in soil, saprolite and bedrock. The enhanced uptake of cations and water from regions of mineral weathering could confer plant and ecosystem resilience during low water years and may impact the rate of bedrock weathering and watershed chemistry during drought. This dataset includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type; a location metadata file (locations.csv); and a samples metadata file (samples.csv). All files are provided as comma-separated values (CSV) files (.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Near-Complete Sampling of Forest Structure from High-Density Drone Lidar Demonstrated by Ray Tracing

Drone lidar has the potential to provide detailed measurements of vertical forest structure throughout large areas, but a systematic evaluation of unsampled forest structure in comparison to independent reference data has not been performed. Here, we used ray tracing on a high-resolution voxel grid to quantify sampling variation in a temperate mountain forest in the southwest Czech Republic. We decoupled the impact of pulse density and scan-angle range on the likelihood of generating a return using spatially and temporally coincident TLS data. We show three ways that a return can fail to be generated in the presence of vegetation: first, voxels could be searched without producing a return, even when vegetation is present; second, voxels could be shadowed (occluded) by other material in the beam path, preventing a pulse from searching a given voxel; and third, some voxels were unsearched because no pulse was fired in that direction. We found that all three types existed, and that the proportion of each of them varied with pulse density and scan-angle range throughout the canopy height profile. Across the entire data set, 98.1% of voxels known to contain vegetation from a combination of coincident drone lidar and TLS data were searched by high-density drone lidar, and 81.8% of voxels that were occupied by vegetation generated at least one return. By decoupling the impacts of pulse density and scan angle range, we found that sampling completeness was more sensitive to pulse density than to scan-angle range. There are important differences in the causes of sampling variation that change with pulse density, scan-angle range, and canopy height. Our findings demonstrate the value of ray tracing to quantifying sampling completeness in drone lidar.

47 OTHER INSTRUMENTATION↗

Hydride and Seek: Comparing Crystallographic Hydride Placement Techniques with an Open-Shell Cobalt Complex

Locating hydrides is crucial in organometallic chemistry but difficult to do accurately using X-ray diffraction. Electron diffraction has been proposed as a way to overcome this problem but has not been systematically compared to neutron diffraction and to quantum crystallography (Hirshfeld atom refinement, HAR) to test this hypothesis. Here, we present a comparative analysis of methods for a terminal cobalt hydride complex by comparing a single-crystal neutron diffraction reference structure to results from single-crystal X-ray diffraction with and without Hirshfeld atom refinement (HAR, NoSpherA2), density functional theory (DFT), and electron diffraction (3D-ED/MicroED) refined under kinematical and dynamical formalisms. Conventional X-ray diffraction gives lower precision than neutron diffraction as expected. Despite expected improvements, HAR gives systematic deviation from the neutron benchmark. Interestingly, optimized DFT equilibrium geometries are closer to the neutron value than the value from HAR. On the other hand, electron diffraction with a high-quality data set coupled with dynamical refinement localizes the hydride in difference maps and gives excellent agreement with the neutron data. Dynamical refinement is crucial, as kinematical refinement does not allow assignment of a hydride peak. This cross-modal comparison defines the conditions under which 3D-ED/MicroED delivers high-precision metal–hydride distances for this open-shell cobalt hydride.

anions↗

Fungal Spore Seasons Advanced Across the US Over Two Decades of Climate Change

Abstract Phenological shifts due to climate change have been extensively studied in plants and animals. Yet, the responses of fungal spores—organisms important to ecosystems and major airborne allergens—remain understudied. This knowledge gap limits our understanding of their ecological and public health implications. To address this, we analyzed a long‐term (2003–2022), large‐scale (the continental US) data set of airborne fungal spores collected by the US National Allergy Bureau. We first pre‐processed the spore data by gap‐filling and smoothing. Afterward, we extracted 10 metrics describing the phenology (e.g., start and end of season) and intensity (e.g., peak concentration and integral) of fungal spore seasons. These metrics were derived using two complementary but not mutually exclusive approaches—ecological and public health approaches, defined as percentiles of total spore concentration and allergenic thresholds of spore concentration, respectively. Using linear mixed‐effects models, we quantified annual shifts in these metrics across the continental US. We revealed a significant advancement in the onset of the spore seasons defined in both ecological (11 days, 95% confidence interval: 0.4–23 days) and public health (22 days, 6–38 days) approaches over two decades. Meanwhile, total spore concentrations in an annual cycle and in a spore allergy season tended to decrease over time. The earlier start of the spore season was significantly correlated with climatic variables, such as warmer temperatures and altered precipitations. Overall, our findings suggest possible climate‐driven advanced fungal spore seasons, highlighting the importance of climate change mitigation and adaptation in public health decision‐making.

Environmental Sciences & Ecology↗

Emerging Technologies for Privacy Preservation in Energy Systems

This study explores the intersection of digitalization and privacy within the energy sector, focusing on the emerging challenges and opportunities presented by integrating Distributed Energy Resources (DERs) and advanced metering infrastructure. The need for robust digital privacy measures has become crucial as the energy industry evolves towards a more decentralized, digitalized, and decarbonized future. This study delves into four cutting-edge privacy-preserving technologies—Homomorphic Encryption (HE), Secure Multiparty Computation (SMPC), Differential Privacy (DP), and Federated Learning (FL)—each offering unique solutions to safeguard consumer data by increasing digital connectivity and data exchange. Through a detailed examination of these methods, the study explains how each technology operates, its applications within the energy sector, and the specific privacy challenges it addresses. Homomorphic Encryption allows for secure computations on encrypted data, enabling data analysis without compromising privacy. Secure Multiparty Computation enables collaborative data analysis across different entities while protecting the confidentiality of the inputs. Differential Privacy introduces randomness into the assembled data set, preventing the identification of individual records in statistical databases. Lastly, Federated Learning offers a paradigm shift in data analysis, where machine learning models are trained at the edge, minimizing the centralization of sensitive data. The research underscores the significance of implementing these privacy-enhancing technologies to comply with strict data protection regulations, foster consumer trust, and enhance the security of the energy infrastructure. By providing a comprehensive overview of these methodologies and their practical implications for the energy sector, this study aims to contribute to the ongoing discourse on digital privacy, offering insights into how the energy industry can navigate the complexities of data privacy in the digital age.

Cali, Umit↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

Machine Learning and Data Science to Advance Laboratory Earthquake Prediction and Illuminate the Mechanics of Precursors to Failure

Earthquakes represent one of our greatest natural hazards and in recent years human induced seismicity is adding to the threat. Even a modest improvement in the ability to forecast devastating large earthquakes or smaller shallow events associated with fluid injection could save thousands of lives and billions of dollars. Current efforts to forecast earthquakes are limited by knowledge of earthquake physics and hampered by a lack of reliable lab or field observations. However, recent work has provided a critical opportunity for advancement. We have found: 1) clear and consistent precursors prior to earthquake-like failure in the laboratory and 2) that lab earthquakes can be predicted using machine learning (ML). These works show that stick-slip failure events –the lab equivalent of earthquakes– are preceded by a cascade of micro-failure events that radiate elastic energy in a manner that foretells catastrophic failure. Remarkably, ML predicts the fault zone stress state, the failure time and in some cases the magnitude of lab earthquakes. In addition, the observations include clear precursors to failure in the form of changes in fault zone properties prior to lab earthquakes. Precursors have been observed in previous laboratory studies but their origin is poorly understood and their possible connection to ML based earthquake prediction is unknown. The work conducted under our project has dramatically expanded these efforts. We have developed an integrated data science approach to illuminate the physics of earthquake precursors and lab earthquake prediction. Our work has accelerated the development of ML, artificial intelligence (AI), and related data science approaches by providing massive data sets that are tightly connected to critical scientific problems and by bringing together leading subject matter experts and data scientists. Earthquake physics involves phenomena that are far from equilibrium. Our work has leveraged data science methods to illuminate these phenomena and investigate how they relate to earthquake prediction. In addition to a large database with many types of labeled events that is available to everyone, our work has advanced the fundamental understanding of seismic forecasting, earthquake physics, and fault rheology

58 GEOSCIENCES↗

Soil Moisture Data for TRACER project (Houston, TX)

The purpose of this study was to collect and distribute ground-truth soil water content and meteorological data in the Houston, TX, area, supporting the Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility, and the 2022 field campaign for the Tracking Aerosol Convection Interactions ExpeRiment (TRACER). The files herein contain soil water content and meteorological data for three stations that the Bureau of Economic Geology at UT Austin installed in the Houston, TX, area during the period of performance. The data includes soil moisture, volumetric water content, electrical conductivity of soil, soil temperature, rain precipitation, air temperature and other parameters. Data files within this data set contain measurements at both sub-hourly and 1-hour resolution measurements. The sub-hourly meteorological data file names end with “TRACER_SubHourly_met.dat”, and the sub-hourly soil data files end with “TRACER_SubHourly_soil.dat”. For the 1-hour resolution data, the files ending with "_Soil_flagged.dat” contain mean hourly volumetric soil water content and temperature measured at 5, 10, 20 and 50 cm depths. The files ending with “_Meteoro_flagged.dat” contain mean hourly measured precipitation, air temperature and humidity, wind speed and direction, and solar radiation. All data have undergone QA/QC procedures that are described by Caldwell et al. (2019) and Dorigo et al. (2013) for the soil-specific data, and EPA (2008) for the meteorological data.

Air temperature↗