Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

3D Continuous Forcing Dataset from 3D Constrained Variational Analysis at SGP

The continuous 3D large-scale forcing (VARANAL3D) data set derived from 3D constrained variational analysis (3DCVA) extends the conventional constrained variational analysis method by incorporating multiple sub-columns within the analysis domain. This advancement introduces spatial variability into the large-scale forcing fields, thereby enriching the data set’s applicability. The VARANAL3D data set spans from 2004 to 2018 and covers a region of 5˚×4.5˚ domain around the ARM SGP site. The analysis domain is divided into 10×9 sub-columns with 0.5˚ resolution. The 3D large-scale forcing data provides necessary variables to drive and evaluate single-column models (SCM), cloud-resolving models (CRM) ,and large-eddy simulations (LES), as well as information for testing model sensitivity to spatial variability of the large-scale forcing data, facilitating more rigorous testing and refinement of physical processes in SCM/CRM/LES.

54 ENVIRONMENTAL SCIENCES↗

Future Intensity‐Duration‐Frequency Curves of Extreme Precipitation in the Midwest United States From Convection‐Permitting Modeling

Abstract During the last four decades, global warming has statistically significant intensified extreme precipitation events in the Midwestern United States (defined here as the region covering Illinois, Indiana, Ohio, and Kentucky), leading to increased risks to human life, property, and infrastructure. To enable climate change adaptation and resilience across various economic and social sectors in this region, updated information about future climate changes, specifically at finer spatial scales, is essential. Leveraging a new 150‐year dynamical downscaling data set at convection‐permitting resolution, this study introduces a framework to construct the projected future intensity‐duration‐frequency (IDF) curves of heavy precipitation, which are prominent tools for infrastructure design and water resources management. This framework generates IDF curves at both sub‐daily and multi‐day duration utilizing hourly in situ observations as well as quantile‐based statistical techniques in bias‐correction and return levels selection. The assumption of non‐stationarity in the distribution parameter fitting process is also implemented in this workflow. Compared to historical IDF curves for 1980–2022, future projected IDF curves for 2058–2100 under Representative Concentration Pathway (RCP) 4.5 and RCP 8.5 scenarios indicate an average intensity increase of approximately 15% and 25%, respectively, across 74 stations, considering both annual and seasonal timescales. Future projections suggest that extreme precipitation events may become more severe across six investigated return periods, with longer return periods showing a greater increase. The frequency of future extreme precipitation events in the Midwest region is also projected to double. Furthermore, current results reveal spatial heterogeneity of future trends across stations owing to the high‐resolution input data set. Plain Language Summary This study investigates the evolving nature of extreme precipitation events in the Midwestern United States under a changing climate. By leveraging a high‐resolution dynamical downscaling data set, we construct projected intensity‐duration‐frequency (IDF) curves for future extreme rainfall events. These curves serve as vital tools for infrastructure planning and water resource management. Our analysis reveals a significant increase in both the intensity and frequency of extreme precipitation events in the region. Future projected IDF curves for the late century indicate an average intensity increase of approximately 15%–25% compared to historical values. Moreover, the frequency of such events is expected to double. Spatial heterogeneity in future trends is observed across different stations within the Midwest, highlighting the importance of high‐resolution modeling in capturing localized climate variability. These findings underscore the urgent need for climate adaptation strategies to mitigate the increasing risks associated with extreme precipitation events in the region. Key Points This study introduces a workflow to construct future intensity‐duration‐frequency (IDF) curves over the Midwest United States using a new convection‐permitting modeling data set The current IDF construction workflow reproduces well the historical observed IDF 30 curves in summer months with median relative errors of 2.4% among 74 stations and 6 investigated durations The projected IDF curves show diverse future trends of extreme precipitation across stations, with intensity increases of approximately 15% and 25% under RCP4.5 and RCP8.5 climate scenarios, respectively, and a doubling of frequency on average

Nguyen, Trung↗

Plan Position Indicator Hydrometeor Field Statistics (PPIHYD) Evaluation Data Product Version 1.0

The PPIHYD evaluation data product provides distinct hydrometeor field statistics calculated from U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility scanning radar plan position indicator (PPI) scans. These statistics include the equivalent reflectivity factor and Doppler spectral width percentiles, min/max values, and first four moments (mean, standard deviation, skewness, and kurtosis) of distinct hydrometeor features (clustered hydrometeor fields). Statistics also include morphological properties, water content and precipitation rate parameterization-based estimates, and thermodynamic properties interpolated using the Interpolated Sonde value-added product (INTERPSONDE VAP). The data set is organized in tabular form and is accompanied by mask arrays with corresponding indices. This straightforward file structure simplifies scanning radar data processing and renders this data set useful for process understanding and model evaluation studies. This report describes the data set and its processing algorithm and provides some examples.

54 ENVIRONMENTAL SCIENCES↗

Prediction and Experimental Verification of Electrolyte Solvation Structure from an OMol25-Trained Interatomic Potential

A molecular-level understanding of electrolyte solvation structure and ion–ion correlations is critical to developing next-generation battery chemistries. Atomistic simulation capabilities with sufficient accuracy, speed, and transferability to deliver reliable structural insights while avoiding arduous system-specific reparameterization are thus highly desirable. Machine learning interatomic potentials (MLIPs) trained on large, chemically diverse data sets are revolutionizing computational chemistry, enabling molecular dynamics simulations of battery electrolytes with near-DFT accuracy over 10,000× faster than DFT. While previous MLIP training data sets with suitable elemental coverage for electrolytes have been based on inorganic materials, the Open Molecules 2025 (OMol25) data set provides large-scale molecular DFT MLIP training data with broad elemental coverage and specifically samples tens of millions of electrolyte configurations. Here, we integrate computational modeling with experimental validation to systematically assess the ability of large-scale MLIPs pretrained on materials data or on OMol25 to accurately resolve nanoscale structural organization and ion-solvation characteristics in Na-ion battery electrolytes across diverse physicochemical conditions and compositional regimes. We find that the OMol25-trained Universal Model of Atoms (UMA-OMol) predicts experimentally measured densities and X-ray structure factors in substantially better agreement compared to state-of-the-art models trained only on inorganic materials data. Using UMA-OMol, we further analyze systematic trends in solvation structure as a function of cation identity, anion chemistry, salt concentration, and solvent topology. We observe that increasing system temperature amplifies the heterogeneity within the solvation environment, perturbing cation–solvent interactions and promoting the formation of contact ion pairs (CIPs). Moreover, subtle variations in the solvent topology of glyme-based electrolytes cause pronounced changes in ion correlations and solvation structure. The experimental agreement and microscopic insights shown here position OMol25-trained MLIPs as a practical route to predictive, high-throughput electrolyte simulations beyond the limits of classical force fields and direct DFT molecular dynamics, serving as a powerful tool for accelerating the design of next-generation Na-ion battery electrolytes and beyond.

MLIPs↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

The Collaborative Seismic Earth Model: Generation 2

Geological interpretations, earthquake source inversions and ground motion modeling, among other applications, require models that jointly resolve crustal and mantle structure. With the second generation of the Collaborative Seismic Earth Model (CSEM2), we present a global multi-resolution tomographic Earth model that serves this purpose. The model evolves through successive regional- and global-scale refinements. While the first generation aggregated regional models, with this study, we ensure consistency between all individual submodels, resulting in a model that accurately explains wave propagation across scales. Recent regional tomographic models were incorporated, comprising continental-scale inversions for Asia and Africa, as well as regional inversions for the Western US, Central Andes, Iran, and Southeast Asia. Across all regional refinements, over 793,000 source-receiver pairs contributed. Moreover, the long-wavelength Earth model (LOWE) introduces large-scale structures outside of pre-existing local refinements. A full-waveform inversion for global anisotropic P-and S-wave speed structure over a total of 194 iterations with a minimum period of 50 s on a large data set of 1 hr of waveform data from 2,423 earthquakes and over 6 million source-receiver pairs ensures that regional updates in the crust and uppermost mantle translate into updates of deeper, global-scale structure. To test the performance of CSEM2, we evaluate waveform fits between observed and synthetic seismograms at 50 s for an independent data set on the global scale, and on the regional scale for lower periods. We accurately simulate waveforms within and across regional refinements, maintaining the original resolution of the submodels embedded in the global framework.

58 GEOSCIENCES↗

Temporal and spatial characterization of a thermogenic, fault-controlled gas hydrate system, Woolsey Mound, Gulf of Mexico

Woolsey Mound, located at Mississippi Canyon Lease Block 118 (MC118), is the site of the Gulf of Mexico hydrate research consortium’s seafloor observatory, where gas hydrates outcrop at the seafloor. The presence of gas hydrates in the mound is confirmed directly by coring and indirectly by 3D seismic reflection data. Craters, pockmarks, chemosynthetic communities, and authigenic carbonates populate the seafloor at Woolsey Mound. Each crater is characterized by a network of shallow crestal faults that connect the hydrate mound to the underlying allochthonous salt body. We characterize the temporal and spatial evolution of gas hydrates at Woolsey Mound under natural perturbations using four collocated 3D seismic reflection data sets that span over 14 years. Data acquisition differences embedded in the data sets arising from variation in geometry, sample rate, and phase are minimized using the “cross-equalization” method. Our results indicate that hydrate formation and dissociation vary temporally and spatially in close connection to the shallow crestal faults. Evidence of gas hydrate dissociation is observed over a period of three years (2000–2003), where major dissociation occurred along the southern portion of the crestal fault in the southeast crater. The dissociation is less prominent in the southwest crater. Evidence of methane venting is observed between 2000 and 2010, which is mostly concentrated in the southeast crater. The residual amplitude anomalies observed between 2000 and 2014 in the mound are mostly positive, implying that the methane venting had increased significantly. The positive anomalies are correlated with the methane seepage recorded in 2011. Our results indicate the evolution of a fault-controlled gas hydrate system in the northern Gulf of Mexico, which would aid in assessing its impact on the seafloor.

Geochemistry & Geophysics↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

Machine-learning-informed scattering correlation analysis of sheared colloids

We have carried out theoretical analysis, Monte Carlo simulations and machine-learning analysis to quantify microscopic rearrangements of dilute dispersions of spherical colloidal particles from coherent scattering intensity. Both monodisperse and polydisperse dispersions of colloids were created and underwent a rearrangement consisting of an affine simple shear and non-affine rearrangement using the Monte Carlo method. We calculated the coherent scattering intensity of the dispersions and the correlation function of intensity before and after the rearrangement and generated a large data set of angular correlation functions for varying system parameters, including number density, polydispersity, shear strain and non-affine rearrangement. Singular value decomposition of the data set shows the feasibility of machine-learning inversion from the correlation function for the polydispersity, shear strain and non-affine rearrangement using only three parameters. A Gaussian process regressor is then trained on the data set and can retrieve the affine shear strain, non-affine rearrangement and polydispersity with relative errors of 3%, 1% and 6%, respectively. Altogether, our model provides a framework for quantitative studies of both steady and non-steady microscopic dynamics of colloidal dispersions using coherent scattering methods.

Gaussian process regression↗

Bayesian Inference for the Seismic Moment Tensor Using Regional Waveforms and Teleseismic- P Polarities with a Data-Derived Distribution of Velocity Models and Source Locations

The largest source of uncertainty in any source inversion is the velocity model used in the transfer function that relates observed ground motion to the seismic moment tensor. However, standard inverse procedure often does not quantify uncertainty in the seismic moment tensor due to error in the Green’s functions from uncertain event location and Earth structure. Here, we incorporate this uncertainty into an estimation of the seismic moment tensor using a data-derived distribution of velocity models based on complementary geophysical data sets, including thickness constraints, velocity profiles, gravity data, surface-wave group velocities, and regional body-wave travel times. The data-derived distribution of velocity models is then used as a prior distribution of Green’s functions for use in Bayesian inference of an unknown seismic moment tensor using regional and teleseismic-P waveforms. The use of multiple data sets is important for gaining resolution to different components of the moment tensor. The combined likelihood is estimated using data-specific error models and the posterior of the seismic moment tensor is estimated and interpreted in terms of the most probable source type.

58 GEOSCIENCES↗

LANL Meteorological Program: 2023 Data Completeness/Quality Report

Los Alamos National Laboratory (LANL) operates seven mesa-top instrumented meteorology towers: Technical Area (TA) 6, TA-49, TA-53, TA-54, TA-63, TA-54B, and TA-16B. An additional instrumented tower is located in Mortandad Canyon (TA-5 MDCN), and there is a rain gauge at North Community (NCOM), located within the town of Los Alamos. The 10 meter (m) towers at TA-63, TA-54B, and TA-16B have been in testing since they were installed in 2021, and will be included in a future data completeness report. A description of the meteorology monitoring network, prior to the installation of the TA-63, TA-54B, and TA-16B is found in Dewart and Boggs (2014). Four of the mesa-top towers (e.g., TA-6, TA-49, TA-53, and TA-54) are instrumented at the 1.2 m, 11.5 m, 23 m, and 46 m levels. In addition, the TA-6 tower is instrumented at the 92 m level. The TA-5 MDCN tower is 10 m in height and is instrumented at 1.2 m and 10 m. Data are collected and averaged every 15 minutes. Range checking is performed on each measurement every 15 minutes; data that are beyond normal ranges are eliminated from the data set and replaced by a code for missing data. In addition, data are reviewed weekly by qualified meteorologists to identify bad data not identified by the range checking technique. The data steward eliminates these data from the data set and replaces them with a code for missing data. The instrument technicians also review that data and schedule instrument replacement, as required. All instruments are calibrated at a frequency that meets the criteria identified in ANSI/ANS-3.11-2015. Data completeness is determined by the number of total 15-minute records available versus the number of possible measurements for the entire year. As a rule, the meteorologists do not attempt to estimate data that are eliminated as bad data. Original datalogger records, including bad data, can be recalled from program archival storage.

54 ENVIRONMENTAL SCIENCES↗

NPFTURBULENCE: Turbulent Parameters by airborne measurements

The original data were collected during the field campaign of “Turbulent layers promoting New Particle Formation” experiment (NPFTURBULENCE; https://www.arm.gov/research/campaigns/aaf2024npfturbulence) over the Atmospheric Radiation Measurement (ARM) user facility's Southern Great Plains (SGP) atmospheric observatory (https://www.arm.gov/capabilities/observatories/sgp ) in north-central Oklahoma. The ARM Aerial Facility ArcticShark uncrewed aerial system (UAS, https://www.arm.gov/capabilities/observatories/aaf/uas) was based at Blackwell–Tonkawa Municipal Airport (IATA: BWL, ICAO: KBKN, FAA LID: BKN, 36.74475° N, 97.34918° W, 313.9 m MSL), for the field campaign from May 5 through May 29, 2024. The ArcticShark UAS performed 11 flights, including 10 research flights over the Central Facility of the ARM SGP to measure atmospheric state, turbulence, surface IR temperature and imagery, and aerosol number concentration and size distribution. The current data set presents a comprehensive collection of turbulent parameters in the atmospheric boundary layer or lower free troposphere based on airborne measurement throughout the field campaign. The primary instruments used to create the current data set were the Aircraft Integrated Meteorological Measurement System (AIMMS-30) and the fine-wire thermocouple probe. For user convenience, the current data set includes several parameters commonly used in turbulent research for normalization and/or scaling: atmospheric boundary-layer height, surface conditions, convective scales for temperature, and velocity, etc.

54 ENVIRONMENTAL SCIENCES↗

Multifrequency radar observations of marine clouds during the EPCAPE campaign

The Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) was a year-round campaign conducted by the US Department of Energy at the Scripps Institution of Oceanography in La Jolla, CA, USA, with a focus on characterizing atmospheric processes at a coastal location. The ground-based prototype of a new Ka-, W-, and G-band (35.75, 94.88, and 238.8 GHz) profiling atmospheric radar, named CloudCube, which was developed at the Jet Propulsion Laboratory, took part in the experiment during 6 weeks in March and April 2023. This article describes the unique data sets that were obtained during the field campaign from a variety of marine clouds and light precipitation. These are, to the best of the authors' knowledge, the first observations of atmospheric clouds using simultaneous multifrequency measurements including 238.8 GHz. These data sets therefore provide an exceptional opportunity to study and analyze hydrometeors with diameters in the millimeter- and submillimeter size range that can be used to better understand cloud and precipitation structure, formation, and evolution. The data sets referenced in this article are intended to provide a complete, extensive, and high-quality collection of G-band data in the form of Doppler spectra and Doppler moments. In addition, Ka-band and W-band reflectivity and Ka-, W-, and G-band reflectivity ratio profiles are included for several cases of interest on 6 different days.

54 ENVIRONMENTAL SCIENCES↗

Fast Response Temperature by airborne measurements over BNF

The original data were collected during the AAF Engineering Flights (AEF2025) in the vicinity of the ARM Bankhead National Forest (BNF) Atmospheric Observatory (https://www.arm.gov/capabilities/observatories/bnf ) in northwestern Alabama in March 2025. The ARM Aerial Facility ArcticShark uncrewed aerial system (UAS, https://www.arm.gov/capabilities/observatories/aaf/uas) was based at the public-use airport of Posey Field, Alabama (FAA LID: 1M4, 34.28027778° N, 87.60055556° W, 283m MSL) from March 9 through March 24, 2025. The ArcticShark UAS performed nine flights, including eight research flights over the AMF3 (BNF Main Site) and Supplemental Facilities to measure atmospheric state, turbulence, surface IR temperature and imagery, and aerosol number concentration and size distribution. The current data set presents fast response temperature in the atmospheric boundary layer and lower free troposphere measured on the airborne platform throughout the field campaign. The primary instruments used to create the current data set were the fine wire thermocouple probe, the Aircraft Integrated Meteorological Measurement System (AIMMS-30), the Pitot-static system (part of UAS flight control), and the infrared gas analyzer sensor for H2O and CO2 (LI-840). All parameters used in temperature calculations (static pressure, True Air Speed, and absolute humidity in form of dew point temperature) were included in the data set. For user convenience, one additional parameter was also included: the type of flight flag (level, up, down, turn, and combination of thereof).

Air temperature, fast response↗

CHELAX-BNF: Fast Response Temperature by airborne measurements

The original data were collected on board the ARM Aerial Facility ArcticShark uncrewed aerial system (UAS; https://www.arm.gov/capabilities/observatories/aaf/uas ) during the “Characterizing HEterogeneous Land-Atmosphere eXchanges at BNF” field campaign (CHEAX-BNF; https://arm.gov/research/campaigns/aaf2025CHELAX-BNF ). The ARM Aerial Facility ArcticShark UAS was based at the public-use airport of Posey Field, AL (FAA LID: 1M4, 34.28027778° N, 87.60055556° W, 283m MSL) from May 28 to June 23, 2025. The ArcticShark UAS performed 5 flights, including 4 research flights over the BNF Main Site (ARM Mobile Facility 3, https://arm.gov/capabilities/observatories/amf ) and Supplemental Facilities to measure atmospheric state, turbulence, surface IR temperature and imagery, aerosol number concentration, and aerosol size distribution. The current data set presents fast response temperature in the atmospheric boundary layer and lower free troposphere measured on the airborne platform throughout the field campaign. The primary instruments used to create the current data set were the fine wire thermocouple probe, the Aircraft Integrated Meteorological Measurement System (AIMMS-30), the Pitot-static system (part of UAS flight control), and the infrared gas analyzer sensor for H2O and CO2 (LI-840). All parameters used in temperature calculations (static pressure, True Air Speed, and absolute humidity in the form of dew point temperature) were included in the data set. For user convenience, one additional parameter was also included: the type of flight flag (level, up, down, turn, and combination of thereof).

Air temperature, fast response↗

Tokamak divertor plasma emulation with machine learning

Abstract Future tokamak devices that aim to create conditions relevant to power plant operations must consider strategies for mitigating damage to plasma facing components in the divertor. One of the goals of MAST-U tokamak operations is to inform these considerations by researching advanced divertor configurations that aid stable plasma detachment. Machine design, scenario planning and detachment control would all greatly benefit from tools that enable rapid calculation of scenario-relevant quantities given some input parameters. This paper presents a method for generating large, simulated scrape-off layer data sets, which was applied to generate a data set of steady-state Hermes-3 simulations of the MAST-U tokamak. A machine learning model was constructed using a Bayesian approach to hyperparameter optimisation to predict diagnosable output quantities given control-relevant input features. The resulting best-performing model, which is based on a feedforward neural network, achieves high accuracy when predicting electron temperature at the divertor target and carbon impurity radiation front position and runs in around 1 ms in inference mode. Techniques for interpreting the predictions made by the model were applied, and a high-resolution parameter scan of upstream conditions was performed to demonstrate the utility of rapidly generating accurate predictions using the emulator. This work represents a step forward in the design of machine learning-driven emulators of tokamak exhaust simulation codes in operational modes relevant to divertor detachment control and plasma scenario design.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Urban Parameters Los Angeles County 100m version 2

132 Urban parameters based on building physical dimensions and location were generated for the city of Los Angeles at 100m resolution using the NATURF model. To use the binary file with WRF, the binary file and the index file must be placed in their own directory in WRF_GEOG and accessed in the same way NUDAPT44 would be accessed. The kmz file can be visualized on Google Earth.

Sweet-Breu, Levi [Baylor University]↗