Search NASA⌕ Search

SEARCH · Search NASA

Results for “database improvement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Validation of an Erythema-Weighted UV Model Using Broadband Solar Irradiance Measurements From Eleven U.S. Sites: Preprint

Erythema-weighted UV solar irradiance (UV-E) has a potential impact on human health if the recommended maximum exposure times are exceeded. In spite of this, it is not measured at most sites that measure Global Horizontal Irradiance (GHI). However, since UV-E is highly correlated with GHI and total ozone content it can be estimated from this information with sufficient accuracy to assisst in public health recommendations. The Power Model (PM) provides a simple method to estimate the erythema-weighted UV irradiance (UV-E) from measured GHI, total ozone column and air mass. In this work, the performance of the PM method is assessed using high-quality data from 11 sites in the continental U.S. (part of SURFRAD and SOLRAD networks) and total ozone estimates publicly available from the MERRA-2 re-analysis database. A three year period (2021- 2023) at 1-minute frequency is considered. The results (for time aggregations of 5 and 60 minutes) show a high Pearson's correlations (> 0.99), consistently positive mean bias deviations (below 13%) and dispersions in the 8-19% range at all sites. Relative values are expressed in terms of the corresponding measurement mean. The performance indicators remain consistent across time resolutions (5 or 60 minutes), suggesting that the model's performance is robust and not significantly affected by short-term variability (which is captured by GHI). Spatial patterns reveal higher biases and RMSD in northern and eastern locations. These values represent a significant improvement over widely used satellite-based global UV-E estimates and open the possibility of using the PM with satellite-based GHI estimates for operational UV-E mapping over the contiguous U.S. territory.

14 SOLAR ENERGY↗

The influence of cloud cover on the reliability of satellite-based solar resource data

Satellite-based solar resource data are often developed and validated by using binary cloudiness categories: clear sky or overcast cloudy sky. To investigate the reliability of solar resource data in partially cloudy conditions, we estimate cloud fraction using two distinct algorithms: a physical retrieval model using surface observed global horizontal irradiance (GHI) and direct normal irradiance (DNI) and a temporal average of cloud mask data estimated by the observed DNI. Our analysis reveals a significant presence of scattered clouds, broken clouds, and mismatches between satellite- and surface-based cloud data at 17 surface sites across the contiguous United States, though confidently clear and cloudy conditions collectively account for more than 70 % of the data. Solar radiation is computed using the National Solar Radiation Database (NSRDB) algorithm and validated using surface observations. Here, our findings suggest that, in the presence of scattered clouds, NSRDB data for clear-sky conditions can be subject to significant overestimation. In cloudy-sky conditions classified by satellite data, DNI computed by the Fast All-sky Radiation Model for Solar applications with DNI (FARMS-DNI) can be underestimated when limited clouds are detected by surface observations. The bias observed in several cloudiness categories indicates that the NSRDB is exceptionally accurate in confidently clear conditions. However, clear-sky conditions with scattered clouds and mismatched cloud data contribute significantly to the overall uncertainties in the NSRDB. Therefore, future improvements in solar resource data should involve development and implementation of satellite-derived cloud fraction and should consider a novel radiative transfer model accounting for amplified cloud reflection. The evaluation within cloudiness categories also provides a physical rationale for the superior performance of FARMS-DNI compared to the Direct Insolation Simulation Code (DISC) in both cloudy-sky and all-sky conditions.

14 SOLAR ENERGY↗

PAVC Gridded 20m Alaska NGEE Tier3 PFTs v1.0

These 20-meter spatial resolution gridded products provide per-pixel fractional cover (%) of Next Generation Ecosystem Experiments (NGEE) Arctic Plant Functional Types (PFTs) Tier 3 across Alaska, north of the boreal treeline. The products were developed for the NGEE Arctic project, which is improving Arctic vegetation representation and parameterization of the E3SM Land Model. This dataset includes 8 files containing fractional cover for NGEE Tier 3 PFTs (https://data.ess-dive.lbl.gov/view/doi:10.15485/2529470): (1) bryophytes; (2) lichens; (3) non-vascular plants, i.e., the sum of lichens and bryophytes; (4) deciduous shrubs, (5) evergreen shrubs, (6) forbs, (7) graminoids, and a non-PFT class, (8) litter. Each pixel contains the percent cover (expressed as a fraction of total ground cover) that was predicted by random-forest regression models. The random-forest models were trained on cover data collected at 978 plots from 2010 to 2021, of which are archived in the Pan-Arctic Vegetation Cover (PAVC) database (https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2483557). The plot cover was linked to 20-meter spatial resolution, satellite-derived predictor variables: Sentinel-2 spectra and Sentinel-1 polarizations averaged over the 2019 growing season, as well as topographical features derived from ArcticDEM. Then, spatio-temporally anomalous plot data that introduced large variability to the regression outcomes were dropped using the Cook’s distance outlier detection method, and the models were re-created using high-quality plots and their associated satellite derived explanatory variables per each PFT. The correlations between plot-observed and satellite-derived fractional cover for all PFTs were well correlated (R2 = 0.69–0.95 and 0.5 for litter) and had low RMSE bias (0.02–0.11). This research was performed as a part of the NGEE Arctic project. The NGEE Arctic project was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.

54 ENVIRONMENTAL SCIENCES↗

Generating An Advanced Cross-section Library For HTGR Pebble Bed Depletion Calculations Using Reduced-Order Model Generation Techniques

For code development, Advanced Reactor Technologies - Gas Cooled Reactors Program (ART-GCR) rely on a collaboration with the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, but the cross sections generation and the methodology definition is part of this program area goals. Based on previous studies in FY23, the size of microscopic cross section libraries increases rapidly with the number of tabulations, requiring significant amount of memory and drastically slowing down the Griffin calculations when evaluating cross sections via the multivariate linear interpolation approach. Rising to these challenges, this work investigates constructing Reduced-order Models (ROMs) for the multi-group microscopic cross sections to accelerate the cross section evaluation in Griffin. A database of multigroup cross sections is first collected considering all possible parameters that a designer could change for optimization. Down-selection of the ROM techniques afterward shows Deep Neural Network (DNN) as the best candidate when jointly consider memory efficiency, predictive accuracy, computational cost, scalability, flexibility and ease of implementation of the algorithms in comparison to the multidimensional interpolation. This work develops a specific interface that enables the cross section predictions using pre-trained DNN models into Griffin leveraging the existing ROM capabilities. DNNs have been trained for all isotopes for use in Griffin. Preliminary Griffin testing shows that DNNs exhibit exceptional predictive accuracy and the use of DNNs provides orders of magnitude improvement in memory efficiency compared to conventional interpolation techniques. With such ROM techniques, it holds great promise to further increase the fidelity of the Pebble Bed Reactor (PBR) simulation by increasing the number of tabulations/state variables during cross section evaluation, while maintaining the computational cost affordable in Griffin.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Forming a database to study reversed magnetic shear from the National Spherical Torus eXperiment using machine learning

Achieving a long-lived reversed magnetic shear (RMS) target plasma in the National Spherical Torus eXperiment Upgrade will require developing various sustainment scenarios. To help with the ongoing plasma control efforts, the development of a new analysis for the motional Stark effect (MSE) diagnostic using a machine learning algorithm, namely, MSE-ML, is described. MSE-ML will be used to identify patterns during RMS discharges, some of which suffer magnetohydrodynamic (MHD) events resulting in current redistribution and monotonic q-profiles. A database consisting of q and magnetic shear profiles is being constructed primarily based on the existing National Spherical Torus eXperiment data with equilibrium reconstructions constrained by the magnetic field pitch angle profile measured using the multi-channel MSE diagnostic. An unsupervised k-means clustering of the data is developed to study the RMS formation as a function of time. The initial clustering from the q-profiles shows significant differences in both amplitude and the duration of the RMS period. As a goal, the clustering results that detect and distinguish shots with substantial and sustained RMS are to be used as a preprocessing step in a supervised algorithm to identify the underlying conditions that lead to long-lasting improved confinement with RMS. Another aim of the MSE-ML study is to identify precursors of RMS-destroying MHD events in either derived data such as the q-profile or directly measured data such as the magnetic field pitch angle profile.

Uzun-Kaymak, I. U. (ORCID:0000000276251493)↗

Endogeneity of pedestrian survival time and emergency medical service response time: Variations across disadvantaged and non-disadvantaged communities

The Vision Zero-Safe Systems Approach prioritizes fast access to Emergency Medical Services (EMS) to improve the survivability of road users in transportation crashes, especially concerning the recent increase in pedestrian-involved crashes. Pedestrian crashes resulting in immediate or early death are considerably more severe than those taking longer. The time gap between injury and fatality is known as survival time, and it heavily relies on EMS response time. The characteristics of the crash location may be associated with EMS response and survival time. A US Department of Transportation initiative identifies communities often facing challenges. Six disadvantaged community (DAC) indicators, including economy, environment, equity, health, resilience, and transportation access, enable an analysis of how survival and EMS response times vary across DACs and non-DACs. To this end, this study created a unique and comprehensive database by linking DACs data with 2017–2021 pedestrian-involved fatal crashes. This study utilizes two-stage residual inclusion models with segmentation for DACs and non-DACs accounting for the endogenous relationship between EMS response and pedestrian survival time. The results indicate that EMS response time is higher and pedestrian survival time is lower in DACs than in non-DACs. A delayed EMS response time is associated with a greater reduction in survival time in DACs compared to non-DACs. Factors, e.g., nighttime and interstate crashes, contribute to higher EMS response time, while pedestrian drugs, driver speeding, and hit-and-run behaviors are associated with a greater reduction in survival time in DACs than non-DACs. Finally, the implications of the findings are discussed in the paper.

60 APPLIED LIFE SCIENCES↗

An overview of 3D field optimization for control of transport and edge instabilities on KSTAR

An international team from several laboratories and universities has made key advances over the last few years in the control of plasma transport and edge instabilities with applied 3D fields in the KSTAR tokamak to optimize long pulse operation scenarios. This overview begins with the optimization of both core and edge resonant magnetic perturbations (RMP) to improve fast ion confinement to avoid excessive limiter heat loads due to fast ion losses and successful modeling of the experimental results. Integrated and advanced plasma control techniques with machine learning (ML) and adaptive control were then used to optimize the 3D field spectrum in real-time to control edge localized modes (ELMs) while avoiding core locked modes that could disrupt the plasma. Accelerating the offline model of 3D fields with a surrogate ML model can optimize ELM suppression in the edge while limiting the impact of the applied RMP fields deeper in the plasma core in real-time. In addition, the impact of the 3D fields on the divertor heat load has been modeled and compared with experimental measurements. An analysis of a multi-machine database including KSTAR has been performed to better understand the metrics for the observed RMP thresholds for ELM suppression and the resulting plasma performance. Predictive modeling of the operational space for ELM suppression and density pumpout due to RMP has shown the importance of magnetic islands in the plasma edge and their impact on plasma turbulence. This research has culminated in the development of successful long pulse operational scenarios on KSTAR while attempting to overcome challenges of the new tungsten divertor.

3D fields↗

Deep learning-based predictive models for laser direct drive at the Omega Laser Facility

The rich and complex physics of inertial confinement fusion provides a unique and challenging space for high-fidelity first-principles modeling. Consequently, simulation codes that are used to design experiments are computationally expensive and lack the predictive capability required for extensive parameter exploration in search of a high-performing design for laser direct drive. In this article, we present two deep-learning-based predictive models intended to address these difficulties. The first model (TL DNN) acts as a fast emulator of simulations as well as experiments at the Omega Laser Facility. This model is trained on a simulation database and subsequently calibrated on experimental data using transfer learning. To facilitate the development of this model, an autoencoder is developed to reduce the dimensionality of the input space by compressing the laser pulse input. The model predicts key experimental scalar observables of Omega experiments with high accuracy and minimal computational cost. This deep neural net enables rapid exploration of a high-dimensional input parameter space for an optimal implosion design. The second model (DNN SM+) aims to extend the statistical modeling work of Lees et al. [Phys. Rev. Lett. 127, 105001 (2021)], by increasing the complexity of the model space and allowing for coupling between degradation terms. Since the model capacity of DNN SM+ is higher than the model of Lees et al., DNN SM+ can potentially provide an improvement in predictive capability, and we use this model to provide insight into complicated degradation dependencies.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Investigation and Diagnosis of Faulty Data Channels in CMS Outer Tracker Module Testing

The High-Luminosity Large Hadron Collider (HL-LHC) is currently undergoing upgrades to improve its luminosity. In parallel, this requires an upgrade to the Compact Muon Solenoid (CMS)’s Outer Tracker, consisting of Pixel-Strip (PS) and Strip-Strip (2S) modules that can accurately track the path of charged particles originating from the collisions. It follows that such complex modules call for extensive testing, requiring a sophisticated Data Acquisition (DAQ) system that can perform specific tests to assess their performance. In addition, errors caused by the hardware of a given testing station, and its associated data channel, need to be accurately identified to guarantee proper testing of modules. We have developed a software extension to the Phase-II Outer Tracker Analyzer of Test Outputs (POTATO), which is a specialized software designed to analyze and grade all of the module tests through a centralized database. This extension categorizes and analyzes module test results by its station and data channel. Its analysis can be used to identify trends in grading that indicate issues in these channels’ grading process rather than in the individual modules. This poster shows our methodology and results for identifying faulty data channels. Using this extension, we can quickly diagnose and address problems in our DAQ system, ensuring proper evaluation corrections for each module.

Chen, Angus [Fermilab]↗

Investigation and Diagnosis of Faulty Data Channels in CMS Outer Tracker Module Testing

The High-Luminosity Large Hadron Collider (HL-LHC) is currently undergoing upgrades to improve its luminosity. In parallel, this requires an upgrade to the Compact Muon Solenoid (CMS)’s Outer Tracker, consisting of Pixel-Strip (PS) and Strip-Strip (2S) modules that can accurately track the path of charged particles originating from the collisions. It follows that such complex modules call for extensive testing, requiring a sophisticated Data Acquisition (DAQ) system that can perform specific tests to assess their performance. In addition, errors caused by the hardware of a given testing station, and its associated data channel, need to be accurately identified to guarantee proper testing of modules. We have developed a software extension to the Phase-II Outer Tracker Analyzer of Test Outputs (POTATO), which is a specialized software designed to analyze and grade all of the module tests through a centralized database. This extension categorizes and analyzes module test results by its station and data channel. Its analysis can be used to identify trends in grading that indicate issues in these channels’ grading process rather than in the individual modules. This poster shows our methodology and results for identifying faulty data channels. Using this extension, we can quickly diagnose and address problems in our DAQ system, ensuring proper evaluation corrections for each module.

Chen, Angus [Fermilab]↗

Development and Evaluation of a General Drag Model for Gas-Solid Flows via Deep Learning

This project presents the development and evaluation of a general drag model for gas–solid multiphase flows using deep learning techniques. A comprehensive database of more than 4,000 experimental and numerical data points for spherical and non spherical particles was compiled, incorporating geometric features such as sphericity, aspect ratio, and orientation. Several predictive approaches—including traditional em pirical correlations, machine learning, and deep neural networks—were benchmarked, with the proposed Drag Coefficient Correlation-aided Deep Neural Network (DCC DNN) demonstrating superior accuracy. To account for particle–particle interactions, additional drag data were generated using CFD-based simulations of packed and flu idized beds, leading to the development of a retrained model capable of incorporat ing volume fraction effects. Integration of the trained model with the MFiX CFD solver was achieved using FTorch, enabling drag predictions during discrete element method (DEM) simulations. Validation against experimental data for single particles and fluidized beds confirmed the model’s improved predictive ability, particularly for non-spherical geometries. While the model performed strongly under fluidized con ditions, limitations remained in unfluidized regimes, suggesting a need for expanded datasets. Overall, this study demonstrates the feasibility of combining deep learning with physics-informed CFD to improve drag modeling for gas–solid flows, with promis ing implications for scaling multiphase simulations in industrial applications.

42 ENGINEERING↗

Robustness of topological persistence in knowledge distillation for wearable sensor data

Topological data analysis (TDA) has shown great success in various applications involving wearable sensor data. However, there are difficulties in leveraging topological features in machine learning and wearable sensors because of the large time consumption and computational resources required to extract the features. To address this problem, knowledge distillation (KD) is utilized to generate a small model and accommodate topological features with persistence image (PI) representations from the raw time series data. Deploying topological knowledge in KD enables the student to achieve better performance compared to the one trained solely on raw time series data. However, it is not yet known if there are coherent characteristics for topological features in PI, which can aid in improving the performance during KD. In this paper, we investigate the suitability and challenges of utilizing topological features in KD for wearable sensor data, thereby contributing to the advancement of the field. Our study explores the impact of transferred topological features by comparing the Teacher-to-Student framework with Multiple Teachers-to-Student where teachers utilize both time series data and persistence images obtained by TDA as inputs. Additionally, we conduct a rigorous examination of topological knowledge effects by testing under various corruptions, knowledge types, and learning strategies in the context of human activity recognition tasks. Our analysis of topological features in KD presents the optimal strategy for incorporating these features. This study includes datasets of varying scales, window lengths, and activity classes, providing a comprehensive evaluation. Our results demonstrate that leveraging topological features in KD to enhance performance across databases.

97 MATHEMATICS AND COMPUTING↗

Southwest Regional Partnership on Carbon Sequestration: Phase III (Final Scientific/Technical Report)

The Southwest Regional Partnership on Carbon Sequestration (SWP) is one of 7 regional partnerships formed in 2003 under the U.S. Department of Energy’s (DOE) Regional Carbon Sequestration Partnerships (RCSPs) initiative. The overall purpose of the initiative was to help determine and implement the technology, infrastructure, and regulations most appropriate to promote carbon storage in different regions of the country. Covering Arizona, Colorado, New Mexico, Oklahoma, Utah, and parts of Texas, Wyoming, and Kansas, the SWP evaluated regional carbon storage and utilization potential and focused on technologies and sites that could complement the region’s strong position in energy production. The project progressed through three phases: • Phase I (2003–2005): Characterized regional geologic formations and CO 2 sources, assessed sequestration potential, and identified pilot test sites. • Phase II (2005–2013): Conducted small-scale field tests to validate sequestration methods, including geologic and terrestrial projects. • Phase III (2008–2022): Demonstrated large-scale CO 2 injection at a commercial oil field to test monitoring, verification, and long-term storage strategies. This report covers Phase III. The final project site, the Farnsworth Unit (FWU) in Texas, provided real-world testing of reservoir characterization, monitoring, and risk evaluation tools and processes that could be used in any commercial scale carbon capture, utilization, and storage (CCUS) project. Extensive data collection and analysis helped refine best practices for reservoir characterization, injection monitoring, and storage verification. The SWP contributed to national databases, DOE best practice manuals, and regional geological assessments to support future sequestration efforts. Key lessons learned include the importance of robust data management, strategic site selection, regulatory navigation, and effective industry collaboration. The project’s findings will inform ongoing and future carbon storage initiatives. Task 1 (Regional Characterization) • The SWP continued to participate in national outreach efforts and NATCARB. • The SWP evaluated multiple potential sites before selecting the FWU as the primary field test location. Task 2 (Public Outreach and Education) • The SWP contributed to national databases, DOE best practice manuals, and regional geological assessments to support future sequestration efforts. Task 3 (Permitting and Regulatory Compliance) • The SWP ensured compliance with federal and state regulations, including National Environmental Policy Act (NEPA) requirements. • The SWP obtained all necessary permits for drilling, injection, and monitoring activities. Task 4 (Site Characterization and Planning) • The SWP developed work plans for four key activities: characterization, simulation, monitoring and verification, and risk evaluation. • The SWP collected and synthesized legacy data from multiple sources to build initial static geological models and dynamic reservoir models demonstrating project feasibility. • The SWP conducted an initial risk evaluation and developed mitigation plans. Task 5 (Field Operations and Data Collection) • The SWP drilled, logged, and cored three characterization wells to gather critical subsurface data. • The SWP conducted multiple geophysical surveys, including 3D seismic, crosswell seismic, and vertical seismic profiling, to improve reservoir characterization. Task 6 (Monitoring and Verification) • The SWP performed extensive geological characterization using data from characterization wells and seismic surveys. • The SWP established a surface monitoring network to track CO 2 flux in soil gas, groundwater chemistry, and near-surface atmospheric CO 2 levels. • The SWP built and refined reservoir models to study the effects of relative permeability on simulation behavior and improve calibration with experimental data. Task 7 (Risk Assessment and Model Refinement) • The SWP conducted multiple studies to evaluate reservoir integrity, predict CO 2 plume behavior and improve predictive modeling capabilities. • The SWP refined geological models and used them to enhance the accuracy of simulation models. • The SWP continued quantitative risk assessment of top-ranked risks and strengthened the link between qualitative and quantitative risk methodologies.

02 PETROLEUM↗

Role of perturbed parallel magnetic field effects in predicting turbulent transport in NSTX

This study presents analysis of gyrokinetic simulations on the National Spherical Torus Experiment (NSTX) to investigate the effects of electromagnetic fields on plasma turbulence and transport. The simulations, performed with varying levels of fidelity using the gyrokinetic CGYRO code, include electrostatic (ES), single-field electromagnetic (EM1), and two-field electromagnetic (EM2) models. A detailed comparison across the simulation database reveals that electromagnetic effects increase both predicted growth rates and quasilinear fluxes, with EM2 simulations producing stronger turbulence than ES and EM1 cases. Quasilinear modeling using QLGYRO demonstrates that while the perturbed parallel magnetic field (δB ∥ ) does not drastically affect the total flux at experimental gradients, it leads to a shift in the dominant instability, altering mode structures from microtearing to kinetic ballooning modes (KBMs). The proximity of the plasma profiles to the KBM threshold is explored, with the experimental conditions being near the onset of KBM-driven transport. The KBM, with its large growth rates, is identified as a potential driver of electron temperature flattening, as it can rapidly transport heat across flux surfaces. Performing stability analysis shows core-localized unstable a low- mode that could contribute to the flattening at the early times of the discharge. TGYRO predictive modeling, incorporating both TGLF and QLGYRO, indicates that the inclusion of δB ∥ significantly improves the accuracy of temperature profile predictions in NSTX high-beta plasmas, although challenges remain in modeling the sharp flux discontinuities caused by KBM-driven instabilities.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Uncertainty quantification for misspecified machine learned interatomic potentials

The use of high-dimensional regression techniques from machine learning has significantly improved the quantitative accuracy of interatomic potentials. Atomic simulations can now plausibly target quantitative predictions in a variety of settings, which has brought renewed interest in robust means to quantify uncertainties. In many practical settings where model complexity is constrained (e.g., due to performance considerations), misspecification — the inability of any one choice of model parameters to exactly match all training data — is a key contributor to errors that is often disregarded. Here, we employ a recent misspecification-aware regression technique to quantify parameter uncertainties, which is then propagated to a broad range of phase and defect properties in tungsten. The propagation is performed through both brute-force resampling and implicit Taylor expansion. The propagated misspecification uncertainties robustly quantify and bound errors on a broad range of material properties. We demonstrate application to recent foundational machine learning interatomic potentials, accurately predicting and bounding errors in MACE-MPA-0 energy predictions across the diverse materials project database.

36 MATERIALS SCIENCE↗

The Pan-Arctic Vegetation Cover (PAVC) database v1.1

The Pan-Arctic Vegetation Cover (PAVC) database contains synthesized field-data observations of vegetation cover from 978 Arctic Alaska plots with observations from 2010 to 2021. The cover datasets contain plot data at both the plant functional type (PFT) and species-level resolution, with standardized PFT definitions and species names. We synthesized publicly available point-intercept and visual estimate plots from the Arctic Vegetation Archive of Alaska, the Alaska Vegetation Plots Database, the North Slope Science Catalog, and the National Ecological Observatory Network; as well as previously unpublished data from the Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic).Users will find four synthesized datasets, 4 associated data descriptor (dd) files, and 1 metadata file in the PAVC database:synthesized_species_fcover.csv contains fractional cover (fcover) for unique accepted species names, where names include vegetation identified at the family, genus, species, subspecies, and variety levels, as well as general functional types across all 5 data sources. The synthesized_species_fcover_dd.csv accompanies this dataset with header information.synthesized_pft_fcover.csv contains fcover for the following PFTs: non-vascular plants with lichen and bryophyte subcategories, trees with deciduous and evergreen subcategories, shrubs with deciduous and evergreen subcategories, graminoids (grasses), and forbs (herbaceous flowering plants) measured as total cover. Litter and “other” cover are also included as total cover. Additional “types” include water and bare ground, which were measured as top cover. The synthesized_pft_fcover_dd.csv accompanies this dataset with header information.species_pft_checklist.csv is a lookup table containing the translation from a dataset species name to an accepted species name and to a PFT. This table can be used to clarify our species to PFT adjudications, and to aid users in assigning their own PFTs. Any issues found in this checklist should be reported in the Issues tab of our github.survey_unit_information.csv contains auxiliary information about the plots synthesized in this database. It contains useful information for filtering plots of interest based on temporal, geospatial, and contextual information about the plot surveys.flmd.csv contains metadata information about each file in the database.This research was performed as a part of the NGEE Arctic project. The NGEE Arctic project was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

PV Reliability and Resilience in Challenging Climates

Challenging climates for Photovoltaics are usually based on climate classification. However, extreme weather events such as high wind, flooding, large hail, extreme snow etc. have become more ubiquitous globally. To study the impact of extraordinary weather events on PV reliability we used two of the largest databases in the USA. First, the National Oceanic and Atmospheric Administration (NOAA) database on extreme weather and secondly, the PV Fleet Data Initiative where we have collected high-resolution PV performance data of more than 8 gigawatts or about 6-7% of all commercial and utility systems in the USA. We analyzed almost 200 systems between 2008-20022 that were immediately impacted by these weather events. The immediate impact (outages) was determined to be about 1% of or a median of approximately 3 days of annual lost production. However, the risk these events pose is exemplified by a long tail where 0.4 % of all systems lost more than 2 weeks annual production. We also found a threshold for high wind (90 km/hr) and hail (25mm), above which we observed significantly higher degradation implying long-term damage to the systems. In addition, we are using satellite imagery to quantify visible damage to PV plants. Finally, we share module, design and installation lessons from some observed case studies to improve extreme weather resilience for PV power systems.

degradation↗