Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

Long Term Cloud Property Datasets From MODIS and AVHRR Using the CERES Cloud Algorithm

Cloud properties play a critical role in climate change. Monitoring cloud properties over long time periods is needed to detect changes and to validate and constrain models. The Clouds and the Earth's Radiant Energy System (CERES) project has developed several cloud datasets from Aqua and Terra MODIS data to better interpret broadband radiation measurements and improve understanding of the role of clouds in the radiation budget. The algorithms applied to MODIS data have been adapted to utilize various combinations of channels on the Advanced Very High Resolution Radiometer (AVHRR) on the long-term time series of NOAA and MetOp satellites to provide a new cloud climate data record. These datasets can be useful for a variety of studies. This paper presents results of the MODIS and AVHRR analyses covering the period from 1980-2014. Validation and comparisons with other datasets are also given.

Minnis, Patrick↗

Validation of the DESI DR2 Ly$α$ forest full-shape analysis

We present the validation of the Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) Lyman-$α$ (Ly$α$) forest full-shape analysis. This analysis combines three-dimensional Ly$α$ forest auto-correlations and cross-correlations with quasars to extract information from both the baryon acoustic oscillation (BAO) feature and the broadband clustering signal, with primary emphasis on the Alcock-Paczynski (AP) measurement. Compared to the DESI DR1 analysis, the DR2 validation uses substantially larger and more realistic mock datasets, including CoLoRe 2LPT and AbacusSummit Ly$α$ forest simulations. The modeling framework is also improved through analytic marginalization over small scales ($<10$$h^{-1}$Mpc) and the impact of ultraviolet background fluctuations. The validation program was completed prior to unblinding and defines quantitative requirements for the cosmological parameters of interest, which are evaluated using hundreds of mock realizations. We further test the analysis through independent fits to the auto- and cross-correlations, multiple catalog splits, and a broad suite of analysis and modeling variations applied to both mocks and blinded observational data. We find that the BAO and AP parameters satisfy all validation requirements and remain stable across all tests. In contrast, mock studies reveal a significant bias in the inferred growth-rate parameter $fσ_8$, leading us to exclude this measurement from the final analysis. The consistency across mocks, data splits, and robustness tests demonstrates that the DR2 Ly$α$ full-shape analysis provides a reliable and substantially improved broadband AP measurement over previous Ly$α$ forest studies.

Herbold, M. [Chicago U., KICP; Ohio State U.] (ORC↗

Implications of stop-and-go traffic on training learning-based car-following control

Learning-based car-following control (LCC) of connected and autonomous vehicles (CAVs) is gaining significant attention with the advancement of computing power and data accessibility. While the flexibility and large model capacity of model-free architecture enable LCC to potentially outperform the model-based car-following (CF) model in improving traffic efficiency and mitigating congestion, the generalizability of LCC for traffic conditions different from the training environment/dataset is not well-understood. Herein, this study seeks to explore the impact of stop-and-go traffic in the training dataset on the generalizability of LCC. It uses the characteristics of lead vehicle trajectories to describe stop-and-go traffic, and links the theory of identifiability (i.e., obtaining a unique parameter estimation result using sensor measurements) to the generalizability of behavior cloning (BC) and policy-based deep reinforcement learning (DRL). Correspondingly, the study shows theoretically that: (i) stop-and-go traffic can enable the property of identifiability and enhance the control performance of BC-based LCC in different traffic conditions; (ii) stop-and-go traffic is not necessary for DRL-based LCC to generalize to different traffic conditions; (iii) DRL-based LCC trained with only constant-speed lead vehicle trajectories (not sufficient to ensure identifiability) can be generalized to different traffic conditions; and (iv) stop-and-go traffic increases variance in the training dataset, which improves the convergence of parameter estimation while negatively impacting the convergence of DRL to the optimal control policy. Numerical experiments validate the above findings, illustrating that BC-based LCC entails comprehensive training datasets for generalizing to different traffic conditions, while DRL-based LCC can achieve generalization with simple free-flow traffic training environments. This further suggests DRL as a more promising and cost-effective LCC approach to reduce operational costs, mitigate traffic congestion, and enhance safety and mobility, which can accelerate the deployment and acceptance of CAVs.

33 ADVANCED PROPULSION SYSTEMS↗

Simulation and Modeling of Hypersonic Turbulent Boundary Layers Subject to Adverse Pressure Gradients due to Concave Streamline Curvature

Direct numerical simulations (DNS) of adverse-pressure-gradient turbulent boundary layers over a planar concave wall are presented for a nominal freestream Mach number of 5, with the objective of assessing the limitations of the currently available Reynolds-averaged Navier-Stokes (RANS) models. The wall geometry and flow conditions of the DNS are representative of the experimental data for a Mach 4.9 turbulent boundary layer that was tested on a two-dimensional planar concave wall model in the high-speed blow-down wind tunnel located at the National Aerothermochemistry Laboratory at Texas A&M University (TAMU). The DNS was validated against the experimental results of TAMU for the same flow conditions and wall geometry. An analysis of the DNS datasets was also conducted to provide an assessment of the validity of Morkovin’s hypothesis and the strong Reynolds analog for turbulence subject to mechanical nonequilibrium. In addition to the DNS results, RANS predictions are obtained by using the Baldwin-Lomax (BL), Spalart-Allmaras (SA), and the k - w SST turbulence models. The comparisons between RANS and DNS showed little impact of an adverse pressure gradient on the accuracy of these models, at least up to an incompressible Clauser pressure gradient parameter of beta(sub inc) 1.22. While the Boussinesq assumption provided reasonable predictions for the Reynolds shear stress, it failed to adequately predict the normal components of the Reynolds stress.

turbulent boundary layers↗

Demonstration of gold nanorod systems for enhanced total efficiency: Experimental and numerical analysis

This study investigates the photothermal performance of gold nanorods engineered to exhibit longitudinal plasmon resonances at 695 nm, 780 nm, and 970 nm. The work combines synthesis, structural characterization, extinction measurements, numerical modeling, and controlled temperature experiments to quantify how nanorod geometry, resonance tuning, concentration, and chamber shape jointly influence heat generation. Transmission electron microscopy confirms that increasing nanorod aspect ratio systematically shifts the longitudinal plasmon peak toward the near-infrared region. Extinction measurements show strong agreement with theoretical predictions, validating the numerical model across two independent datasets. Three chamber geometries were tested under laser excitation at 640 nm, 808 nm, and 980 nm: an ascending stepped base, a flat base, and a descending stepped base. Without nanorods, the ascending geometry produced the highest efficiency due to enhanced natural convection. After introducing gold nanorods, all geometries exhibited substantial thermal enhancement, with total efficiencies exceeding 20%. The strongest improvement was obtained for nanorods resonant at 780 nm with a mass concentration of 4.6 mg/mL implemented on the descending stepped-base geometry. This performance resulted from the combined effect of spectral overlapping with the 808 nm laser, the highest nanorod concentration, and localized heat accumulation that intensified buoyancy-driven flow. The findings demonstrate that total efficiency is governed by a synergistic interplay between optical resonance, nanoparticle concentration, and macroscopic chamber design, revealing the system-level coupling between nanoscale plasmonic absorption and macroscale heat-transfer phenomena. The results provide a validated framework for tuning nanoscale plasmonic absorbers and optimizing thermal systems for applications requiring efficient light-to-heat conversion.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

The Zooplankton International Geospatial (ZIG) dataset: A global repository of spatiotemporal freshwater zooplankton community composition data to support ecological research

Zooplankton play critical roles in aquatic ecosystem function and food webs. Nevertheless, global syntheses of their abundance and community dynamics are challenging due to methodological differences across monitoring programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we assembled, curated, validated, and harmonized the Zooplankton International Geospatial (ZIG) dataset, which includes co-located and contemporaneous zooplankton, water chemistry, and limnological data from 307 lakes and reservoirs. ZIG includes waterbodies from each major lake thermal region and range in size from 0.8-2,805,8600 hectares. Temporal coverage for individual waterbodies ranges between 1-60 years of data (median = 4 years) with sampling from once annually to weekly. ZIG is publicly available and can be used to understand freshwater biodiversity change and its drivers at unprecedented scales, and we consider it to be a cornerstone for future investigations of freshwater biology, chemistry, and ecology.

Figary, Stephanie [Cornell University, Ithaca, NY]↗

Improved Estimates of Pentad Precipitation through the Merging of Independent Precipitation Datasets

Three independent, quasi-global, gridded datasets of precipitation (a rain gauge-based dataset, the satellite-only component of the NASA Integrated Multi-satellitE Retrievals for Global Precipitation Measurement mission [IMERG] Final Run precipitation product, and precipitation estimates derived from NASA Soil Moisture Active Passive [SMAP] soil moisture retrievals), are objectively combined into a single pentad precipitation dataset at 36-km resolution using a unique approach based on extended triple collocation. The quality of each of the four datasets is then evaluated against independent observations. When a global land surface model at 36-km resolution is integrated four times, once utilizing the merged precipitation forcing and once with each of the three contributing datasets, the near-surface soil moisture variations produced with the merged forcing validate best against independent satellite-based soil moisture fields. In addition, the merged dataset is found to be more consistent, relative to each contributor, with estimates of air temperature variations across the globe. The merged dataset thus appears to draw successfully on the complementary strengths of each contributor: the particularly high quality of the rain gauge-based dataset in areas of high gauge density, the more uniform accuracy across the globe of the IMERG data, and the moderate accuracy, particularly in semi-arid regions, of the soil moisture retrieval-based data. Plain Language Summary Obtaining measurements of precipitation across the globe can be challenging. Rain gauges in some ways provide the most accurate measurements, but gauges are absent in many parts of the world, and even where they exist, they only measure precipitation at the gauge itself and therefore may not provide an accurate large-scale average. Satellite-based estimates of precipitation largely overcome these problems, but such data have their own issues, notably a “snapshot” (rather than a time-average) character of the measurements and difficulty associated with interpreting the measured radiances in the presence of complex land surfaces. In the present paper, we use a novel approach to generate a “merged” dataset, one that optimally combines the gauge precipitation information and the satellite-based precipitation information with a third set of estimates derived from soil moisture retrievals. The merged precipitation dataset and each of the three contributors (aggregated here to 5-day averages at a spatial resolution of about 36-km) are then evaluated for consistency with independent geophysical fields. The merged dataset is found to perform best, a clear indication that it takes proper advantage of the complementary strengths of each contributor and, accordingly, that the presented approach for merging the different contributors is indeed viable.

Precipitation↗

Datasets and U-Net Model for "A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: a Case Study for California and Oklahoma"

This dataset has results and the model associated with the publication Ciulla et al., (2024). It contains a U-Net semantic segmentation model (unet_model.h5) and associated code implemented in tensorflow 2.0 for the model training and identification of oil and gas well symbols in USGS historical topographic maps (HTMC). Given a quadrangle map (7.5 minutes), downloadable at this url: https://ngmdb.usgs.gov/topoview/, and a list of coordinates of the documented wells present in the area, the model returns the coordinates of oil and gas symbols in the HTMC maps. For reproducibility of our workflow, we provide a sample map in California and the documented well locations for the entire State of California (CalGEM_AllWells_20231128.csv) downloaded from https://www.conservation.ca.gov/calgem/maps/Pages/GISMapping2.aspx. Additionally, the locations of 1,301 potential undocumented orphaned wells identified using our deep learning framework or the counties of Los Angeles and Kern in California, and Osage and Oklahoma in Oklahoma are provided in the file found_potential_UOWs.zip. The results of the visual inspection of satellite imagery in Osage County is in the file visible_potential_UOWs.zip. The dataset also includes a custom tool to validate the detected symbols in the HTMC maps (vetting_tool.py). More details about the methodology can be found in the associated paper: Ciulla, F., Santos, A., Jordan, P., Kneafsey, T., Biraud, S.C., and Varadharajan, C. (2024) A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: a Case Study for California and Oklahoma. Accepted for publication in Environmental Science and Technology. The geographical coordinates provided correspond to the locations of potential undocumented orphaned oil and gas wells (UOWs) extracted from historical maps. The actual presence of wells need to be confirmed with on-the-ground investigations. For your safety, do not attempt to visit or investigate these sites without appropriate safety training, proper equipment, and authorization from local authorities. Approaching these well sites without proper personal protective equipment (PPE) may pose significant health and safety risks. Oil and gas wells can emit hazardous gasses including methane, which is flammable, odorless and colorless, as well as hydrogen sulfide, which can be fatal even at low concentrations. Additionally, there may be unstable ground near the wellhead that may collapse around the wellbore. This dataset was prepared as an account of work sponsored by the United States Government. While this document is believed to contain correct information, neither the United States Government nor any agency thereof, nor the Regents of the University of California, nor any of their employees, makes any warranty, express or implied, or assumes any legal responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by its trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or the Regents of the University of California. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof or the Regents of the University of California.

Artificial Intelligence↗

Burst pressure models and validations for thick-walled pipelines containing corrosion defects

Corrosion is one major threat to pipeline integrity. Over the past decades, many corrosion models have been developed for determining the remaining strength of corroded pipelines, including ASME B31.G, Modified B31.G, LPC, PCORRC and their modified models. All these corrosion models are applicable only to large diameter, thin-walled pipelines with a diameter to wall thickness ratio D/t ≥ 20. In practice, many pipelines have a small diameter and thick wall with a D/t ratio < 20, and thus an adequate corrosion model is needed for assessing remaining strength for corroded thick-walled pipelines. This paper briefly reviews the theoretical burst pressure models for defect-free thin and thick-walled pipelines and four representative corrosion assessment models for thin-walled corroded pipelines. On this basis, two modified corrosion models are proposed to thick-walled pipelines in terms of the average shear stress yield theory. To verify the proposed corrosion models, comprehensive validations are performed. Numerical validations include the elastic-plastic finite element analysis to determine burst pressure for pipelines without and with corrosion defects and the model evaluation using a large dataset of available FEA results of burst pressure for machined defects. Experimental validations include a set of burst pressure tests for defect-free thick-walled pipes with different thicknesses and the model evaluation using one large burst dataset for machined defects with flat bottoms and another large dataset for real corrosion defects with curved river bottom profiles. Both numerical and experimental validations show that the proposed corrosion models can more accurately predict the remaining strength for corroded thin and thick-walled pipelines.

Pipeline↗

EXPERIMENTAL VALIDATION OF THEORETICAL BURST STRENGTH SOLUTION FOR DEFECT-FREE THICK-WALLED PIPES

The burst pressure of line pipes is an important strength property required in pipeline design and integrity management. Historically, the Barlow formula in conjunction with the ultimate tensile stress (UTS) of pipeline steels were utilized to estimate the burst strength of line pipes. However, the Barlow formula did not consider the plastic flow effect for ductile steels and is applicable only to thin-walled pipes. In 2006, the present author proposed a new multiaxial plastic yield theory and obtained a theoretical Zhu-Leis solution of burst strength for defect-free thin-walled pipes in term of UTS and strain hardening exponent n of pipeline steels. The Zhu-Leis solution has been validated by various burst test data for thin-walled pipelines for a wide range of steel grades from Grade B to X120. Recently, the present author extended the Zhu-Leis theory of plasticity to thick-walled pipes and obtained the Zhu-Leis solution of burst pressure for thick-walled pipes. The proposed burst pressure solution is applicable to both thin and thick-walled pipes. To experimentally validate the proposed theoretical burst pressure solution, this paper obtains a set of burst test data for three thick-walled pipes in Grade B carbon steel with a nominal diameter of 2.375 inches and three nominal wall thicknesses, resulting in D/t = 15.4, 10.9, 6.9. Through comparisons, these burst data validate the theoretical burst pressure solution for thick-walled pipes. Moreover, two additional burst test datasets collected from literature for thin and thick-walled pipes further validate the proposed burst pressure solution for both thin and thick-walled pipes.

Zhu, Xiankui↗

WTK-LED: The WIND Toolkit Long-Term Ensemble Dataset

To satisfy a wide group of stakeholders across various wind energy disciplines, including but not limited to stakeholders in the distributed and utility scale wind industry, the new emerging airborne wind energy field, grid integration, power systems modeling, environmental modeling, and researchers in academia, and to close some of the gaps that current public datasets have, we aimed at developing an updated version of the meteorological WIND Toolkit, named WIND Toolkit Long-term Ensemble Dataset (WTK-LED), which is a meteorological dataset providing time series every 5 min and 2 km, including model uncertainty of wind speed at every modeling grid point so that users are provided with a range of possible wind speeds every 2 km. The data were produced using the Weather Research and Forecasting Model (WRF). The vertical grid used in WTK-LED includes many vertical layers in the atmospheric boundary layer to provide information of atmospheric quantities across the rotor layer of utility scale and distributed wind turbines. The WTK-LED includes: 1) Numerical simulations covering the continental United States, Alaska, and Hawaii, with high-resolution data being available for 3 years (2018-2020). 2) Climate simulations from Argonne National Laboratories covering the North American continent, including Alaska, Canada, and most of Mexico and the Caribbean Islands. These simulations complement the new WTK-LED to offer a 4-km dataset covering 20 years, from 2001-2020. 3) Specific long-term,high-resolution offshore simulations have been conducted separately for the US coasts, Hawaii, and the Great Lakes, leading to the 2023 National Offshore Wind data set. This report focuses on a description of the land-based WTK-LED for CONUS, Hawaii, and Alaska, for the 3-year 2-km/5-min dataset and the 20-year 4-km/hourly dataset, as well as the uncertainty quantification method. We also provide limited validation results. Based on our results to date, we suggest use cases and applications for each dataset of the WTK-LED.

17 WIND ENERGY↗

Application of Machine Learning and Data Augmentation Algorithms in the Discovery of Metal Hydrides for Hydrogen Storage

The development of efficient and sustainable hydrogen storage materials is a key challenge for realizing hydrogen as a clean and flexible energy carrier. Among various options, metal hydrides offer high volumetric storage density and operational safety, yet their application is limited by thermodynamic, kinetic, and compositional constraints. In this work, we investigate the potential of machine learning (ML) to predict key thermodynamic properties—equilibrium plateau pressure, enthalpy, and entropy of hydride formation—based solely on alloy composition using Magpie-generated descriptors. We significantly expand an existing experimental dataset from ~400 to 806 entries and assess the impact of dataset size and data augmentation, using the PADRE algorithm, on model performance. Models including Support Vector Machines and Gradient Boosted Random Forests were trained and optimized via grid search and cross-validation. Results show a marked improvement in predictive accuracy with increased dataset size, while data augmentation benefits are limited to smaller datasets and do not improve accuracy in underrepresented pressure regimes. Furthermore, clustering and cross-validation analyses highlight the limited generalizability of models across different material classes, though high accuracy is achieved when training and testing within a single hydride family (e.g., AB2). The study demonstrates the viability and limitations of ML for accelerating hydride discovery, emphasizing the importance of dataset diversity and representation for robust property prediction.

augmentation↗

Characterizing in-stream turbulent flow for tidal energy converter siting in Cook Inlet, Alaska

Cook Inlet in Alaska is the most promising location for tidal energy development in the U.S. due to its significant tidal range of approximately 10 meters and high volume flux. The inlet's unique geometry and flow characteristics make it the most energetic tidal stream in the nation, with GW-scale potential energy capacity. With the growing interest in tidal energy converter (TEC) deployment in this area, we implemented a regional-scale, 3D hydrodynamic modeling framework to predict tidal current and turbulence characteristics that can assist TEC designers and project managers. We validated the model results extensively using various datasets collected with bottom-mounted acoustic Doppler current profilers and velocimeters. The comparison between the model outputs and observational data highlighted the effectiveness of the 3D FVCOM model and the Mellor-Yamada Level 2.5 Turbulence Model in accurately assessing macro-scale kinetic energy, turbulence intensity, and the production and dissipation rates at a prospective TEC site. Using two months of model simulation data, we examined the channel cross-section for TEC deployment, focusing on undisturbed power density and macro-scale turbulent properties. Further, our findings indicate that understanding the turbulence characteristics and flow properties can enhance Stage I/II resource characterization by identifying optimal locations for TECs and their layouts within the channel. Furthermore, we demonstrated that TEC designers can utilize macro-scale turbulence data from 3D coastal models as boundary conditions for other turbulence models, allowing for a more detailed resolution of the turbulence structure at TEC siting locations. Ultimately, this work emphasizes the importance of estimating flow and turbulence conditions in energetic systems to understand turbulent sites better and improve resource characterization.

16 TIDAL AND WAVE POWER↗

DECADE+ DES Y3 Weak Lensing Mass Map: A 13,000 deg $\^{} 2$ View of Cosmic Structure from 270 Million Galaxies

We present the largest galaxy weak lensing mass map of the late-time Universe, reconstructed from 270 million galaxies in the DECADE and DES Year 3 datasets, covering 13,000 square degrees. We validate the map through systematic tests against observational conditions (depth, seeing, etc.), finding the map is statistically consistent with no contamination. The large area covered by the mass map makes it a well-suited tool for cosmological analyses, cross-correlation studies and the identification of large-scale structure features. We demonstrate its potential by detecting cosmic filaments directly from the mass map for the first time and validating them through their association with galaxy clusters selected using the Sunyaev-Zeldovich effect from Planck and ACT DR6.

Gatti, M.↗

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel↗

High-Fidelity Dataset Generation for Sensor Anomalies in Power Grids using Hardware-in-the-Loop Testbed

Sensor anomalies in power grids can have significant impacts on the operation of the grid due to the increased reliance of the grid operation on data-driven applications. However, there is a lack of datasets that accurately capture these anomalies as many of the anomalies go undetected using the current bad data detectors. High-fidelity labeled datasets are essential for developing robust applications that can detect and mitigate the impacts of anomalies. In this paper, we propose a hardware-in-the-loop testbed model that can emulate the grid behavior with high-fidelity. This testbed is used to inject anomalies at various levels in the grid architecture and generate labeled datasets. These high-fidelity datasets can be used for development and validation of data-driven applications for detection and mitigation of anomalies in grids and other cyber-physical systems.

Hyder, Burhan↗

TCR-H: explainable machine learning prediction of T-cell receptor epitope binding on unseen datasets

Artificial-intelligence and machine-learning (AI/ML) approaches to predicting T-cell receptor (TCR)-epitope specificity achieve high performance metrics on test datasets which include sequences that are also part of the training set but fail to generalize to test sets consisting of epitopes and TCRs that are absent from the training set, i.e., are ‘unseen’ during training of the ML model. We present TCR-H, a supervised classification Support Vector Machines model using physicochemical features trained on the largest dataset available to date using only experimentally validated non-binders as negative datapoints. TCR-H exhibits an area under the curve of the receiver-operator characteristic (AUC of ROC) of 0.87 for epitope ‘hard splitting’ (i.e., on test sets with all epitopes unseen during ML training), 0.92 for TCR hard splitting and 0.89 for ‘strict splitting’ in which neither the epitopes nor the TCRs in the test set are seen in the training data. Furthermore, we employ the SHAP (Shapley additive explanations) eXplainable AI (XAI) method for post hoc interrogation to interpret the models trained with different hard splits, shedding light on the key physiochemical features driving model predictions. TCR-H thus represents a significant step towards general applicability and explainability of epitope:TCR specificity prediction.

60 APPLIED LIFE SCIENCES↗