Search NASA⌕ Search

SEARCH · Search NASA

Results for “accurate phase labels”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Phase identification using co‐association matrix ensemble clustering

Calibrating distribution system models to aid in the accuracy of simulations such as hosting capacity analysis is increasingly important in the pursuit of the goal of integrating more distributed energy resources. The recent availability of smart meter data is enabling the use of machine learning tools to automatically achieve model calibration tasks. This research focuses on applying machine learning to the phase identification task, using a co‐association matrix‐based, ensemble spectral clustering approach. The proposed method leverages voltage time series from smart meters and does not require existing or accurate phase labels. This work demonstrates the success of the proposed method on both synthetic and real data, surpassing the accuracy of other phase identification research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Semiempirical Estimate of Aircraft Wing Weight

Computational method estimates weight of aircraft wings from theoretical relationships and empirical data. Permits comparison of alternative materials, methods of construction, and design philosophies. Method used to make tradeoffs in preliminary design phases on basis of simple input data and for more accurate calculations in later phases when more data are available.

York, P.↗

Deep learning-based segmentation of lithium-ion battery microstructures enhanced by artificially generated electrodes

Accurate 3D representations of lithium-ion battery electrodes, in which the active particles, binder and pore phases are distinguished and labeled, can assist in understanding and ultimately improving battery performance. Here, we demonstrate a methodology for using deep-learning tools to achieve reliable segmentations of volumetric images of electrodes on which standard segmentation approaches fail due to insufficient contrast. We implement the 3D U-Net architecture for segmentation, and, to overcome the limitations of training data obtained experimentally through imaging, we show how synthetic learning data, consisting of realistic artificial electrode structures and their tomographic reconstructions, can be generated and used to enhance network performance. We apply our method to segment x-ray tomographic microscopy images of graphite-silicon composite electrodes and show it is accurate across standard metrics. We then apply it to obtain a statistically meaningful analysis of the microstructural evolution of the carbon-black and binder domain during battery operation.

25 ENERGY STORAGE↗

Emulation of seismic-phase traveltimes with machine learning

SUMMARY We present a machine learning (ML) method for emulating seismic-phase traveltimes that are computed using a global-scale 3-D earth model and physics-based ray tracing. Accurate traveltime predictions based on 3-D earth models are known to reduce the bias of event location estimates, increase our ability to assign phase labels to seismic detections and associate detections to events. However, practical use of 3-D models is challenged by slow computational speed and the unwieldiness of pre-computed lookup tables that are often large and have prescribed computational grids. In this work, we train a ML emulator using pre-computed traveltimes, resulting in a compact and computationally fast way to approximate traveltimes that are based on a 3-D earth model. Our model is trained using approximately 850 million P-wave traveltimes that are based on the global LLNL-G3D-JPS model, which was developed for more accurate event location. The training-set consists of traveltimes between 10 393 global seismic stations and randomly sampled event locations that provide a prescribed, distance-dependent geographic sample density for each station. Prediction accuracy is dependent on event-station distance and whether the station was included in the training set. For stations included in the training set the mean absolute deviation (MAD) of the difference between traveltimes computed using ray tracing through the 3-D model and the ML emulator for local, regional, and teleseismic distances are 0.090, 0.125 and 0.121 s, respectively. For tested station locations not included in the training set, MAD values for the three distance ranges increase to 0.173, 0.219 and 0.210 s, respectively. Empirical traveltime residuals for a global reference data are indistinguishable when ML emulation or the 3-D model is used to compute traveltimes. This result holds regardless of whether the recording station is used in ML training or not.

58 GEOSCIENCES↗

Deep learning classification of lipid droplets in quantitative phase images

We report the application of supervised machine learning to the automated classification of lipid droplets in label-free, quantitative-phase images. By comparing various machine learning methods commonly used in biomedical imaging and remote sensing, we found convolutional neural networks to outperform others, both quantitatively and qualitatively. We describe our imaging approach, all implemented machine learning methods, and their performance with respect to computational efficiency, required training resources, and relative method performance measured across multiple metrics. Overall, our results indicate that quantitative-phase imaging coupled to machine learning enables accurate lipid droplet classification in single living cells. As such, the present paradigm presents an excellent alternative of the more common fluorescent and Raman imaging modalities by enabling label-free, ultra-low phototoxicity, and deeper insight into the thermodynamics of metabolism of single cells.

59 BASIC BIOLOGICAL SCIENCES↗

Ice Phase Classification Made Easy with Score-Based Denoising

Accurate identification of ice phases is essential for understanding various physicochemical phenomena. However, such classification for structures simulated with molecular dynamics is complicated by the complex symmetries of ice polymorphs and thermal fluctuations. For this purpose, both traditional order parameters and data-driven machine learning approaches have been employed, but they often rely on expert intuition, specific geometric information, or large training data sets. In this work, we present an unsupervised phase classification framework that combines a score-based denoiser model with a subsequent model-free classification method to accurately identify ice phases. Further, the denoiser model is trained on perturbed synthetic data of ideal reference structures, eliminating the need for large data sets and labeling efforts. The classification step utilizes the smooth overlap of atomic position (SOAP) descriptors as the atomic fingerprint, ensuring Euclidean symmetries and transferability to various structural systems. Our approach achieves a remarkable 100% accuracy in distinguishing ice phases of test trajectories using only seven ideal reference structures of ice phases as model inputs. This demonstrates the generalizability of the score-based denoiser model in facilitating phase identification for complex molecular systems. The proposed classification strategy can be broadly applied to investigate structural evolution and phase identification for a wide range of materials, offering new insights into the fundamental understanding of water and other complex systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accurate and Data‐Efficient Micro X‐ray Diffraction Phase Identification Using Multitask Learning: Application to Hydrothermal Fluids

Traditional analysis of highly distorted micro X‐ray diffraction (μ‐XRD) patterns from hydrothermal fluid environments is a time‐consuming process, often requiring substantial data preprocessing and labeled experimental data. Herein, the potential of deep learning with a multitask learning (MTL) architecture to overcome these limitations is demonstrated. MTL models are trained to identify phase information in μ‐XRD patterns, minimizing the need for labeled experimental data and masking preprocessing steps. Notably, MTL models show superior accuracy compared to binary classification convolutional neural networks. Additionally, introducing a tailored cross‐entropy loss function improves MTL model performance. Most significantly, MTL models tuned to analyze raw and unmasked XRD patterns achieve close performance to models analyzing preprocessed data, with minimal accuracy differences. This work indicates that advanced deep learning architectures like MTL can automate arduous data handling tasks, streamline the analysis of distorted XRD patterns, and reduce the reliance on labor‐intensive experimental datasets.

97 MATHEMATICS AND COMPUTING↗

Dynamics of cell proliferation in the adult dentate gyrus of two inbred strains of mice

The output potential of proliferating populations in either the developing or the adult nervous system is critically dependent on the length of the cell cycle (T(c)) and the size of the proliferating population. We developed a new approach for analyzing the cell cycle, the 'Saturate and Survive Method' (SSM), that also reveals the dynamic behaviors in the proliferative population and estimates of the size of the proliferating population. We used this method to analyze the proliferating population of the adult dentate gyrus in 60 day old mice of two inbred strains, C57BL/6J and BALB/cByJ. The results show that the number of cells labeled by exposure to BUdR changes dramatically with time as a function of the number of proliferating cells in the population, the length of the S-phase, cell division, the length of the cell cycle, dilution of the S-phase label, and cell death. The major difference between C57BL/6J and BALB/cByJ mice is the size of the proliferating population, which differs by a factor of two; the lengths of the cell cycle and the S-phase and the probability that a newly produced cell will die within the first 10 days do not differ in these two strains. This indicates that genetic regulation of the size of the proliferating population is independent of the genetic regulation of cell death among those newly produced cells. The dynamic changes in the number of labeled cells as revealed by the SSM protocol also indicate that neither single nor repeated daily injections of BUdR accurately measure 'proliferation.'.

Non-NASA Center↗

An Open Combinatorial Diffraction Dataset Including Consensus Human and Machine Learning Labels with Quantified Uncertainty for Training New Machine Learning Models

Modern machine learning and autonomous experimentation schemes in materials science rely on accurate analysis of the data ingested by these models. Unfortunately, accurate analysis of the underlying data can be difficult, even for domain experts, complicating the training of the models intended to drive experiments. This is especially true when the goal is to identify the presence of weak signatures in diffraction or spectroscopic datasets. In this work, we examine a set of as-obtained diffraction data that track the phase transition from monoclinic to tetragonal in a Nb-doped VO2 film as a function of temperature and dopant concentration. We then task a set of domain experts and a set of machine learning experts with identifying which phase is present in each diffraction pattern manually and algorithmically, respectively; in both cases, the labels can vary dramatically, especially at the phase boundaries. We use the mode of the labels and the Shannon entropy as a method to capture, preserve and propagate consensus labels and their variance. Further we use the expert labels as a benchmark and demonstrate the use of Shannon entropy weighted scoring to test the performance of machine learning generated labels. Finally, we propose a material data challenge centered around generating improved labeling algorithms. This real-world dataset curated with expert labels can act as test bed for new algorithms. The raw data, annotations and code used in this study are all available online at data.gov and the interested reader is encouraged to replicate and improve the existing models

97 MATHEMATICS AND COMPUTING↗

Multi-Class Anomaly Detection in Flight Data using Semi-Supervised Explainable Deep Learning Model

Identifying precursor for safety incidents in aviation data is a crucial task, yet extremely challenging. The main approach, in practice, leverages domain expertise to define expected tolerances in system’s behavior and alarm exceedance from such safety margins. However, this approach is incapable of identifying unknown risk and vulnerabilities. Machine learning has been long studied and deployed to identify precursors for such anomalies, with the great challenge of the need for sufficient labelled set of data to achieve a reliable and accurate performance. In this article, we develop an explainable deep semi-supervised model for anomaly detection in aviation, building upon recent advancements in the machine learning literature. The proposed model combines feature engineering and classification in the feature space, while leveraging all available data (labelled and unlabeled). Validating on two case studies of anomaly detection in take-off and landing phases of commercial aircraft, we show that our model is able to outperform state-of-the-art supervised anomaly detection model and reach significantly high accuracy and low false alarm with minimum amount of available labelled data.

Anomaly Detection↗

Large-scale physically accurate modelling of real proton exchange membrane fuel cell with deep learning

Proton exchange membrane fuel cells, consuming hydrogen and oxygen to generate clean electricity and water, suffer acute liquid water challenges. Accurate liquid water modelling is inherently challenging due to the multi-phase, multi-component, reactive dynamics within multi-scale, multi-layered porous media. In addition, currently inadequate imaging and modelling capabilities are limiting simulations to small areas (<1 mm 2 ) or simplified architectures. Herein, an advancement in water modelling is achieved using X-ray micro-computed tomography, deep learned super-resolution, multi-label segmentation, and direct multi-phase simulation. The resulting image is the most resolved domain (16 mm 2 with 700 nm voxel resolution) and the largest direct multi-phase flow simulation of a fuel cell. This generalisable approach unveils multi-scale water clustering and transport mechanisms over large dry and flooded areas in the gas diffusion layer and flow fields, paving the way for next generation proton exchange membrane fuel cells with optimised structures and wettabilities.

25 ENERGY STORAGE↗

PHASE: Personalized Head-based Automatic Simulation for Electromagnetic properties in 7T MRI

Accurate and individualized human head models are becoming increasingly important for electromagnetic (EM) simulations. These simulations depend on precise anatomical representations to realistically model electric and magnetic field distributions, particularly when evaluating Specific Absorption Rate (SAR) within safety guidelines. State of the art simulations use the Virtual Population due to limited public resources and the impracticality of manually annotating patient data at scale. Here, this paper introduces Personalized Head-based Automatic Simulation for EM properties (PHASE), an automated open-source toolbox that generates high-resolution, patient-specific head models for EM simulations using paired T1-weighted (T1w) magnetic resonance imaging (MRI) and computed tomography (CT) scans with 14 tissue labels. To evaluate the performance of PHASE models, we conduct semi-automated segmentation and EM simulations on 15 real human patients, serving as the gold standard reference. The PHASE model achieved comparable global SAR and localized SAR averaged over 10 grams of tissue (SAR-10g), demonstrating its potential as a promising tool for generating large-scale human model datasets in the future. The code and models of PHASE toolbox have been made publicly available: https://github.com/hrlblab/PHASE.

Deep learning↗

Intercomparison of flood inundation models across land use types and hydrological flood stages

Flood Inundation Mapping (FIM) model selection is a key operational decision because accurate, rapid mapping underpins early warning and resource allocation. FIM performance is context-dependent and can vary with hydrograph phase, land-use/land-cover (LULC), and the evaluation benchmark. Intercomparison studies typically assess a single near-peak snapshot against one reference dataset. Here, we provide a context-stratified intercomparison across (i) multiple hydrograph phases, (ii) LULC classes, and (iii) benchmark types, for five FIM approaches spanning a wide range of physical complexity and operational cost (TRITON, LISFLOOD-FP, HEC-RAS 2D, ARC-Curve2Flood, and OWP HAND-FIM). We use the Hurricane Matthew flood (2016) in the Neuse River Basin, North Carolina, USA, as a case study. Using high-resolution remote sensing-derived flood inundation maps, hand-labeled points, and building footprints, we assess model skill across two rising and two falling hydrograph limbs and across major LULC types. Results show that model rankings shift systematically across contexts: LISFLOOD-FP ranks highest in three of four flood phases, while TRITON leads during one rising limb phase; LISFLOOD-FP performs best in vegetated areas, whereas HEC-RAS improves relative performance in agricultural and urban areas; and benchmark choice influences conclusions, with LISFLOOD-FP performing best for flooded-building detection in the late falling limb, while TRITON ranks highest against hand-labeled points. We also report representative wall-clock runtimes for each workflow to provide use-case context for operational feasibility. Together, these results offer transferable guidance for model selection and for designing large-scale, benchmark-aware FIM intercomparison studies.

Nikrou, Parvaneh [University of Alabama]↗

Optimal surface-tension isotropy in the Rothman-Keller color-gradient lattice Boltzmann method for multiphase flow

The Rothman-Keller color-gradient (CG) lattice Boltzmann method is a popular method to simulate two-phase flow because of its ability to deal with fluids with large viscosity contrasts and a wide range of interfacial tensions. Here, two fluids are labeled red and blue, and the gradient in the color difference is used to compute the effect of interfacial tension. It is well known that finite-difference errors in the color-gradient calculation lead to anisotropy of interfacial tension and errors such as spurious currents. Here, we investigate the accuracy of the CG calculation for interfaces between fluids with several radii of curvature and find that the standard CG calculations lead to significant inaccuracy. Specifically, we observe significant anisotropy of the color gradient of order 7% for high curvature of an interface such as when a pinchout occurs. We derive a second order accurate color gradient and find that the diagonal nearest neighbors can be weighted differently than in the usual color-gradient calculation such that anisotropy is minimized to a fraction of a percent. The optimal weights that minimize anisotropy for the smallest radius of curvature interface are found to be w = (0.298, 0.284, 0.275) for diagonal nearest neighbors for the cases of the interface smoothing parameter β = (0.5, 0.7, 0.99), somewhat higher than the w = 0.25 value derived by Leclaire et al. [Leclaire, Reggio, and Trepanier, Computers and Fluids 48, 98 (2011)] based on obtaining isotropic errors to second order. We find that use of these optimal w values yields over a factor of 10 decrease in anisotropy and over a factor of 30 decrease in mean anisotropy relative to using the standard w = 1 value. And we find a factor of about 2 decrease in the anisotropic error and up to factor 15 decrease in mean anisotropic error relative to the choice of w = 0.25 for small radius of curvature interfaces. The improved CG calculations will allow the method to be more reliably applied to studies of phenomenology and pore scale processes such as viscous and capillary fingering, and droplet formation where surface-tension isotropy of narrow fingers and small droplets plays a crucial role in correctly capturing phenomenology. We present an example illustrating how different phenomena can be captured using the improved color-gradient method. Namely, we present simulations of a wetting fluid invading a fluid filled pipe where the viscosity ratio of fluids is unity in which droplets form at the transition to fingering using the improved CG calculations that are not captured using the standard CG calculations. We present an explanation of why this is so which relates to anisotropy of the surface tension, which inhibits the pinchouts needed to form droplets.

58 GEOSCIENCES↗

Machine Learning Automated Analysis of Enormous Synchrotron X-ray Diffraction Datasets

X-ray diffraction (XRD) data analysis can be a time-consuming and laborious task. Deep neural network (DNN) based models trained with synthetic XRD patterns have been proven to be a highly efficient, accurate, and automated method for analyzing common XRD data collected from solid samples in ambient environments. However, it remains unclear whether synthetic XRD-based models can be effective in solving micro(μ)-XRD mapping data for in situ experiments involving liquid phases, which always have lower quality and significant artifacts. In this study, we collected μ-XRD mapping data from a LaCl 3 -calcite hydrothermal fluid system and trained two categories of models to analyze the experimental XRD patterns. Here, the models trained solely with synthetic XRD patterns showed low accuracy (as low as 64%) when solving experimental μ-XRD mapping data. However, the accuracy of the DNN models significantly improved (90% or above) when we trained them with a data set containing both synthetic and a small number of labeled experimental μ-XRD patterns. This study highlights the importance of labeled experimental patterns in training DNN models to solve μ-XRD mapping data from in situ experiments involving liquid phases.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Synthesis, Processing, and Use of Isotopically Enriched Epitaxial Oxide Thin Films

Isotopic engineering has emerged as a key approach to study the nucleation, diffusion, phase transitions, and reactions of materials at an atomic level. It aims to uncover mass transport pathways, kinetics, and operational and failure mechanisms of functional materials and devices. Understanding these phenomena leads to deeper insights into important physical processes, such as the transport of ions in energy conversion and storage devices and the role of active sites and supports during heterogeneous catalytic reactions. Likewise, isotopic engineering is being pursued as a means of modifying functionality to enable future technological applications. In this Account, we summarize our recent work employing isotope labeling (e.g., 18 O 2 and 57 Fe) during thin film synthesis and postgrowth processing to reveal growth mechanisms, defect chemistry, and elemental diffusion under working and extreme conditions. Isotope-resolved analysis techniques with nanometer-scale spatial resolution, such as time-of-flight secondary ion mass spectrometry and atom probe tomography, facilitate the accurate quantification of isotopic placement and concentration in our well-defined heterostructures with precisely positioned, isotope-enriched layers. By measuring the nanometer-scale redistribution between natural abundance and isotopically enriched oxygen layers during the deposition of Fe 2 O 3 and Cr 2 O 3 by molecular beam epitaxy, we identified intermixing processes driven by surface adatoms occurring both at the film growth surface and within the first few layers below the surface. Further insights into synthesis mechanisms were gained by studying the tungsten oxide thin films grown by evaporating WO 3 powder in the presence of background 18 O 2 , revealing minimal incorporation of background oxygen during the film formation process. Thermal and radiation-enhanced diffusion in epitaxial Fe and Cr oxides were precisely tracked using 18 O and 57 Fe tracer layers incorporated into model epitaxial oxide thin films. This approach has allowed us to access thermal diffusion behavior at lower temperatures than previously measured, revealing a potential changeover in diffusion mechanism. Understanding radiation-enhanced diffusion in model oxides that represent the surface layers on the structural components of nuclear reactors informs our understanding of their corrosion behavior under irradiation. Isotopic labeling can also provide unique insights into the surface exchange reactions and defect chemistry of electrocatalysts. For instance, tracking the change in 18 O concentration at the surface of an epitaxial LaNiO 3 thin film after the electrocatalytic oxygen evolution reaction revealed the participation of lattice oxygen, confirming a hypothesis that had been proposed previously. Lastly, we highlight a new direction wherein we perform in situ processing studies utilizing isotopic tracers in conjunction with model epitaxial thin films within the atom probe tomography instrument. Additionally, this Account illustrates the great potential of isotopic engineering to enable fundamental mechanistic insights into physical processes and engineer functional properties in epitaxial films, heterostructures, and superlattices.

36 MATERIALS SCIENCE↗

A Simple and Sensitive LC-MS/MS Method for the Determination of Free 8-Hydroxy-2'-Deoxyguanosine in Human Urine

Urinary free 8-hydroxy-2'-deoxyguanosine (8OHdG), an oxidized product of DNA, and is frequently chosen as a biomarker of oxidative stress in humans, including studies of oxidative DNA damage during space flight. It is challenging to accurately and efficiently quantify urinary free 8OHdG in large scale human studies. LC-MS/MS is emerging as a preferable analytical technique owing its high sensitivity, selectivity and efficiency, compared to some traditional methods such as ELISA and HPLC. A simple and sensitive LC-MS/MS method has been developed for the determination of free 8OHdG in human urine. Sample preparation was done by solid phase extraction with a Waters Oasis HLB 96 well plate. A Waters Alliance 2795 HT Separation Module combined with a Quattro Micro tandem mass spectrometer was used as the LC-MS/MS system. The runtime of one injection can be less than 5 minutes using a reversed phase C18 column and an isocratic flow of methanol/water. ESI positive ions were quantified in the multiple reaction modes (MRM) using m/z 284 yields 168 for 8OHdG and m/z 289 yields173 for stable isotope labeled internal standard [(15)N5] 8OHdG. With this method for 8OHdG, a lower limit of quantitation of 1.0 nM (0.28 ng/mL) has been achieved using 100 microliter urine sample. The analytical range is between 1.0 and 100 nM with a correlation coefficient greater than or equal to 0.99. Good reproducibility can be obtained with intra-assay and inter-assay CVs less than or equal to 10% for 8OHdG spiked urine QC samples. This method can be used in high-throughput routine analysis of free 8OHdG in human urine.

Wang, Zuwei↗

Water Across Synthetic Aperture Radar Data (WASARD): SAR Water Body Classification for the Open Data Cube

The detection of inland water bodies from Synthetic Aperture Radar (SAR) data provides a great advantage over water detection with optical data, since SAR imaging is not impeded by cloud cover. Traditional methods of detecting water from SAR data involves using thresholding methods that can be labor intensive and imprecise. This paper describes Water Across Synthetic Aperture Radar Data (WASARD): a method of water detection from SAR data which automates and simplifies the thresholding process using machine learning on training data created from Geoscience Australia’s WOFS algorithm. Of the machine learning models tested, the Linear Support Vector Machine was determined to be optimal, with the option of training using solely the VH polarization or a combination of the VH and VV polarizations. WASARD was able to identify water in the target area with a correlation of 97% with WOFS. Sentinel-1, Open Data Cube, Earth Observations, Machine Learning, Water Detection 1. INTRODUCTION Water classification is an important function of Earth imaging satellites, as accurate remote classification of land and water can assist in land use analysis, flood prediction, climate change research, as well as a variety of agricultural applications [2]. The ability to identify bodies of water remotely via satellite is immensely cheaper than contracting surveys of the areas in question, meaning that an application that can accurately use satellite data towards this function can make valuable information available to nations which would not be able to afford it otherwise. Highly reliable applications for the remote detection of water currently exist for use with optical satellite data such as that provided by LANDSAT. One such application, Geoscience Australia’s Water Observations from Space (WOFS) has already been ported for use with the Open Data Cube [6]. However, water detection using optical data from Landsat is constrained by its relatively long revisit cycle of 16 days [5], and water detection using any optical data is constrained in that it lacks the ability to make accurate classifications through cloud cover [2]. The alternative solution which solves these problems is water detection using SAR data, which images the Earth using cloud-penetrating microwaves. Because of its advantages over optical data, much research has been done into water detection using SAR data. Traditionally, this has been done using the thresholding method, which involves picking a polarization band and labeling all pixels for which this band’s value is below a certain threshold as containing water. The thresholding method works since water tends to return a much lower backscatter value to the satellite than land [1]. However, this method can be flawed since estimating the proper threshold is often imprecise, complicated, and labor intensive for the end user. Thresholding also tends to use data from only one SAR polarization, when a combination of polarizations can provide insight into whether water is present. [2] In order to alleviate these problems, this paper presents an application for the Open Data Cube to detect water from SAR data using support vector machine (SVM) classification. 2. PLATFORM WASARD is an application for the Open Data Cube, a mechanism which provides a simple yet efficient means of ingesting, storing, and retrieving remote sensing data. Data can be ingested and made analysis ready according to whatever specifications the researcher chooses, and easily resampled to artificially alter a scene’s resolution. Currently WASARD supports water detection on scenes from ESA’s Sentinel-1 and JAXA’s ALOS. When testing WASARD, Sentinel-1 was most commonly used due to its relatively high spatial resolution and its rapid 6 day revisit cycle [5]. With minor alterations to the application's code, however, it could support data from other satellites. 3. METHODOLOGY Using supervised classification, WASARD compares SAR data to a dataset pre-classified by WOFS in order to train an SVM classifier. This classifier is then used to detect water in other SAR scenes outside the training set. Accuracy was measured according to the following metrics:  Precision: a measure of what percentage of the points WASARD labels as water are truly water  Recall: a measure of what percentage of the total water cover WASARD was able to identify.  F1 Score: a harmonic average of the precision and recall scores Both precision and recall are calculated at the end of the training phase, when the trained classifier is compared to a testing dataset. Because the WOFS algorithm’s classifications are used as the truth values when training a WASARD classifier, when precision and recall are mentioned in this paper, they are always with respect to the values produced by WOFS on a similar scene of Landsat data, which themselves have a classification accuracy of 97% [6]. Visual representations of water identified by WASARD in this paper were produced using the function wasard_plot(), which is included in WASARD. 3.1 Algorithm Selection The machine learning model used by WASARD is the Linear Support Vector Machine (SVM). This model uses a supervised learning algorithm to develop a classifier, meaning it creates a vector which can be multiplied by the vector formed by the relevant data bands to determine whether a pixel in a SAR scene contains water. This classifier is trained by comparing data points from selected bands in a SAR scene to their respective labels, which in this case are “water” or “not water” as given by the WOFS algorithm. The SVM was selected over the Random Forest model, which outperformed the SVM in training speed, but had a greater classification time and lower accuracy, and the Multilayer Perceptron Artificial Neural Network, which had a slightly higher average accuracy than the SVM, but much greater training and classification times. Figure 1: Visual representation of the SVM Classifier. Each white point represents a pixel in a SAR scene. In Figure 1, the diagonal line separating pixels determined to be water from those determined not to be water represents the actual classification vector produced by the SVM. It is worth noting that once the model has been trained, classification of pixels is done in a similar manner as in the thresholding method. This is especially true if only one band was used to train the model. 3.1 Feature Selection Sentinel-1 collects data from two bands: the Vertical/Vertical polarization (VV) and the Vertical/Horizontal polarization (VH). When 100 SVM classifiers were created for each polarization individually, and for the combination of the two, the following results were achieved: Figure 2: Accuracy of classifiers trained using different polarization bands. Precision and Recall were measured with respect to the values produced by WOFS. Figure 2 demonstrates that using both the VV and VH bands trades slightly lower recall for significantly greater precision when compared with the VH band alone, and that using the VV band alone is inferior in both metrics. WASARD therefore defaults to using both the VV and VH bands, and includes the option to use solely the VH band. The VV polarization’s lower precision compared to the VH polarization is in contrast to results from previous research and may merit further analysis [4]. 3.2 Training a Classifier The steps in training a classifier with WASARD are 1. Selecting two scenes (one SAR, one optical) with the same spatial extents, and acquired close to each other in time, with a preference that the scenes are taken on the same day. 2. Using the WOFS algorithm to produce an array of the detected water in the scene of optical data, to be used as the labels during supervised learning 3. Data points from the selected bands from the SAR acquisition are bundled together into an array with the corresponding labels gathered from WOFS. A random sample with an equal number of points labeled “Water” and “Not Water” is selected to be partitioned into a training and a testing dataset 4. Using Scikit-Learn’s LinearSVC object, the training dataset is used to produce a classifier, which is then tested against the testing dataset to determine its precision and recall The result is a wasard_classifier object, which has the following attributes: 1. f1, recall, and precision: 3 metrics used to determine the classifier’s accuracy 2. Coefficient: Vector which the SVM uses to make its predictions. The classifier detects water when the dot product of the coefficient and the vector formed by the SAR bands is positive 3. Save(): allows a user to save a classifier to the disk in order to use it without retraining 4. wasard_classify(): Classifies an entire xarray of SAR data using the SVM classifier All of the above steps are performed automatically when the user creates a wasard_classifier object. 3.3 Classifying a Dataset Once the classifier has been created, it can be used to detect water in an xarray of SAR data using wasard_classify(). By taking the dot product of the classifier’s coefficients and the vector formed by the selected bands of SAR data, an array of predictions is constructed. A classifier can effectively be used on the same spatial extents as the ones where it was trained, or on any area with a similar landscape. While

Kreiser, Zachary↗