Search NASA⌕ Search

SEARCH · Search NASA

Results for “masking algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluating Automated Face Identity-Masking Methods with Human Perception and a Deep Convolutional Neural Network

Face de-identification (or “masking”) algorithms have been developed in response to the prevalent use of video recordings in public places. Here, we evaluated the success of face identity masking for human perceivers and a deep convolutional neural network (DCNN). Eight de-identification algorithms were applied to videos of drivers’ faces, while they actively operated a motor vehicle. These masks were pre-selected to be applicable to low-quality video and to maintain coarse information about facial actions. Humans studied high-resolution images to learn driver identities and were tested on their recognition of active drivers in low-resolution videos. Faces in the videos were either unmasked or were masked by one of the eight algorithms. When participants were tested immediately after learning (Experiment 1), all masks reduced identification, with six of eight masks reducing identification to extremely poor performance. In a second experiment, two of the most effective masks were tested after a delay of 7 or 28 days. The delay did not further reduce identification of the masked faces. In all masked conditions, participants maintained stringent decision criteria, with low confidence in recognition, further indicating the effectiveness of the masks. Next, the DCNN performed an identity-matching task between high-resolution images and masked videos—a task analogous to that done by humans. The pattern of accuracy for the DCNN mirrored some, but not all, aspects of human performance, highlighting the need to test the effectiveness of identity masking for both humans and machines. The DCNN was also tested on its ability to match identity between masked and unmasked versions of the same video, based only on the face. DCNN performance for the eight masks offers insight into the nature of the information in faces that is coded in these networks.

97 MATHEMATICS AND COMPUTING↗

Automating Genetic Algorithm Mutations for Molecules Using a Masked Language Model

Inspired by the evolution of biological systems, genetic algorithms have been applied to generate solutions for optimization problems in a variety of scientific and engineering disciplines. For a given problem, a suitable genome representation must be defined along with a mutation operator to generate subsequent generations. Unlike natural systems which display a variety of complex rearrangements (e.g. mobile genetic elements), mutation for genetic algorithms commonly utilizes only random point-wise changes. Furthermore, generalizing beyond point-wise mutations poses a key difficulty as useful genome rearrangements depend on the representation and problem domain. To move beyond the limitations of manually defined point-wise changes, here we propose the use of techniques from masked language models to automatically generate mutations. As a first step, common subsequences within a given population are used to generate a vocabulary. The vocabulary is then used to tokenize each genome. A masked language model is trained on the tokenized data in order to generate possible rearrangements (i.e. mutations). In order to illustrate the proposed strategy, we use string representations of molecules and use a genetic algorithm to optimize for drug-likeness and synthesizability. Finally, our results show that moving beyond random point-wise mutations accelerates genetic algorithm optimization.

59 BASIC BIOLOGICAL SCIENCES↗

DMTN-197: Streak Masking in DM Image Processing

Streaks caused by satellites are a persistent problem in optical images, and their occurrence will likely increase in the coming years. Many Hyper Suprime-Cam images are already affected by streaks, and the same is expected for Rubin Observatory data. In the LSST Data Release Production (DRP), most streaks and other artifacts are already detected and masked by the algorithm CompareWarpAssembleCoadd. However, some streaks are not caught at this stage and can contaminate the final coadds and detection catalogs. To find and mask these, we adopt a morphologically-based method for detecting streaks, which uses the Kernel-Based Hough Transform to detect straight lines. Once the lines are detected, the streak profile is fit and the affected portion of the image is masked out. This algorithm, maskStreaks, is included in the meas_algorithms package and is used in CompareWarpAssembleCoadd to remove streaks from coadds. Similar implementation in the Alert Production Pipeline is possible but has not been implemented.

79 ASTRONOMY AND ASTROPHYSICS↗

Supporting information for Few-Shot Learning Enables Population-Scale Analysis of Leaf Traits in Populus trichocarpa

In this work, we use few-shot learning to segment the body and vein architecture of P. trichocarpa leaves from high-resolution scans obtained in the UC Davis common garden. Leaf and vein segmentation are formulated as separate tasks, in which convolutional neural networks (CNNs) are used to iteratively expand partial segmentations until reaching stopping criteria. Our leaf and vein segmentation approaches use just 50 and 8 manually traced images for training, respectively, and are applied to a set of 2,634 top and bottom leaf scans. We show that both methods achieve high segmentation accuracy, in some cases exceeding even human-level segmentation. The leaf and vein segmentations are subsequently used to extract 68 morphological traits using traditional open-source image processing tools, which are validated using real-world physical measurements. For a biological perspective, we perform a genome-wide association study using the vein density trait to discover novel genetic architectures associated with multiple physiological processes relating to leaf development and function. In addition to sharing all of the few-shot learning code (see https://github.com/jlager/few-shot-leaf-segmentation), we are releasing all images, manual segmentations, model predictions, 68 extracted leaf phenotypes, and a new set of SNPs called against the v4 P. trichocarpa genome for 1,419 genotypes. The data folder includes all images, ground truth segmentations, predicted segmentations, and extracted leaf traits. All images encode the sample ID in the file name by indicating the treatment, block, row, position, and leaf side, respectively. For example, the file, C_1_1_2_bot.jpeg, indicates the control treatment, block 1, row 1, position 2, and the bottom side of the leaf. Tabulated results include position IDs as well as the corresponding genotype IDs. The images folder includes the 2,906 high-resolution leaf scans taken in the field. The leaf_masks folder includes 50 ground truth segmentations used for training the leaf tracing algorithm. The leaf_preds folder includes the 2,906 predicted segmentations from the leaf tracing algorithm. The vein_masks folder includes 8 ground truth segmentations used for training the vein growing algorithm. The vein_preds folder includes the 1,453 predicted segmentations from the vein growing algorithm. The vein_probs folder includes the 1,453 predicted probability maps from the vein growing algorithm before thresholding. The genomes folder includes the set of SNPs called against the v4 P. trichocarpa genome for 1,419 genotypes with a README file detailing the steps taken. The results folder includes: (i) raw values of the 68 predicted leaf traits in digital_traits.tsv, (ii) manually measured values of petiole length and width in manual_traits.tsv, (iii) thin plate spline (TPS) adjusted values of the vein density trait in vein_density_tps_adj.tsv, (iv) best linear unbiased prediction (BLUP) adjusted values of the vein density trait in vein_density_blups.tsv, and (v) GWAS results for the vein density trait, including chromosome positions and corresponding P values, in gwas_results.csv.

09 BIOMASS FUELS↗

KAZR Hydrometeor and Insect Masks

The Ka-band ARM zenith pointing radar (aka, KAZR) is so sensitive that it detects cloud particles and individual insects. While detecting insects with Ka-band radar is beneficial and desirable to advance radar entomology, this sensitivity can be detrimental to radar meteorology because insects could be interpreted as clouds or precipitation. For example, misclassifying insects as clouds has been a problem for the ARM Active Remote Sensing of Clouds (ARSCL) Value Added Product since its inception (Clothiaux et al. 2000). Based on cloud particle and insect radar scattering properties, an algorithm was developed that identifies clouds, raindrops, ice particles, and insects in KAZR co- and cross-polarimeteric Doppler velocity spectra. The algorithm produces affirmative masks in the KAZR native time and height resolution indicating time-height locations of hydrometeors and insects. The hydrometeor mask contains binary information (e.g., yes/no hydrometeor presence), and the insect mask includes a proxy for insect activity that increases when more insects are detected in the Doppler velocity spectra. The algorithm was developed using KAZR medium sensitivity mode (MD) observations and was applied to two summer seasons of KAZR observations at the Southern Great Plains (SGP) Central Facility: May-October 2018 and 2019. In the future, this data set will be expanded to include other KAZR operating modes and observations from other ARM field sites. Details of the algorithm and data set can be found in: Williams, C.R., K.L. Johnson, S.E. Giangrande, J. C. Hardin, R. Oktem, and D. M. Romps, 2021: Identifying Insects, Clouds, and Precipitation using Vertically Pointing Polarimetric Radar Doppler Velocity Spectra. Atmospheric Measurement Techniques, submitted 6-Feb-2021. https://amt.copernicus.org/preprints/amt-2021-27/#discussion.For more information on the ARSCL VAP, see Clothiaux, E. E., T. P. Ackerman, G. G. Mace, K. P. Moran, R. T. Marchand, M. A. Miller, and B. E. Martner, 2000; J. Appl. Meteor., 39, 645-665.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of cloud height, optical thickness, and phase retrievals from the CHROMA algorithm applied to Sentinel-3 OLCI data

We previously developed the Cloud Height Retrieval from O 2 Molecular Absorption (CHROMA) algorithm for the Ocean Color Instrument (OCI) on the new NASA Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission. Here, we apply CHROMA to observations from the Ocean Land Colour Instrument (OLCI) to guide expectations for PACE, as it will take some time to obtain large-scale validation data for OCI. We use cloud top height (CTH), phase, and (for liquid clouds) cloud optical thickness (COT) data from the ground-based Atmospheric Radiation Measurement (ARM) network to evaluate the OLCI retrievals. We found that OLCI and Moderate Resolution Imaging Spectroradiometer (MODIS) CTH compare similarly well to the ARM reference. OLCI has a tendency to underestimate CTH as CTH increases, and algorithm assumptions about cloud geometric thickness may contribute to this. ARM COT from multifilter shadowband radiometers (MFRSR) and Sun photometers are well-correlated with one another, albeit with a roughly 30 % offset on average; OLCI and MODIS COT agree more closely with the MFRSR data. OLCI retrieval uncertainty estimates show skill at telling low-uncertainty cases from high-uncertainty ones, although CTH uncertainties are underestimated. Additionally, we compare the OLCI data to satellite retrievals based on thermal infrared measurements from MODIS and Sea and Land Surface Temperature Radiometer (SLSTR) data. Differences are broadly consistent with physical expectations based on the A-band vs. thermal techniques, although one key challenge in such aggregated comparisons is different cloud masking sensitivities and algorithm failure rates meaning additional sampling differences are introduced. We conclude by discussing the transition to and possible enhancements for PACE OCI.

Sayer, Andrew M. [Univ. of Maryland Baltimore Coun↗

Disclosure of the XRD pipeline software

The XRD pipeline software is a program to reduce 2D X-ray diffraction data from area detectors with advanced algorithms on automatic masking and image segmentation, which help characterization of multiple phases recorded in the data and facilitate data analysis.SF-25-114

XU, WENQIAN [Argonne National Laboratory (ANL), Ar↗

Micropulse Lidar Cloud Mask Machine-Learning Value-Added Product Report

Cloud detection algorithms of various techniques have been developed and applied to atmospheric ground-based lidar data to identify cloud boundaries and produce clouds masks. While these algorithms are able to identify a wide variety of cloud types and conditions, it is often observed that the algorithms can still fail to accurately detect clouds that are readily discernible when inspecting the lidar imagery. Based on this observation, an alternative approach for cloud detection is to take advantage of machine-learning capabilities and the trained human eye as an interpreter of lidar images, and in turn, to train a neural network to recognize the desired features in the lidar data.

54 ENVIRONMENTAL SCIENCES↗

NREL GOES Reference Document

The purpose of this document is to create a single reference document representing abridged versions of the PATMOS-x Algorithm Theoretical Basis Documents (ATBDs) specifically relevant to GOES 16-18 processing. This document has been compiled from several documents: Cloud Mask ATBD, AWG Cloud Height Algorithm (ACHA) ATBD, Daytime Cloud Optical and Microphysical Properties (DCOMP) ATBD, and CLAVR-x User's Guide.

14 SOLAR ENERGY↗

Deep learning for morphological identification of extended radio galaxies using weak labels

Abstract The present work discusses the use of a weakly-supervised deep learning algorithm that reduces the cost of labelling pixel-level masks for complex radio galaxies with multiple components. The algorithm is trained on weak class-level labels of radio galaxies to get class activation maps (CAMs). The CAMs are further refined using an inter-pixel relations network (IRNet) to get instance segmentation masks over radio galaxies and the positions of their infrared hosts. We use data from the Australian Square Kilometre Array Pathfinder (ASKAP) telescope, specifically the Evolutionary Map of the Universe (EMU) Pilot Survey, which covered a sky area of 270 square degrees with an RMS sensitivity of 25–35 $\mu$ Jy beam $^{-1}$ . We demonstrate that weakly-supervised deep learning algorithms can achieve high accuracy in predicting pixel-level information, including masks for the extended radio emission encapsulating all galaxy components and the positions of the infrared host galaxies. We evaluate the performance of our method using mean Average Precision (mAP) across multiple classes at a standard intersection over union (IoU) threshold of 0.5. We show that the model achieves a mAP $_{50}$ of 67.5% and 76.8% for radio masks and infrared host positions, respectively. The network architecture can be found at the following link: https://github.com/Nikhel1/Gal-CAM

Astronomy & Astrophysics↗

Adaptive language model training for molecular design

Abstract The vast size of chemical space necessitates computational approaches to automate and accelerate the design of molecular sequences to guide experimental efforts for drug discovery. Genetic algorithms provide a useful framework to incrementally generate molecules by applying mutations to known chemical structures. Recently, masked language models have been applied to automate the mutation process by leveraging large compound libraries to learn commonly occurring chemical sequences (i.e., using tokenization) and predict rearrangements (i.e., using mask prediction). Here, we consider how language models can be adapted to improve molecule generation for different optimization tasks. We use two different generation strategies for comparison, fixed and adaptive. The fixed strategy uses a pre-trained model to generate mutations; the adaptive strategy trains the language model on each new generation of molecules selected for target properties during optimization. Our results show that the adaptive strategy allows the language model to more closely fit the distribution of molecules in the population. Therefore, for enhanced fitness optimization, we suggest the use of the fixed strategy during an initial phase followed by the use of the adaptive strategy. We demonstrate the impact of adaptive training by searching for molecules that optimize both heuristic metrics, drug-likeness and synthesizability, as well as predicted protein binding affinity from a surrogate model. Our results show that the adaptive strategy provides a significant improvement in fitness optimization compared to the fixed pre-trained model, empowering the application of language models to molecular design tasks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automatic Code Generation for High-Performance Graph Algorithms

Graph problems are common across fields of scientific computing and social sciences. However, despite their importance, implementing graph algorithms effectively on modern computing systems is a challenging task that requires significant programming effort and generally results in customized implementations. Current computing and memory hierarchies are not architected for irregular computations resulting in challenges for graph algorithms to achieve high performance on those architectures. In this paper, we present GraphX, a novel compiler framework and DSL designed to simplify the development of efficient graph algorithms and achieve high performance on modern computing systems. GraphX consists of a DSL for efficient implementation of graph algorithms, various optimizations, such as support for sparse linear algebra and workspace transformations, optimized graph primitives, including semiring and masking, and a high-performance code generation engine. Using GraphX, users can implement graph algorithms using a semantically-rich language with graph-oriented operators. GraphX uses these semantics to automatically generate efficient code for target architectures, increasing performance and portability across architectures. The composable nature of GraphX makes it possible to extend the set of optimizations and architectures without modifying the source code. We demonstrate GraphX outperforms state-of-the-art graph libraries, such as LAGraph, up to $3.7 speedup in semiring operations, $2.19 speedup in an important sparse computational kernel, and $9.05 speedup in graph processing algorithms.

compiler, graph algorithms, semiring, masking, wor↗

Algorithms for Non-Negative Matrix Factorization on Noisy Data With Negative Values

Non-negative matrix factorization (NMF) is a dimensionality reduction technique that has shown promise for analyzing noisy data, especially astronomical data. For these datasets, the observed data may contain negative values due to noise even when the true underlying physical signal is strictly positive. Prior NMF work has not treated negative data in a statistically consistent manner, which becomes problematic for low signal-to-noise data with many negative values. In this paper we present two algorithms, Shift-NMF and Nearly-NMF, that can handle both the noisiness of the input data and also any introduced negativity. Both of these algorithms use the negative data space without clipping or masking and recover non-negative signals without any introduced positive offset that occurs when clipping or masking negative data. We demonstrate this numerically on both simple and more realistic examples, and prove that both algorithms have monotonically decreasing update rules.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A neural network for cloud masking in WorldView satellite imagery

This repository includes the cloud detection algorithm used in the manuscript 'Topography controls variability in circumpolar permafrost thaw pond expansion' by Abolt et al. It also includes a demonstration of the algorithm's use at a survey area in northern Alaska and a demonstration of training the algorithm. This dataset includes .pt place-text, .py python, .tif image, .ipynb Jupyter notebook, and .txt text files.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic) was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

DeepGhostBusters: Using Mask R-CNN to Detect and Mask Ghosting and Scattered-Light Artifacts from Optical Survey Images

Wide-field astronomical surveys are often affected by the presence of undesirable reflections (often known as "ghosting artifacts" or "ghosts") and scattered-light artifacts. The identification and mitigation of these artifacts is important for rigorous astronomical analyses of faint and low-surface-brightness systems. However, the identification of ghosts and scattered-light artifacts is challenging due to a) the complex morphology of these features and b) the large data volume of current and near-future surveys. In this work, we use images from the Dark Energy Survey (DES) to train, validate, and test a deep neural network (Mask R-CNN) to detect and localize ghosts and scattered-light artifacts. We find that the ability of the Mask R-CNN model to identify affected regions is superior to that of conventional algorithms and traditional convolutional neural networks methods. We propose that a multi-step pipeline combining Mask R-CNN segmentation with a classical CNN classifier provides a powerful technique for the automated detection of ghosting and scattered-light artifacts in current and near-future surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

TempestExtremes v2.1: a community framework for feature detection, tracking, and analysis in large datasets

TempestExtremes (TE) is a multifaceted framework for feature detection, tracking, and scientific analysis of regional or global Earth system datasets on either rectilinear or unstructured/native grids. Version 2.1 of the TE framework now provides extensive support for examining both nodal (i.e., pointwise) and areal features, including tropical and extratropical cyclones, monsoonal lows and depressions, atmospheric rivers, atmospheric blocking, precipitation clusters, and heat waves. Available operations include nodal and areal thresholding, calculations of quantities related to nodal features such as accumulated cyclone energy and azimuthal wind profiles, filtering data based on the characteristics of nodal features, and stereographic compositing. This paper describes the core algorithms (kernels) that have been added to the TE framework since version 1.0, including algorithms for editing pointwise trajectory files, composition of fields around nodal features, generation of areal masks via thresholding and nodal features, and tracking of areal features in time. Several examples are provided of how these kernels can be combined to produce composite algorithms for evaluating and understanding common atmospheric features and their underlying processes. These examples include analyzing the fraction of precipitation from tropical cyclones, compositing meteorological fields around extratropical cyclones, calculating fractional contribution to poleward vapor transport from atmospheric rivers, and building a climatology of atmospheric blocks.

58 GEOSCIENCES↗

Sensitive detection of structural dynamics using a statistical framework for comparative crystallography

Chemical and conformational changes are crucial to protein function and its pharmacological control. X-ray crystallography can reveal these changes in atomic detail, but standard analysis methods, which refine separate datasets, often overlook differences that are subtle or arise in only a subset of molecules. Direct comparison of crystallographic datasets is, in principle, more powerful, but systematic errors (“scales”) often mask changes in the crystallographic observables (“structure factors”). Machine learning algorithms that jointly estimate scales and structure factors can address this limitation. Here, we augment this approach with multivariate, structured priors derived from crystallographic theory, implemented in the variational deep learning framework Careless. Doing so strongly improves the detection of protein dynamics, element-specific anomalous signals, and the binding of drug candidates, offering a robust approach to comparative crystallography and, potentially, to detection of protein dynamics by other structure determination methods.

Hekstra, Doeke R. [Harvard Univ., Cambridge, MA (U↗

SNAP diffraction dataset for 2023 SMC data challenge

The data provided this challenge is ice under high pressure measured using the Spallation Neutrons and Pressure Diffractometer (SNAP) at the Spallation Neutron Source (SNS) at Oak Ridge National Laboratory. The data is stored in a hdf5 file following the NeXus standard and can be read with tools built for either. While the NeXus format is self-describing, there is benefit to explaining some details. The data is stored in a single NXdata entry within a single NXentry. The NXdata has several fields denoting the 3-dimensional data (signal), the axes (D0 is the Qx axis, D1 is the Qy axis, and D2 is the Qz axis), and fields for the uncertainties and masking information. The data can be quickly viewed using the LoadMD algorithm and slice viewer in the Mantid workbench https://www.mantidproject.org.

36 MATERIALS SCIENCE↗