Search NASA⌕ Search

SEARCH · Search NASA

Results for “neighbor selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Unconventional Wells Interference: Supervised Machine Learning for Detecting Fracture Hits

The primary objective of the study was development of a machine learning (ML)-based workflow for fracture hit (“frac hit”) detection and monitoring using shale oil-field data such as drilling surveys, production history (oil and produced water), pressure, and fracking start time and duration records. The ML method takes advantage of long short-term memory (LSTM) and multilayer perceptron (MLP) neural networks to identify the frac hits due to hydraulic communication between the fracking child well(s) and the producing parent well(s) within the same pad (intra-pad interaction) and/or on different pads (inter-pad interaction). It utilizes time series of pressure and production data from within a pad and from adjacent pads. The workflow can capture time variable features of frac hits when the model architecture is deep and wide enough, with enough trainable parameters for deep learning and feature extraction, as demonstrated in this paper by using training and testing subsets of the field data from selected neighboring pads with over a couple of hundred wells. The study was focused on frac-hit interaction among paired wells and demonstrated that the ML model, once trained, can predict the frac-hit probability.

58 GEOSCIENCES↗

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

A TCN-Based Hybrid Forecasting Framework for Hours-Ahead Utility-Scale PV Forecasting

This paper presents a Temporal Convolutional Network (TCN) based hybrid PV forecasting framework for enhancing hours-ahead utility-scale PV forecasting. The hybrid framework consists of two forecasting models: a physics-based trend forecasting (TF) model and a data-driven fluctuation forecasting (FF) model. Three TCNs are integrated in the framework for: i) blending the inputs from different Numerical Weather Prediction sources for the TF model to achieve superior performance on forecasting hourly PV profiles, ii) capturing spatial-temporal correlations between detector sites and the target site in the FF model to achieve more accurate forecast of intra- hour PV power drops, and iii) reconciling TF and FF results to obtain coherent hours-ahead PV forecast with both hourly trends and intra-hour fluctuations well preserved. To automatically identify the most contributive neighboring sites for forming a detector network, a scenario-based correlation analysis method is developed, which significantly improves the capability of the FF model on capturing large power fluctuations caused by cloud movements. Here, the framework is developed, tested, and validated using actual PV data collected from 95 PV farms in North Carolina. Simulation results show that the performance of 6 hours ahead PV power forecasting is improved by 20% - 30% compared with state-of-the-art methods.

42 ENGINEERING↗

Models and Measurements Quantify Photon Recycling, Charge-Carrier Diffusion and Photon Scattering Contributions to Photoluminescence in InP Nanowire Arrays

Nanowire arrays present many unique advantages for solar-to-chemical energy conversion. One possible advantage is that photon recycling between neighboring nanowires has the potential to increase solar energy conversion efficiencies. Here, in this work, we explore three underlying mechanisms of optical and electronic coupling between neighboring nanowires─incident photon scattering, photon recycling, and charge-carrier transport from the photoexcited nanowire to the neighboring nanowire via the underlying substrate─using single nanowire-level microscopy and spectroscopy measurements. We present a comprehensive analysis of light absorption and emission of a single nanowire at open circuit, and subsequent re-absorption and re-emission by a neighboring nanowire. We developed a novel correlated single nanowire microspectroscopy and widefield imaging methodology to spatially resolve photon communication pathways between neighboring nanowires and selectively image re-emitted and reflected photons. We developed unique multiphysics models to couple wave optics and semiconductor photophysics to especially isolate contributions from photon recycling and electronic transport to photon emission from neighboring nanowires. By systematically varying the morphologies of the nanowires modeled, we identified pathways to maximize photon recycling between neighboring nanowires. We concluded that the measured photoluminescence is more strongly influenced by the diffusion of charge carriers as compared to photon recycling in materials with moderate-to-large charge-carrier mobilities (>10 cm 2 V –1 s –1 ), and that photon recycling dictates photoluminescence intensity only when the charge-carrier mobility is low (<1 cm 2 V –1 s –1 ). The experimental and simulation platforms developed herein for photon management strategies can be leveraged by the semiconductor photocatalysis community to enhance solar-to-chemical conversion efficiencies in semiconductor nanowire arrays.

25 ENERGY STORAGE↗

Generating a Simulated Fluid Flow Over an Aircraft Surface Using Anisotropic Diffusion

A fluid-flow simulation over a computer-generated aircraft surface is generated using a diffusion technique. The surface is comprised of a surface mesh of polygons. A boundary-layer fluid property is obtained for a subset of the polygons of the surface mesh. A pressure-gradient vector is determined for a selected polygon, the selected polygon belonging to the surface mesh but not one of the subset of polygons. A maximum and minimum diffusion rate is determined along directions determined using a pressure gradient vector corresponding to the selected polygon. A diffusion-path vector is defined between a point in the selected polygon and a neighboring point in a neighboring polygon. An updated fluid property is determined for the selected polygon using a variable diffusion rate, the variable diffusion rate based on the minimum diffusion rate, maximum diffusion rate, and angular difference between the diffusion-path vector and the pressure-gradient vector.

Rodriguez, David L.↗

Generating a Simulated Fluid Flow over a Surface Using Anisotropic Diffusion

A fluid-flow simulation over a computer-generated surface is generated using a diffusion technique. The surface is comprised of a surface mesh of polygons. A boundary-layer fluid property is obtained for a subset of the polygons of the surface mesh. A gradient vector is determined for a selected polygon, the selected polygon belonging to the surface mesh but not one of the subset of polygons. A maximum and minimum diffusion rate is determined along directions determined using the gradient vector corresponding to the selected polygon. A diffusion-path vector is defined between a point in the selected polygon and a neighboring point in a neighboring polygon. An updated fluid property is determined for the selected polygon using a variable diffusion rate, the variable diffusion rate based on the minimum diffusion rate, maximum diffusion rate, and the gradient vector.

Rodriguez, David L.↗

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence↗

Cryo-EM structures of the small-conductance Ca 2+ -activated K Ca 2.2 channel

Small-conductance Ca 2+ -activated K + (K Ca 2.1-K Ca 2.3) channels modulate neuronal and cardiac excitability. We report cryo-electron microscopy structures of the K Ca 2.2 channel in complex with calmodulin and Ca 2+ , alone or bound to two small molecule inhibitors, at 3.18, 3.50, 2.99 and 2.97 angstrom resolution, respectively. Extracellular S3-S4 loops in β-hairpin configuration form an outer canopy over the pore with an aromatic box at the canopy’s center. Each S3-S4 β-hairpin is tethered to the selectivity filter in the neighboring subunit by inter-subunit hydrogen bonds. This hydrogen bond network flips the aromatic residue (Tyr362) in the filter’s GYG signature by 180°, causing the outer selectivity filter to widen and water to enter the filter. Disruption of the tether by a mutation narrows the outer selectivity filter, realigns Tyr362 to the position seen in other K + channels, and significantly increases unitary conductance. UCL1684, a mimetic of the bee venom peptide apamin, sits atop the canopy and occludes the opening in the aromatic box. AP14145, an analogue of a therapeutic for atrial fibrillation, binds in the central cavity below the selectivity filter and induces closure of the inner gate. These structures provide a basis for understanding the small unitary conductance and pharmacology of K Ca 2.x channels.

59 BASIC BIOLOGICAL SCIENCES↗

Part-scale microstructure prediction for laser powder bed fusion Ti-6Al-4V using a hybrid mechanistic and machine learning model

Laser powder bed fusion (LPBF) Ti-6Al-4V is widely studied for use in structural applications in aerospace and medical industries, but mechanical anisotropy and microstructural inhomogeneity prohibits its wider adoption. Although successful microstructure prediction models have been developed, a remaining challenge is their limited integration across length/time scales and validation by experimental studies. Here, this work proposes a physics-augmented machine learning surrogate model to unite predictions of LPBF temperature, β phase morphology and texture, and α/α’ formation into a single framework that is calibrated and validated with experiments. First, a phase field (PF) model of the martensitic β→α’ transformation is developed and calibrated using data from in-situ synchrotron cyclic heating/cooling studies quantifying the variation of α phase fraction with time. In parallel, an established finite difference-Monte Carlo (FDMC) model predicts the part-scale temperature profile and β grain formation during solidification. A dataset is developed using LPBF cyclic temperature descriptors from the FDMC model as inputs and corresponding α/α’ phase fraction and width from the PF model as outputs. Five machine learning (ML) regression models are tested and optimized, having mean absolute error in testing ≤ 4 %, and the k-nearest neighbors (KNN) model is selected as the best performing. The KNN model is called at the nodal level during post-processing of the FDMC model to replace and downscale the response of the PF model. The combined agility and accuracy of the hybrid FDMC-ML model enables part-scale microstructure predictions that can be further used for property predictions to accelerate AM process optimization.

36 MATERIALS SCIENCE↗

HLA-Clus: HLA class I clustering based on 3D structure

In a previous paper, we classified populated HLA class I alleles into supertypes and subtypes based on the similarity of 3D landscape of peptide binding grooves, using newly defined structure distance metric and hierarchical clustering approach. Compared to other approaches, our method achieves higher correlation with peptide binding specificity, intra-cluster similarity (cohesion), and robustness. Here we introduce HLA-Clus, a Python package for clustering HLA Class I alleles using the method we developed recently and describe additional features including a new nearest neighbor clustering method that facilitates clustering based on user-defined criteria. The HLA-Clus pipeline includes three stages: First, HLA Class I structural models are coarse grained and transformed into clouds of labeled points. Second, similarities between alleles are determined using a newly defined structure distance metric that accounts for spatial and physicochemical similarities. Finally, alleles are clustered via hierarchical or nearest-neighbor approaches. We also interfaced HLA-Clus with the peptide:HLA affinity predictor MHCnuggets. By using the nearest neighbor clustering method to select optimal allele-specific deep learning models in MHCnuggets, the average accuracy of peptide binding prediction of rare alleles was improved. The HLA-Clus package offers a solution for characterizing the peptide binding specificities of a large number of HLA alleles. This method can be applied in HLA functional studies, such as the development of peptide affinity predictors, disease association studies, and HLA matching for grafting. HLA-Clus is freely available at our GitHub repository (https://github.com/yshen25/HLA-Clus).

59 BASIC BIOLOGICAL SCIENCES↗

A solution to the problem of SAR range curvature

When synthetic aperture radar systems are pushed to attain finer resolution at larger ranges than was previously the case for remote sensing purposes, the geometric signal aberration known as range curvature arises. Known techniques for correcting range curvature are exact at only one selected range, thus forcing neighboring ranges to use the same correction as an approximation. A solution to the problem is proposed that is exact at all ranges, thus simplifying and improving the image processing for such systems.

Raney, R. K.↗

A development of grid generation procedure for multicomponent aerodynamic configuration

Two approaches for solving the transonic flow in a multi-block grid were explored. The first approach examines a method involving "zonal decomposition" wherein block boundaries are treated as true boundary surfaces separating interfacing grids. The issues investigated involve techniques for matching solutions at a block boundary. A feasibility study was completed and the results are presented. The second approach involves overlapping grids for differencing across a block boundary near an artificially induced coordinate singularity occurring at a fictitious corner. This approach selects a set of neighboring nodes for the fictitious corner such that the resulting physical cells for a node are topologically the same as any other node on the airfoil surface.

Chen, H. C.↗

Statistics of associations among IR galaxies

In the course of expanding the search of Kleinmann et. al. (1988) for distant, infrared-luminous objects, the authors noticed (as is often remarked) that a large number of infrared-selected galaxies have close neighbors or show merger characteristics (e.g., tidal tails, distorted disks). Because the sample size is large (567 infrared galaxies and 2182 field galaxies), this sample is ideal for statistically examining the importance of interactions among infrared galaxies. In particular, the authors compare the nearest-neighbor distribution and the two-point correlation function of their sample with that of a control sample of field galaxies.

Gallimore, Jack F.↗

Evidence for Black Holes in Green Peas from WISE Colors and Variability

We explore the presence of active galactic nuclei (AGNs)/black holes in Green Pea galaxies (GPs), motivated by the presence of high-ionization emission lines such as He ii and [Ne iii] in their optical spectra. In order to identify AGN candidates, we used mid-infrared (MIR) photometric observations from the all-sky Wide-field Infrared Survey Explorer (WISE) mission for a sample of 1004 GPs. Considering only >5σ detections with no contamination from neighboring sources in AllWISE, we select 31 GPs out of 134 as candidate AGNs based on a stringent three-band WISE color diagnostic. Using multi-epoch photometry in W1 and W2 bands based on time-resolved unWISE coadd images, we find two sources exhibiting variability in both the WISE bands among 112 GPs with W1 ≤16 mag and no contamination from neighboring sources in unWISE. These two variable sources were selected as AGNs by the WISE three-band color diagnostic as well. Compared to variable AGN fractions observed among low-mass galaxy samples in previous studies, we find a higher fraction (~1.8%) of MIR variable sources among GPs, which demonstrates the uniqueness and importance of studying these extreme objects. Through this work, we demonstrate that MIR diagnostics are promising tools to select AGNs that may be missed by other selection techniques (including optical emission-line ratios and X-ray emission) in star formation-dominated, low-mass, low-metallicity galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗