Search NASASearch

SEARCH · Search NASA

Results for “feature vectors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

End-to-End Automated Segmentation Framework for Four-Dimensional Scanning Transmission Electron Microscopy Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) is powerful for rapidly characterizing arrays of nanoparticles produced via high-throughput synthesis. However, such 4D-STEM datasets typically contain thousands of nanoparticles, each characterized by thousands of diffraction patterns spatially distributed across the nanoparticle, necessitating efficient and comprehensive analysis. We propose an end-to-end segmentation framework to automatically segment each nanoparticle into regions with distinct composition/orientation of crystal grains, using only the 4D-STEM data. Bragg disk information is extracted in a physics-informed manner from the diffraction patterns at each spatial location and combined with the real space coordinates to form feature vectors. These feature vectors are then used as inputs to a Gaussian mixture model (GMM) to segment the nanoparticle into distinct regions. We also develop two visualization tools based on the GMM outputs to infer the interface transition and the degree of superposition. Our framework comprehensively integrates machine learning tools and physics knowledge, and provides a basis for substantially compressing enormous 4D-STEM datasets, e.g., by replacing the full 4D-STEM dataset for each nanoparticle with only a single set of Bragg disk features for each distinct crystal grain identified in the nanoparticle. In this article, we demonstrate the power of our framework by presenting results for real, complex datasets.

47 OTHER INSTRUMENTATION

Characterizing skyrmion flow phases with principal component analysis

Principal component analysis (PCA) is a powerful method that can identify patterns in large, complex data sets by constructing low-dimensional order parameters from higher-dimensional feature vectors. There are increasing efforts to use space-and-time-dependent PCA to detect transitions in nonequilibrium systems that are difficult to characterize with equilibrium methods. Here, we demonstrate that feature vectors incorporating the position and velocity information of driven skyrmions moving through random disorder permit PCA to resolve different types of disordered skyrmion motion as a function of driving force and the ratio of the Magnus force to the dissipation. Since the Magnus force creates gyroscopic motion and a finite Hall angle, skyrmions can exhibit a greater range of flow phases than what is observed in overdamped driven systems with quenched disorder. We show that in addition to identifying previously known skyrmion flow phases, PCA detects several additional phases, including different types of channel flow, moving fluids, and partially ordered states. Guided by the PCA analysis, we further characterize the disordered flow phases to elucidate the different microscopic dynamics and show that the changes in the PCA-derived order parameters can be connected to features in bulk transport measures, including the transverse and longitudinal velocity-force curves, differential conductivity, topological defect density, and changes in the skyrmion Hall angle as a function of drive. We discuss how asymmetric feature vectors can be used to improve the resolution of the PCA analysis, and how this technique can be extended to find disordered phases in other nonequilibrium systems with time-dependent dynamics.

36 MATERIALS SCIENCE

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (

Detecting damage in composites using volume decomposition analysis of tomographic data

Detection of damage in a single tow ceramic matrix composite specimen has been achieved using orthogonal decomposition of volumetric tomographic datasets collected at four tensile loads. This decomposition approach has been applied at two different length scales: (i) individual fibres and (ii) bulk volumes containing fibres and matrix material. Volumes were first decomposed to feature vectors, orders of magnitude smaller than the original volume they describe, and then comparisons between datasets at different load levels were made in feature vector space. The results show quantitative measurements of damage location, damage morphology and the relative growth of this damage with increased load when compared with a dataset with less or no damage. No prior knowledge of the dataset or training of algorithms is required for damage to be detected, it is only necessary that at least two datasets are available for comparison, e.g. from in situ or repeated scanning measurements. Results are generated on significantly shorter timescales when compared with previous automated approaches to tomography data processing. This approach has the potential to be applied to damage detection in a range of materials through comparisons of volumetric datasets from a range of measurement or computational techniques.

Middleton, Ceri A.

Machine-learning-enabled on-the-fly analysis of RHEED patterns during thin film deposition by molecular beam epitaxy

Thin film deposition is a fundamental technology for the discovery, optimization, and manufacturing of functional materials. Deposition by molecular beam epitaxy (MBE) typically employs reflection high-energy electron diffraction (RHEED) as a real-time in situ probe of the growing film. However, the state-of-the-art for RHEED analysis during deposition requires human observation. Here, we present an approach using machine learning (ML) methods to monitor, analyze, and interpret RHEED images on-the-fly during thin film deposition. In the analysis workflow, RHEED pattern images are collected at one frame per second and featurized using a pretrained deep convolutional neural network. The feature vectors are then statistically analyzed to identify changepoints; these changepoints can be related to changes in the deposition mode from initial film nucleation to a transition regime, smooth film deposition, and in some cases, an additional transition to a rough, islanded deposition regime. The feature vectors are additionally analyzed via graph analysis and community classification. The graph is quantified as a stabilization plot, and we show that inflection points in the stabilization plot correspond to changes in the growth regime. The full RHEED analysis workflow is termed RHAAPsody and includes data transfer and output to a visual dashboard. We demonstrate the functionality of RHAAPsody by analyzing the precaptured RHEED images from epitaxial depositions of anatase TiO2 on SrTiO3(001) and show that the analysis workflow can be executed in less than 1 s. Our approach shows promise as one component of ML-enabled real-time feedback control of the MBE deposition process.

36 MATERIALS SCIENCE

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]

Contrasting Time-Frequency Representations for Unknown Waveform Detection

In real-world applications like spectrum management and interference detection, dealing with unseen electromagnetic waveforms is critical. Although some methods attempt to simulate open set data using generator models, they face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. This results in difficulties capturing distinctive features across classes, especially in dynamic scenarios where new classes emerge. To detect unseen waveforms, we propose combining time and frequency domain features with cosine similarity loss to enhance feature distinctiveness and enabling more accurate predictions. This approach efficiently captures more comprehensive information than single-domain representations or approaches without cosine loss. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10\% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS

A knowledge-informed large language model framework for U.S. nuclear power plant shutdown initiating event classification for probabilistic risk assessment

Identifying and classifying shutdown initiating events (SDIEs) is critical for developing shutdown probabilistic risk assessment for nuclear power plants. Existing computational approaches cannot achieve satisfactory performance due to the challenges of unavailable large, labeled datasets, imbalanced event types, and label noise. To address these challenges, we propose a hybrid pipeline that integrates a knowledge-informed machine learning model to prescreen non-SDIEs and a large language model (LLM) to classify SDIEs into four types. In the prescreening stage, we proposed a set of 44 SDIE text patterns that consist of the most salient keywords and phrases from six SDIE types. Text vectorization based on the SDIE patterns generates feature vectors that are highly separable by using a simple binary classifier. The second stage builds Bidirectional Encoder Representations from Transformers (BERT)-based LLM, which learns generic English language representations from self-supervised pretraining on a large dataset and adapts to SDIE classification by fine-tuning it on an SDIE dataset. The proposed approaches are evaluated on a dataset with 10,928 events using precision, recall ratio, F 1 score, and average accuracy. In conclusion, the results demonstrate that the prescreening stage can exclude more than 97% non-SDIEs, and the LLM achieves an average accuracy of 95.1% for SDIE classification.

99 - GENERAL AND MISCELLANEOUS

Q-Cluster: Quantum Error Mitigation Through Noise-Aware Unsupervised Learning

Quantum error mitigation (QEM) is critical in reducing the impact of noise in the pre-fault-tolerant era, and is expected to complement error correction in fault-tolerant quantum computing (FTQC). In this work, we propose a novel QEM approach, Q-Cluster, that uses unsupervised learning (clustering) to reshape the measured bit-string distribution. Our approach starts with a simplified bit-flip noise model. It first performs clustering on noisy measurement results, i.e., bit-strings, based on the Hamming distance. The centroid of each cluster is calculated using a qubit-wise majority vote. Next, the noisy distribution is adjusted with the clustering outcomes and the bitflip error rates using Bayesian inference. Our simulation results show that Q-Cluster can mitigate high noise rates (up to 40% per qubit) with the simple bit-flip noise model. However, real quantum computers do not fit such a simple noise model. To address the problem, we (a) apply Pauli twirling to tailor the complex noise channels to Pauli errors, and (b) employ a machine learning model, ExtraTrees regressor, to estimate an effective bit-flip error rate using a feature vector consisting of machine calibration data (gate & measurement error rates), circuit features (number of qubits, numbers of different types of gates, etc.) and the shape of the noisy distribution (entropy). Our experimental results show that our proposed Q-Cluster scheme improves the fidelity by a factor of 1.46x, on average, compared to the unmitigated output distribution, for a set of low-entropy benchmarks on five different IBM quantum machines. Our approach outperforms the state-of-art QEM approaches RZNE [28], M3 [24], Hammer [35], and QBEEP [33] by 1.26x,1.29x,1.47x, and 2.65 x, respectively.

42 ENGINEERING

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

Training machine learning models for radioisotope identification using gamma spectroscopy remains an elusive challenge for many practical applications, largely stemming from the difficulty of acquiring and labeling large, diverse experimental datasets. Simulations can mitigate this challenge, but the accuracy of models trained on simulated data can deteriorate substantially when deployed to an out-of-distribution operational environment. In this study, we demonstrate that unsupervised domain adaptation (UDA) can improve the ability of a model trained on synthetic data to generalize to a new testing domain, provided unlabeled data from the target domain are available. Conventional supervised techniques are unable to utilize this data because the absence of isotope labels precludes defining a supervised classification loss. Instead, we first pretrain a spectral classifier using labeled synthetic data and subsequently leverage unlabeled target data to align the learned feature representations between the source and target domains. We compare a range of different UDA techniques, finding that minimizing the maximum mean discrepancy (MMD) between source and target feature vectors yields the most consistent improvement to testing scores. For instance, using a custom transformer-based neural network, we achieved a testing accuracy of $0.904 \pm 0.022$ on an experimental LaBr test set after performing unsupervised feature alignment via MMD minimization, compared to $0.754 \pm 0.014$ before alignment. Overall, our results highlight the potential of using UDA to adapt a radioisotope classifier trained on synthetic data for real-world deployment.

Lalor, Peter W.

Reduced Order Modeling conditioned on monitored features for response and error bounds estimation in engineered systems

Reduced Order Models (ROMs) form essential tools across engineering domains by virtue of their function as surrogates for computationally intensive digital twinning simulators. Although purely data-driven methods are available for ROM construction, schemes that allow to retain a portion of the physics tend to enhance the interpretability and generalization of ROMs. However, physics-based techniques can adversely scale when dealing with nonlinear systems that feature parametric dependencies. This study introduces a generative physics-based ROM that is suited for nonlinear systems with parametric dependencies and is additionally able to provide numerical error bounds associated with the respective estimates. A main contribution of this work is the conditioning of these parametric ROMs to features that can be derived from monitoring measurements, feasibly in an online fashion. This is contrary to most existing ROM schemes, which remain restricted to the prescription of the physics-based, and usually a priori unknown, system parameters. Our work utilizes conditional Variational Autoencoders to continuously map the required reduction bases to a feature vector extracted from limited output measurements, while additionally allowing for a probabilistic assessment of the ROM-estimated Quantities of Interest. An auxiliary task using a neural network-based parametrization of suitable probability distributions is introduced to re-establish the link with physical model parameters. We verify the proposed scheme on a series of simulated case studies incorporating effects of geometric and material nonlinearity under parametric dependencies related to system properties and input load characteristics.

Conditional VAEs

Accelerating the Discovery of New, Single Phase High Entropy Ceramics via Active Learning

High-entropy ceramics have garnered interest due to their remarkable hardness, compressive strength, thermal stability, and fracture toughness; yet the discovery of new high-entropy ceramics (out of a tremendous number of possible elemental permutations) still largely requires costly, inefficient, trial-and-error experimental and computational approaches. The entropy forming ability (EFA) factor was recently proposed as a computational descriptor that positively correlates with the likelihood that a 5-metal high-entropy carbide (HECs) will form the desired single phase, homogeneous solid solution; however, discovery of new compositions is computationally expensive. If you consider 8 candidate metals, the HEC EFA approach uses 49 optimizations for each of the 56 unique 5-metal carbides, requiring a total of 2744 costly density functional theory calculations. Here, we describe an orders-of-magnitude more efficient active learning (AL) approach for identifying novel HECs. To begin, we compared numerous methods for generating composition-based feature vectors (e.g., magpie and mat2vec), deployed an ensemble of machine learning (ML) models to generate an average and distribution of predictions, and then utilized the distribution as an uncertainty. Here we then deployed an AL approach to extract new training data points where the ensemble of ML models predicted a high EFA value or was uncertain of the prediction. Our approach has the combined benefit of decreasing the amount of training data required to reach acceptable prediction qualities and biases the predictions toward identifying HECs with the desired high EFA values, which are tentatively correlated with the formation of single phase HECs. Using this approach, we increased the number of 5-metal carbides screened from 56 to 15,504, revealing 4 compositions with record-high EFA values that were previously unreported in the literature. Our AL framework is also generalizable and could be modified to rationally predict optimized candidate materials/combinations with a wide range of desired properties (e.g., mechanical stability, thermal conductivity).

36 MATERIALS SCIENCE

Structure Sensitive Reaction Kinetics of Chiral Molecules on Intrinsically Chiral Surfaces

Enantiospecific heterogeneous catalysis utilizes chiral surfaces to resolve enantiomers via structure sensitive surface chemistry. The catalyst design challenge is the identification of chiral surface structures that maximize enantiospecificity. Herein, we develop data driven models for the enantiospecificity of tartaric acid reactions on chiral Cu(hkl) R&S surfaces. Measurements of enantiospecific rate constants were obtained by using curved Cu(hkl) R&S surfaces that enable kinetic measurements on hundreds of chiral surface orientations. One model uses feature vectors derived from generalized coordination numbers to capture the local structure around Cu atoms exposed by the Cu(hkl) R&S surfaces. The second model introduces the use of chiral cubic harmonic functions to capture the symmetry constraints of the face-centered cubic Cu structure. The model using 58 generalized coordination numbers has a fitting error similar to that of the model using only 5 cubic harmonic functions. The two models predict maxima in the enantiospecificity on surfaces with very similar surface orientations. The models developed in this work are applicable for any enantiospecific reaction happening on any chiral material with a cubic lattice structure, opening the way to understanding the surface structure sensitivity of the enantiospecific reaction kinetics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Predictive Indicators of the Performance of Large Language Models

In several mission contexts, it is desirable to estimate the performance of large language models (LLMs) on tasks that we cannot run directly. In light of published “scaling laws” our hypothesis is that some tasks should be consistently more challenging than others based on characteristics of the task. The goal of this project was to begin quantifying how much information about LLM performance can be gained from the features of a model and a task. Two of our statistical models struggled to converge. Pass/fail test results may provide limited information for inference beyond model quality and task difficulty, but we see no evidence at this time for significant feature interaction effect sizes, arguing for simple models. Future work extending the models to capitalize on perplexity of ground truth answers is suggested. This project also introduces “Depth of Knowledge Variant Testing” as a strategy for more finely assessing language models on open domain question and answer tasks. We developed sets of questions that ask a language model to produce similar information while demonstrating increasing depth of knowledge, and also relabeled existing Q&A test questions with their depth of knowledge. Our results suggest further consideration of Bloom’s taxonomy and further refinement of prompts to properly elicit information at varying depths. In the course of this work, we set up a basic infrastructure for standardizing tasks and testing many language models on these tasks. In addition to testing the predictive quality of model features and performance across test suites, with this project we have introduced two new task features to contextualize each test question: the Dewey Classification main category of information covered, and the Bloom’s taxonomy level that corresponds to the depth of knowledge probed by the question. Splits across these and other features produced over five hundred task subtypes with distinct feature vectors, which we tested on half a dozen models.

97 MATHEMATICS AND COMPUTING

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory

Leveraging design of experiments to build chemometric models for the quantification of uranium (VI) and HNO3 by Raman spectroscopy

Partial least squares regression (PLSR) and support vector regression (SVR) models were optimized for the quantification of U(VI) (10–320 g L −1 ) and HNO 3 (0.6–6 M) by Raman spectroscopy with optimized calibration sets chosen by optimal design of experiments. The designed approach effectively minimized the number of samples in the calibration set for PLSR and SVR by selecting sample concentrations with a quadratic process model, despite complex confounding and covarying spectral features in the spectra. The top PLS2 model resulted in percent root mean square errors of prediction for U(VI), HNO 3 , and NO 3 − of 3.7%, 3.6%, and 2.9%, respectively. PLS1 models performed similarly despite modeling an analyte with a majority linear response (i.e., uranyl symmetric stretch) and another with more covarying vibrational modes (i.e., HNO 3 ). Partial least squares (PLS) model loadings and regression coefficients were evaluated to better understand the relationship between weaker Raman bands and covarying spectral features. Support vector machine models outperformed PLS1 models, resulting in percent root mean square error of prediction values for U(VI) and HNO 3 of 1.5% and 3.1%, respectively. The optimal nonlinear SVR model was trained using a similar number of samples (11) compared with the PLSR model, even though PLS is a linear modeling approach. The generic D-optimal design presented in this work provides a robust statistical framework for selecting training set samples in disparate two-factor systems. This approach reinforces Raman spectroscopy for the quantification of species relevant to the nuclear fuel cycle and provides a robust chemometric modeling approach to bolster online monitoring in challenging process environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

LandScan HD: a high-resolution gridded ambient population methodology for the world

Unwarned population distributions accounting for routine human activities are needed to address many global human security challenges, including disasters, conflict, and infrastructure demand. LandScan High Definition (LSHD) supports this need through gridded ambient population estimates that measure average human presence between daytime and nighttime at a high spatial resolution of 3 arcseconds (approximately 90 m). Although LSHD has traditionally been produced on a country-specific basis, advances in global foundational data and computational resources now enable scaling its methodology to the world. Combining aspects of top-down and bottom-up gridded population methods, LSHD allocates subnational population totals from authoritative statistics to built-up areas based on occupancy estimates for multiple facility types (e.g., residential, commercial) and then reaggregates these estimates to a global population grid. We scale this approach by organizing the LSHD data stack into a 1° resolution tileset of vector analytic features, enabling an efficient and repeatable workflow for all countries worldwide. Examining the Philippines as an output of the global LSHD baseline dataset, we contrast unwarned and residential (WorldPop) population distributions by (1) exploring a practical application of flood risk assessment and (2) evaluating their congruence with outcomes of collective human activities (subnational CO 2 emissions). Finally, we discuss plans to address current LSHD limitations through data/modeling and uncertainty quantification improvements and provide outlook for workflow automation and extending the model to social, demographic and economic population characteristics.

Building morphology