Search NASASearch

SEARCH · Search NASA

Results for “Classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Third-Party Supplier Risk Re-Classification Using Multi-Model Semantic Voting and External Web Augmentation

Risk decisions in many third-party risk management (TPRM) workflows rely on static inherent risk questionnaires (IRQ). These static forms provide a snapshot of the vendor from the business users’ perspective, as these requests are processed without cross-referencing for evidence. Consequently, responses can be misinformed or embellished with inaccuracies, thereby masking the vendor’s true risk to the enterprise. This paper presents a multi-stage verification framework to augment IRQs with web evidence and a deterministic ensemble of large language model assessors to reclassify risk. In a case study of 100 submissions previously misclassified as low risk, the proposed framework correctly identified 76% of the cases as high risk, while the existing workflow identified none. McNemar’s continuity corrected statistics of 74 were obtained with a two sided p-value of 2.65 × 10-23, indicating a significantly more effective workflow compared to the legacy model.

99 - GENERAL AND MISCELLANEOUS

Open Set Recognition for Unknown Waveform Classification

This presentation applies open set recognition to classify unknown waveforms, enabling systems to not only identify known types but also reliably detect when waveforms fall outside the training distribution. This approach enhances robustness by avoiding forced misclassification of novel or anomalous signals.

99 - GENERAL AND MISCELLANEOUS

Graph Identification of Proteins in Tomograms (GRIP-Tomo) 2.0: Topologically aware classification for proteins

Cryo-electron tomography (cryo-ET) enables structural characterization of biomolecules under near-native conditions. Existing approaches for interpreting the resulting three-dimensional volumes are computationally expensive and have difficulty interpreting density associated with small proteins/complexes. To explore alternate approaches for identifying proteins in cryo-ET data we pursued a Graph Network and topologically invariant approach. Here, we report on a fast algorithm that classifies particles by searching for nuances of evolutionarily conversed motifs and the geometrical characteristics of protein structure. GRIP-Tomo 2.0 is a machine-learning pipeline that extracts interpretable topological features of protein structures within noisy experimental backgrounds. Compared to version 1.0, the new pipeline includes three upgrades that significantly improve performance including synthetic tomogram generation simulating realistic noise, graph-based persistent feature extraction as protein fingerprints, and high-performance computing acceleration. GRIP-Tomo 2.0 achieves over 90% accuracy in classifying between proteins and noise using both real and synthetic datasets which represents a foundational step toward advancing cryo-ET workflows and empowering automated visual proteomics.

Li, Chengxuan

GMFOLD: Subgraph matching for high-throughput DNA-aptamer secondary structure classification and machine learning interpretability

Aptamers are oligonucleotide receptors that bind to their targets with high affinity. Here, we consider aptamers comprised of single-stranded DNA that undergo target-binding-induced conformational changes, giving rise to unique secondary and tertiary structures. Given a specific aptamer primary sequence, there are well-established computational tools (notably mfold) to predict the secondary structure via free energy minimization algorithms. While mfold generates secondary structures for individual sequences, there is a need for a high-throughput process whereby thousands of DNA structures can be predicted in real-time for use in an interactive setting, when combined with aptamer selections that generate candidate pools that are too large to be experimentally interrogated. We developed a new Python code for high-throughput aptamer secondary structure determination (GMfold). GMfold uses subgraph matching methods to group aptamer candidates by secondary structure similarities. We also improve an open-source code, SeqFold, to incorporate subgraph matching concepts. We represent each secondary structure as a lowest-energy bipartite subgraph matching of the DNA graph to itself. These new tools enable thousands of DNA sequences to be compared based on their secondary structures, using machine-learning algorithms. This process is advantageous when analyzing sequences that arise from aptamer selections via systematic evolution of ligands by exponential enrichment (SELEX). This work is a building block for future machine-learning-informed DNA-aptamer selection processes to identify aptamers with improved target affinity and selectivity and advance aptamer biosensors and therapeutics.

Aptamer

MARLOWE: An Untargeted Proteomics, Statistical Approach to Taxonomic Classification for Forensics

General proteomics research for fundamental science typically addresses laboratory- or patient-derived samples of known origin and composition. However, in a few research areas, such as environmental proteomics, clinical identification of infectious organisms, archeology, art/cultural history, and forensics, attributing the origin of a protein-containing sample to the organisms that produced it is a central focus. A small number of groups have approached this problem and developed software tools for taxonomic characterization and/or identification using bottom-up proteomics. Most such tools identify peptides via database search, and many rely on organism-specific peptides as markers. Our group recently introduced MARLOWE, a software tool for taxonomic characterization of unknown samples based on de novo peptide identification and signal-erosion-resistant strong peptides, which are shared peptides distributed in a taxonomy-dependent manner. In the current work, we further characterize the utility of MARLOWE using publicly available proteomics data from forensically-relevant samples. MARLOWE characterizes samples based on their protein profile, and returns ranked organism lists of potential contributors and taxonomic scores based on shared strong peptides between organisms. Overall, the correct characterization rate ranges between 44 and 100%, depending on the sample type and data acquisition parameters (with lower numbers associated with lower-quality data sets). MARLOWE demonstrates successful characterization of true contributors and close relatives, and provides sufficient specificity to distinguish certain microbial species. MARLOWE demonstrates its ability to provide insight into potential taxonomic sources for a wide range of sample types without prior assumptions about sample contents. As a result, this approach can find utility in forensic science and also broadly in bioanalytical applications that utilize proteomics approaches for taxonomic characterization.

Bacteria

Altermagnetism classification

Altermagnets are magnetic states with fully compensated spins and broken PT (PT: parity times time reversal) symmetry (i.e., spin-split bands). We classify three kinds of altermagnets in terms of broken P and T. Furthermore, strong altermagnets have spin-split bands without spin-orbit coupling (SOC), and weak altermagnets has spin-split bands only with non-zero SOC. These strong vs. weak altermagnets can be identified from the total number of symmetric spin rotation operations.

Altermagnets

Plasma confinement state classification in fusion power plants: Profile reflectometer and ensemble diagnostics

As Fusion Pilot Plants (FPPs) are increasingly viewed as within reach, many engineering challenges remain. Not many diagnostics are expected to be available in a reactor environment. Survivability, maintainability, and limited port space substantially restrict the number of FPP-relevant diagnostics. One remaining challenge is developing tools and devices to extract plasma state information necessary for controlling an FPP from a limited subset of diagnostics. This work is part of an overarching project to address this challenge. The specific diagnostic subset to be used in FPPs is still under debate. We take the approach of developing machine-learning-based tools for different significant plasma state parameters, using already known FPP-viable diagnostics. Previously we developed a plasma confinement mode classifier utilizing the Electron Cyclotron Emission (ECE) diagnostic. Here, we expand on this by developing a Profile Reflectometer (PR) based classifier with 97% test accuracy, and an ensemble model that combines the ECE and PR models into a single model, achieving 99% test accuracy.

Clark, Randall [Univ. of California, San Diego, CA

Classification of compact objects and model comparison using EOS knowledge

Nuclear theory and experiments, alongside astrophysical observations, constrain the equation of state (EOS) of supranuclear-dense matter. Conversely, knowledge of the EOS allows an improved interpretation of nuclear or astrophysical data. In this article, we use several established constraints on the EOS and the new NICER measurement of PSR J0437-4715 to comment on the nature of the primary companion in GW230529 and the companion of PSR J0514-4002E. We find that, with a probability of ≳84% and ≳68%, respectively, both objects are black holes. These likelihoods increase to above 95% when one uses GW170817’s remnant as an upper limit on the TOV mass. We also demonstrate that the current knowledge of the EOS substantially disfavors high masses and radii for PSR J⁢0030+0451, inferred recently when combining NICER with XMM-Newton background data and using particular hot-spot models. Lastly, we also use our obtained EOS knowledge to comment on measurements of the nuclear symmetry energy, finding that the large value predicted by the PREX-II measurement displays some mild tension with other constraints on the EOS.

79 ASTRONOMY AND ASTROPHYSICS

A Compound Data Poisoning Technique with Significant Adversarial Effects on Transformer-based Sentiment Classification Tasks

Transformer-based models have demonstrated much success in various natural language processing tasks. However, they are often vulnerable to adversarial attacks, such as data poisoning, which can intentionally fool the model into generating incorrect results. In this article, we present a novel, compound variant of a data poisoning attack on a transformer-based model that maximizes the poisoning effect while minimizing the scope of poisoning. Here we do so by combining the established data poisoning technique (label flipping) with a novel adversarial artifact selection and insertion technique aimed at minimizing detectability and the scope of the poisoning footprint. We find that by using a combination of these two techniques, we achieve a state-of-the-art attack success rate of approximately 90% while poisoning only 0.5% of the original training set, thus minimizing the scope and detectability of the poisoning action. These findings have the potential to advance the development of better data poisoning detection methods.

97 MATHEMATICS AND COMPUTING

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]

Hyperdimensional computing for image classification (HDC) v1.0

This is an implementation of the hyperdimensional computing technique to classify images. It consists of a python script that trains the system for a set of images from a set of images (dataset) specified by the user. This training produces hardware configuration parameters and description vectors that are then loaded into the hardware description part of the project. The hardware description consists of hardware described in Verilog (a well known language for this purpose) that is synthesizable and can be implemented in a real chip. This hardware received the training information generated by python, and then is able to accept images to produce answers for each image on which category (class) from the pre-=trained ones the image belongs to. The hardware and python training scripts are configurable and documented. The advantage of hyperdimensional computing is its robustness to errors and the easy capability for online learning (refining the training during inference slowly over time), which this implementation supports.

Michelogiannakis, Georgios [Lawrence Berkeley Nati

GRinding Automated Classification Engine

This work is an ML-driven framework for automated surface analysis of microscopy images. We create a training dataset by imaging stainless steel samples to benchmark four developed deep neural network architectures. These models, based on a YOLOv8n-cls backend, integrate image features and process metadata using various fusion methods to distinguish between acceptable and unacceptable surface finishes. This code is associated with publication "Classifying Alloy Surface Preparation Quality with Metadata-Infused Machine Learning for Rapid Alloy Discovery" for project APEX LDRD-ER (25-ERD-039)

Gongora, AldairE [Lawrence Livermore National Labo