Search NASA⌕ Search

SEARCH · Search NASA

Results for “Selective classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Deep learning uncertainty quantification for clinical text classification

Machine learning algorithms are expected to work side-by-side with humans in decision-making pipelines. Thus, the ability of classifiers to make reliable decisions is of paramount importance. Deep neural networks (DNNs) represent the state-of-the-art models to address real-world classification. Although the strength of activation in DNNs is often correlated with the network’s confidence, in-depth analyses are needed to establish whether they are well calibrated. In this paper, we demonstrate the use of DNN-based classification tools to benefit cancer registries by automating information extraction of disease at diagnosis and at surgery from electronic text pathology reports from the US National Cancer Institute (NCI) Surveillance, Epidemiology, and End Results (SEER) population-based cancer registries. In particular, we introduce multiple methods for selective classification to achieve a target level of accuracy on multiple classification tasks while minimizing the rejection amount—that is, the number of electronic pathology reports for which the model’s predictions are unreliable. We evaluate the proposed methods by comparing our approach with the current in-house deep learning-based abstaining classifier. Overall, all the proposed selective classification methods effectively allow for achieving the targeted level of accuracy or higher in a trade-off analysis aimed to minimize the rejection rate. On in-distribution validation and holdout test data, with all the proposed methods, we achieve on all tasks the required target level of accuracy with a lower rejection rate than the deep abstaining classifier (DAC). Interpreting the results for the out-of-distribution test data is more complex; nevertheless, in this case as well, the rejection rate from the best among the proposed methods achieving 97% accuracy or higher is lower than the rejection rate based on the DAC. We show that although both approaches can flag those samples that should be manually reviewed and labeled by human annotators, the newly proposed methods retain a larger fraction and do so without retraining—thus offering a reduced computational cost compared with the in-house deep learning-based abstaining classifier.

59 BASIC BIOLOGICAL SCIENCES↗

Non-averaged single-molecule tertiary structures reveal RNA self-folding through individual-particle cryo-electron tomography

Large-scale and continuous conformational changes in the RNA self-folding process present significant challenges for structural studies, often requiring trade-offs between resolution and observational scope. Here, we utilize individual-particle cryo-electron tomography (IPET) to examine the post-transcriptional self-folding process of designed RNA origami 6-helix bundle with a clasp helix (6HBC). By avoiding selection, classification, averaging, or chemical fixation and optimizing cryo-ET data acquisition parameters, we reconstruct 120 three-dimensional (3D) density maps from 120 individual particles at an electron dose of no more than 168 e - Å -2 , achieving averaged resolutions ranging from 23 to 35 Å, as estimated by Fourier shell correlation (FSC) at 0.5. Each map allows us to identify distinct RNA helices and determine a unique tertiary structure. Statistical analysis of these 120 structures confirms two reported conformations and reveals a range of kinetically trapped, intermediate, and highly compacted states, demonstrating a maturation folding landscape likely driven by helix-helix compaction interactions.

36 MATERIALS SCIENCE↗

Data-driven search for promising intercalating ions and layered materials for metal-ion batteries

The rise in demand for lithium-ion batteries has led to a large-scale search for electrode materials and intercalating ion species to meet the demands of next-generation energy technologies. Recent efforts largely focus on searching for cathodes that can accommodate large amounts of intercalating ions, but similar work on anodes is relatively limited. This study utilizes machine learning methods to find alternative two-dimensional (2D) materials and intercalating ions beyond Li for metal-ion batteries with high-power efficiencies. The approach first uses density functional theory (DFT) calculations to estimate the theoretical capacities and voltages of various metal ions on 2D materials. The DFT-generated data also provide insights into the local structural accommodation upon ion intercalation on various 2D materials. Significant changes to the lattice can result in irreversible changes to the bonding environments in the anode material, resulting in poor cycling stability. Next, this study develops a binding energy and structural accommodation-based classification model to screen anode materials for next-generation batteries. The classification model selects intercalating ions and 2D material pairs suitable for batteries based on the calculated voltage and volumetric changes in the 2D material upon intercalation. Finally, this study builds a regression model to accurately predict the binding energies of the various intercalating ions on 2D materials. The approach highlights the importance of different elemental and structural features for classification and regression tasks. In conclusion, the insights gained from this study on the role of involved features, such as electronegativities of the constituent ions and the presence of unfilled electronic levels, will help to streamline further studies towards the search for future layered battery materials.

36 MATERIALS SCIENCE↗

Toward memory-efficient melt pool monitoring: a classification framework using event-based imaging and sparse sensing technique

Vision sensors like CMOS and CCD cameras are often used for in-process monitoring of melt pools in laser-based additive and welding processes, but they require transferring large amounts of data and computational processing resources. Event-based neuromorphic imagery, on the other hand, detects only the change in pixel intensity, thus potentially reducing the data amount and latency. With an event imager, this study develops a framework for melt pool condition classification, including image construction, time scale selection, optimal pixel selection, and sparse classification, to achieve a highly memory-efficient scheme. These are based on sparse sensing techniques with singular value decomposition (SVD) and QR pivoting, the two fundamental matrix transformations for linear dimensionality reduction. The framework is then validated by classifying a controlled experiment by exciting various mode shapes of liquid gallium pools of varying depths (3, 6, and 8 mm). At 200 pixels, the classifier can reach overall accuracy of 75%, while at 2000 pixels (0.013% of the total possible pixels), the accuracy is nearly 90% (89.86%). At the same number of pixels, random selection can only achieve 46% and 67%, respectively. The memory savings of the sparsely sampled event data compared to a conventional imager is about 500 times. In addition to performance, implementation and limitations of the framework are also discussed.

42 ENGINEERING↗

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

Identification of integrated proteomics and transcriptomics signature of alcohol-associated liver disease using machine learning

Distinguishing between alcohol-associated hepatitis (AH) and alcohol-associated cirrhosis (AC) remains a diagnostic challenge. In this study, we used machine learning with transcriptomics and proteomics data from liver tissue and peripheral mononuclear blood cells (PBMCs) to classify patients with alcohol-associated liver disease. The conditions in the study were AH, AC, and healthy controls. We processed 98 PBMC RNAseq samples, 55 PBMC proteomic samples, 48 liver RNAseq samples, and 53 liver proteomic samples. First, we built separate classification and feature selection pipelines for transcriptomics and proteomics data. The liver tissue models were validated in independent liver tissue datasets. Next, we built integrated gene and protein expression models that allowed us to identify combined gene-protein biomarker panels. For liver tissue, we attained 90% nested-cross validation accuracy in our dataset and 82% accuracy in the independent validation dataset using transcriptomic data. We attained 100% nested-cross validation accuracy in our dataset and 61% accuracy in the independent validation dataset using proteomic data. For PBMCs, we attained 83% and 89% accuracy with transcriptomic and proteomic data, respectively. The integration of the two data types resulted in improved classification accuracy for PBMCs, but not liver tissue. We also identified the following gene-protein matches within the gene-protein biomarker panels: CLEC4M-CLC4M, GSTA1-GSTA2 for liver tissue and SELENBP1-SBP1 for PBMCs. In this study, machine learning models had high classification accuracy for both transcriptomics and proteomics data, across liver tissue and PBMCs. The integration of transcriptomics and proteomics into a multi-omics model yielded improvement in classification accuracy for the PBMC data. The set of integrated gene-protein biomarkers for PBMCs show promise toward developing a liquid biopsy for alcohol-associated liver disease.

60 APPLIED LIFE SCIENCES↗

MARVEL Hazard Evaluation ECAR-6440

MARVEL Hazards Evaluation evaluated the impacts of MARVEL operations, hazards, and postulated accidents. The hazard evaluation of MARVEL events and associated operations was performed for selection and evaluation of safety classification of systems, structures, and components (SSCs) and SSC safety functions, and for selection of design basis accidents (DBAs) applicable to the MARVEL microreactor design.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau↗

Image processing pipeline for AI-driven nanoparticle megalibrary characterization

Recent innovations have made it possible to produce megalibraries, millions of structurally and compositionally distinct nanoparticles on a chip. These megalibraries yield vast volumes of data that are impossible to analyze manually, necessitating the development of automated tools. In previous work, we created a binary classification machine learning model to select quality nanoparticle images for downstream analysis. In this work, we show that adding a custom image processing step before training can produce significantly higher-performing models in a fraction of the time and make them more robust to different image noise levels and microscope acquisition settings. The image processing pipeline proposed here effectively cleans raw nanoparticle images, enhances key features, and allows us to use much lower resolution images and simpler neural network model architectures. These features result in higher performance and significant cost savings. Experiments demonstrate superior performance relative to baseline, including an 18.2% improvement in recall and a 13.1% increase in accuracy. Given the high cost of downstream analysis, it is critical to minimize false positives, and our best-performing model reaches a precision of 95.9% and a weighted F-score of 95.1% on an unseen test set. Additionally, model training time is reduced from hours to less than a minute. We also show that, using this custom image processing pipeline, model performance is significantly improved at lower pixel resolutions compared to downsizing alone. We expect that adopting this pipeline for AI-driven automated nanoparticle characterization will allow researchers to rapidly and accurately analyze much greater volumes of data, thereby accelerating materials discovery.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Long-term measurements of ice nucleating particles at Atmospheric Radiation Measurement (ARM) sites worldwide

Ice nucleating particles (INPs) play a critical role in cloud microphysics and precipitation formation, yet long-term, spatially extensive observational datasets remain limited. Here, we present one of the most comprehensive publicly available datasets of immersion-mode INP concentrations using a single analytical method, generated through the U.S. Department of Energy's (DOE) Atmospheric Radiation Measurement (ARM) user facility. INP filter samples have been collected across a broad range of environments – including agricultural plains, Arctic coastlines, high-elevation mountain sites, marine regions, and urban areas – via fixed observatories, mobile facility deployments, and vertically-resolved tethered balloon system operations. We describe the standardized processing and quality assurance pipeline, from filter collection and processing using the Ice Nucleation Spectrometer to final data products archived on the ARM Data Discovery portal. The dataset includes both total INP concentrations and selectively treated samples, allowing for classification of biological, organic, and inorganic INP types. It features a continuous 5-year record of INP measurements from a central U.S. site, with data collection still ongoing. Seasonal and site-specific differences in INP concentrations are illustrated through intercomparisons at −10 and −20 °C, revealing distinct regional sources and atmospheric drivers. We also outline mechanisms for researchers to access existing data, request additional sample analyses, and propose future field campaigns involving ARM INP measurements. This dataset supports a wide range of scientific applications, from observational and mechanistic studies to model development, and provides critical constraints on aerosol-cloud interactions across diverse atmospheric regimes (Creamean et al., 2024, 2020b; https://doi.org/10.5439/1770816).

Creamean, Jessie M. [Colorado State Univ., Fort Co↗

Streamlined spatial and environmental expression signatures characterize the minimalist duckweed Wolffia australiana

Single-cell genomics permits a new resolution in the examination of molecular and cellular dynamics, allowing global, parallel assessments of cell types and cellular behaviors through development and in response to environmental circumstances, such as interaction with water and the light–dark cycle of the Earth. Here, we leverage the smallest, and possibly most structurally reduced, plant, the semiaquaticWolffia australiana, to understand dynamics of cell expression in these contexts at the whole-plant level. We examined single-cell-resolution RNA-sequencing data and foundWolffiacells divide into four principal clusters representing the above- and below-water-situated parenchyma and epidermis. Although these tissues share transcriptomic similarity with model plants, they display distinct adaptations thatWolffiahas made for the aquatic environment. Within this broad classification, discrete subspecializations are evident, with select cells showing unique transcriptomic signatures associated with developmental maturation and specialized physiologies. Assessing this simplified biological system temporally at two key time-of-day (TOD) transitions, we identify additional TOD-responsive genes previously overlooked in whole-plant transcriptomic approaches and demonstrate that the core circadian clock machinery and its downstream responses can vary in cell-specific manners, even in this simplified system. Distinctions between cell types and their responses to submergence and/or TOD are driven by expression changes of unexpectedly few genes, characterizingWolffiaas a highly streamlined organism with the majority of genes dedicated to fundamental cellular processes.Wolffiaprovides a unique opportunity to apply reductionist biology to elucidate signaling functions at the organismal level, for which this work provides a powerful resource.

Biochemistry & Molecular Biology↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

In real-world applications like spectrum management and interference detection, dealing with unseen electromagnetic waveforms is critical. Although some methods attempt to simulate open set data using generator models, they face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. This results in difficulties capturing distinctive features across classes, especially in dynamic scenarios where new classes emerge. To detect unseen waveforms, we propose combining time and frequency domain features with cosine similarity loss to enhance feature distinctiveness and enabling more accurate predictions. This approach efficiently captures more comprehensive information than single-domain representations or approaches without cosine loss. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10\% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗

Are light curve classification metrics good proxies for SN Ia cosmological constraining power?

Context. When selecting a light curve classifier for use as part of a photometric supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, such as the contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would eliminate the computational expense of a full cosmology forecast in the analysis pipeline design process. Aims. This study tests the assumption that light curve classification metrics are an appropriate proxy for cosmology metrics. Methods. We emulated photometric SN Ia cosmology light curve samples with controlled contamination rates of individual contaminant classes and evaluated each of them under a set of classification metrics. We then derived cosmological parameter constraints from all samples under two common analysis approaches and quantified the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results. We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are shown to be insensitive to the latter. Conclusions. Based on these findings, we discourage any exclusive reliance on light curve classification-based metrics for analysis design decisions, which (counterintuitively) include but are not limited to the classifier choice. Instead, we recommend optimising science analysis pipeline design choices using a metric of the information gained about the physical parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Chaconne: A Statistical Approach to Nonlocal Compression for Supervised Learning, Semi-Supervised Learning, and Anomaly Detection

This project developed a novel statistical understanding of compression analytics (CA), which has challenged and clarified some core assumptions about CA, and enabled the development of novel techniques that address vital challenges of national security. Specifically, this project has yielded the development of novel capabilities including 1. Principled metrics for model selection in CA, 2. Techniques for deriving/applying optimal classification rules and decision theory to supervised CA, including how to properly handle class imbalance and differing costs of misclassification, 3. Two techniques for handling nonlocal information in CA, 4. A novel technique for unsupervised CA that is agnostic with regard to the underlying compression algorithm, 5. A framework for semisupervised CA when a small number of labels are known in an otherwise large unlabeled dataset. 6. The academic alliance component of this project has focused on the development of a novel exemplar-based Bayesian technique for estimating variable length Markov models (closely related to PPM [prediction by partial matching] compression techniques). We have developed examples illustrating the application of our work to text, video, genetic sequences, and unstructured cybersecurity log files.

99 GENERAL AND MISCELLANEOUS↗