Search NASA⌕ Search

SEARCH · Search NASA

Results for “Classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37

ChemEcho v1.0

ChemEcho is a tool that converts tandem mass spectra into embeddings used to build machine learning (ML) models with fully explainable predictions. It provides an API for transforming raw tandem mass spectral data into embeddings, along with functions for training and validating ML models. Additionally, it includes utilities for retrieving and cleaning training data. ChemEcho is broadly applicable in ML pipelines that use tandem mass spectra for a variety of tasks, such as chemical classification or bioactivity mining. While there are existing methods to generate embeddings from fragmentation data, ChemEcho's approach ensures that predictions remain interpretable, enabling experts to evaluate results and generate hypotheses about the underlying data.

Harwood, Thomas [Lawrence Berkeley National Labora↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

VHClass

The code is used to predict the taxonomic source of an antibody heavy chain sequence. The code assigns a binary label to the input set of sequences - camelid or human. This prediction is generated using a random-forest based classification algorithm which is the backbone of the code. A complementary code splits the antibody sequence into antibody features - framework regions and CDR regions.

Davis, Anastasiia↗

Genomic Language model for Annotation of Repetitive Elements (GLARE) v1.0

GLARE (Genomic Language model for Annotation of Repetitive Elements) is a tool that classifies transposable elements (TEs)—the mobile, repetitive DNA sequences that make up large fractions of eukaryotic genomes. GLARE fine-tunes the NTv3-650M genomic language model on a harmonized collection of curated TE sequences from the PanTEon and Repbase reference databases, assigning each input sequence to one of 11 orders and 32 superfamilies in a Wicker-compatible taxonomy. Features. From nucleotide FASTA input, GLARE outputs per-sequence predictions, class summaries, composition figures, and an annotated FASTA. It provides calibrated confidence scores with optional abstention and runs on CPU or GPU. Uses. GLARE serves as a classification component in genome-annotation pipelines, downstream of TE discovery, supporting genome annotation and comparative and evolutionary genomics. Advantages. GLARE is the first repeat-element classifier to leverage a pretrained genomic language model. Combined with multi-database training, this approach outperformed all nine classifiers in the PanTEon benchmark, generalized better to unseen taxonomic clades, and remained robust to sequence orientation—a common failure mode of existing tools.

Bruna, Tomas [Lawrence Berkeley National Laborator↗

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

The NASA ACTIVATE Mission

The NASA Aerosol Cloud Meteorology Interactions over the Western Atlantic Experiment (ACTIVATE) conducted 162 joint flights with two aircraft over the northwest Atlantic to study aerosol–cloud interactions (ACIs), which represent the largest uncertainty in estimating total anthropogenic radiative forcing. The combination of a high-flying King Air and low-flying HU-25 Falcon, equipped with remote sensing and in situ instruments, characterized trace gases, aerosol particles, clouds, and meteorological variables with data collected nearly simultaneously below, within, and above marine boundary layer (MBL) clouds. Flights spanning warm and cold seasons across 3 years (2020–22) provided a broad range of conditions associated with aerosol particles, cloud properties (including particle size and phase), and meteorology, ideally suited for robust ACI calculations and assessing how well models simulate a wide range of MBL clouds from stratiform to cumulus. ACTIVATE data suggest that drivers of cloud droplet number concentration N d , including aerosol particles and MBL dynamics, vary between winter and summer months with a stronger potential to convert aerosol particles into cloud droplets in winter. Models of varying complexity not only highlight some skills in simulating winter and summer cloud types but also identify challenges that still need to be addressed such as treatment of turbulence, wet scavenging, and mesoscale organization. Remote sensing advances range from new retrieval methods for N d , cloud phase classification, vertically resolved aerosol and cloud condensation nuclei number concentration, and ocean surface wind speed. This work describes these scientific and technological advances along with efforts in outreach and open data science.

aerosol indirect effect↗

Clusters of Regional Precipitation Seasonality Change in the Community Earth System Model, Version 2

Abstract The likely changes to precipitation seasonality with warming are both impactful and not well understood. This work aims to describe areas that experience similar changes to seasonal precipitation irrespective of the original underlying precipitation seasonality. We train a self-organizing map on the difference between the seasonal cycle of precipitation in the past and in a high-warming future climate as represented by the Community Earth System Model, version 2, to create regions with similar changes in precipitation seasonality. This method is applied separately over land and ocean surfaces because of the differing processes leading to precipitation over each. This method indicates that future changes in seasonal precipitation are most varied in the tropics because of a southward shift in the intertropical convergence zone. The seasonal shifts found over midlatitude oceans indicate a poleward shift in atmospheric river activity. We find a correspondence between certain land-based precipitation changes and Köppen climate classification. The seasonality of large-scale and convective precipitation is examined for each region. The relationship between the seasonal changes to precipitation and associated atmospheric processes is discussed. These processes include atmospheric rivers, the intertropical convergence zone, tropical cyclones, and monsoons.

54 ENVIRONMENTAL SCIENCES↗

Assessment of U.S. Urban Surface Temperature Using GOES-16 and GOES-17 Data: Urban Heat Island and Temperature Inequality

Abstract This study utilizes hourly land surface temperature (LST) data from the Geostationary Operational Environmental Satellite (GOES) to analyze the seasonal and diurnal characteristics of surface urban heat island intensity (SUHII) across 120 largest U.S. cities and their surroundings. Distinct patterns emerge in the classification of seasonal daytime SUHII and nighttime SUHII. Specifically, the enhanced vegetation index (EVI) and albedo (ALB) play pivotal roles in influencing these temperature variations. The diurnal cycle of SUHII further reveals different trends, suggesting that climate conditions, urban and nonurban land covers, and anthropogenic activities during nighttime hours affect SUHII peaks. Exploring intracity LST dynamics, the study reveals a significant correlation between urban intensity (UI) and LST, with LST rising as UI increases. Notably, populations identified as more vulnerable by the social vulnerability index (SVI) are found in high UI regions. This results in discernible LST inequality, where the more vulnerable communities are under higher LST conditions, possibly leading to higher heat exposure. This comprehensive study accentuates the significance of tailoring city-specific climate change mitigation strategies, illuminating LST variations and their intertwined societal implications.

Environmental Sciences & Ecology↗

Interregional Electricity Transmission in the United States: Realized Savings and Opportunities for Increased Value, 2014 to 2023

Interregional electricity trade in the United States reduces total energy costs, but falls short of maximizing the economic potential of existing transmission infrastructure. This analysis of thirty-two U.S. electricity system interfaces crossing grid or market seams finds an average achieved cost savings of at least $1,226 million per year. This positive impact is offset by an average annual cost of uneconomic interchange of $551 million. Additional value of at least $238 million per year could be enabled by increasing utilization. Notably, uneconomic flows and low utilization rates can persist even when the price spread between regions is large. These conclusions are derived from historical (2014–2023) data on the magnitude and direction of interregional price spreads, the magnitude and direction of net energy transfers, and how these variables coincide, but do not account for all system constraints. These findings motivate the development and deployment of solutions that improve interregional transmission operations, in parallel with robust transmission infrastructure planning. JEL Classification: Q40 Energy: General

Q40 Energy: General↗

An Updated Synthesis of the Projectile Point Typology and Chronology of Eastern Idaho

The eastern Idaho archaeological record is a unique confluence of the Great Basin, Columbia Plateau, and Great Plains, resulting in a diverse and complex projectile point sequence. This study presents a revised typology and chronology of projectile points in the region, grounded in a comprehensive review of over 750 diagnostic examples from 16 stratified sites and 110 associated radiocarbon dates. We reevaluate existing classifications and propose a revised chronological framework with regionally appropriate types. This study provides a thorough background to eastern Idaho projectile points and contributes to broader discussions of projectile point typology and chronology in the Desert West, offering a robust tool for future archaeological research in the region.

99 - GENERAL AND MISCELLANEOUS↗

Plasma proteomic biomarkers of physical frailty in heart failure: a propensity score matched discovery-based pilot study

Background: Physical frailty is highly prevalent in heart failure (HF), but we lack an understanding of the underlying pathophysiology. Proteomics evaluation of plasma samples may elucidate potential mechanisms and biomarkers of physical frailty in HF. We aimed to identify plasma proteomic biomarkers that are differentially expressed between physically frail and non physically frail adults with HF. Methods: This was a secondary analysis of a subset of data and plasma samples from a study of frailty among patients with New York Heart Association (NYHA) Functional Classification I-IV HF. Physical frailty was measured using the Frailty Phenotype Criteria. Propensity score matching was used to match pairs of physically frail (n = 20) vs. non-physically frail (n = 20) patients on clinical characteristics. Plasma samples were processed using a sensitive liquid chromatography mass spectrometry platform, utilizing a multiplexed tandem mass tag-labeled quantitative proteomics approach. Differentially expressed proteins were quantified individually using paired t tests with associated log fold change of 0.3 and Fisher’s combined p values. Results: The sample (n = 40) was 62.8±16.9 years old, 58% female, and 55% NYHA Class III/IV. Proteomics analysis revealed 7 proteins differentially expressed using full differential criteria: matrix metalloproteinase-14 was downregulated in frailty, and copine-1, low affinity immunoglobulin gamma Fc region receptor III-A and III-B, probable non-functional immunoglobulin kappa variable 2D-24, glutathione S-transferase Mu 1, and argininosuccinate lyase were upregulated in frailty. Conclusions: Proteomic biomarkers related to the immune system, stress response, and detoxification were differentially expressed between physically frail and non-physically frail adults with HF.

Biomarkers↗

Risk of longer-term endocrine and metabolic conditions in the Deepwater Horizon Oil Spill Coast Guard cohort study – five years of follow-up

Abstract Introduction Long-term endocrine and metabolic health risks associated with oil spill cleanup exposures are largely unknown, despite the endocrine-disrupting potential of crude oil and oil dispersant constituents. We aimed to investigate risks of longer-term endocrine and metabolic conditions among U.S. Coast Guard (USCG) responders to the Deepwater Horizon (DWH) oil spill. Methods Our study population included all active duty DWH Oil Spill Coast Guard Cohort members ( N = 45,224). Self-reported spill exposures were ascertained from post-deployment surveys. Incident endocrine and metabolic outcomes were defined using International Classification of Diseases (9th Revision) diagnostic codes from military health encounter records up to 5.5 years post-DWH. Using Cox proportional hazards regression, we estimated adjusted hazard ratios (aHR) and 95% confidence intervals (CIs) for various incident endocrine and metabolic diagnoses (2010–2015, and separately during 2010–2012 and 2013–2015). Results The mean baseline age was 30 years (~ 77% white, ~ 86% male). Compared to non-responders ( n = 39,260), spill responders ( n = 5,964) had elevated risks for simple and unspecified goiter (aHR = 2.09, 95% CI: 1.29–3.38) and disorders of lipid metabolism (aHR = 1.09, 95% CI: 1.00–1.18), including its subcategory other and unspecified hyperlipidemia (aHR = 1.10, 95% CI: 1.01–1.21). The dysmetabolic syndrome X risk was elevated only during 2010–2012 (aHR = 2.07, 95% CI: 1.22–3.51). Responders reporting ever ( n = 1,068) vs. never ( n = 2,424) crude oil inhalation exposure had elevated risks for disorders of lipid metabolism (aHR = 1.24, 95% CI: 1.00–1.53), including its subcategory pure hypercholesterolemia (aHR = 1.71, 95% CI: 1.08–2.72), the overweight, obesity and other hyperalimentation subcategory of unspecified obesity (aHR = 1.52, 95% CI: 1.09–2.13), and abnormal weight gain (aHR = 2.60, 95% CI: 1.04–6.55). Risk estimates for endocrine/metabolic conditions were generally stronger among responders reporting exposure to both crude oil and dispersants (vs. neither) than among responders reporting only oil exposure (vs. neither). Conclusion In this large cohort of active duty USCG responders to the DWH disaster, oil spill cleanup exposures were associated with elevated risks for longer-term endocrine and metabolic conditions.

Denic-Roberts, Hristina↗

NANO.PTML model for read-across prediction of nanosystems in neurosciences. computational model and experimental case of study

Abstract Neurodegenerative diseases involve progressive neuronal death. Traditional treatments often struggle due to solubility, bioavailability, and crossing the Blood-Brain Barrier (BBB). Nanoparticles (NPs) in biomedical field are garnering growing attention as neurodegenerative disease drugs (NDDs) carrier to the central nervous system. Here, we introduced computational and experimental analysis. In the computational study, a specific IFPTML technique was used, which combined Information Fusion (IF) + Perturbation Theory (PT) + Machine Learning (ML) to select the most promising Nanoparticle Neuronal Disease Drug Delivery (N2D3) systems. For the application of IFPTML model in the nanoscience, NANO.PTML is used. IF-process was carried out between 4403 NDDs assays and 260 cytotoxicity NP assays conducting a dataset of 500,000 cases. The optimal IFPTML was the Decision Tree (DT) algorithm which shown satisfactory performance with specificity values of 96.4% and 96.2%, and sensitivity values of 79.3% and 75.7% in the training (375k/75%) and validation (125k/25%) set. Moreover, the DT model obtained Area Under Receiver Operating Characteristic (AUROC) scores of 0.97 and 0.96 in the training and validation series, highlighting its effectiveness in classification tasks. In the experimental part, two samples of NPs (Fe 3 O 4 _A and Fe 3 O 4 _B) were synthesized by thermal decomposition of an iron(III) oleate (FeOl) precursor and structurally characterized by different methods. Additionally, in order to make the as-synthesized hydrophobic NPs (Fe 3 O 4 _A and Fe 3 O 4 _B) soluble in water the amphiphilic CTAB (Cetyl Trimethyl Ammonium Bromide) molecule was employed. Therefore, to conduct a study with a wider range of NP system variants, an experimental illustrative simulation experiment was performed using the IFPTML-DT model. For this, a set of 500,000 prediction dataset was created. The outcome of this experiment highlighted certain NANO.PTML systems as promising candidates for further investigation. The NANO.PTML approach holds potential to accelerate experimental investigations and offer initial insights into various NP and NDDs compounds, serving as an efficient alternative to time-consuming trial-and-error procedures.

60 APPLIED LIFE SCIENCES↗

When does global attention help: a unified empirical study on atomistic graph learning

Graph neural networks (GNNs) are widely used as surrogates for costly experiments and first-principles simulations to study the behavior of compounds at atomistic scale, and their architectural complexity is constantly increasing to enable the modeling of complex physics. While most recent GNNs combine more traditional message passing neural networks (MPNNs) layers to model short-range interactions with more advanced graph transformers (GTs) with global attention mechanisms to model long-range interactions, it is still unclear when global attention mechanisms provide real benefits over well-tuned MPNN layers due to inconsistent implementations, features, or hyperparameter tuning. We introduce the first unified, reproducible benchmarking framework–built on HydraGNN–that enables seamless switching among four controlled model classes: MPNN, MPNN with chemistry/topology encoders, GPS-style hybrids of MPNN with global attention, and fully fused localglobal models with encoders. Using seven diverse open-source datasets for benchmarking across regression and classification tasks, we systematically isolate the contributions of message passing, global attention, and encoder-based feature augmentation. Our study shows that encoder-augmented MPNNs form a robust baseline, while fused localglobal models yield the clearest benefits for properties governed by long-range interaction effects. We further quantify the accuracycompute trade-offs of attention, reporting its overhead in memory. Together, these results establish the first controlled evaluation of global attention in atomistic graph learning and provide a reproducible testbed for future model development.

Equivariant graph neural networks↗

Identification of shared viral sequences in peat moss metagenomes reveals elements of a possible Sphagnum core virome

Viruses are an understudied component of plant microbiomes. Identifying viruses that are shared between individual plants, or members of the “core virome”, could reveal stable viral populations with the potential to modulate the composition and function of the microbiome. Here, we examined the virome associated with Sphagnum mosses, a keystone species that has direct influence over the fate of peatland carbon stores. We analyzed bulk metagenomes and metatranscriptomes generated from Sphagnum field samples collected over a ten-month period to identify virus-like sequences shared among plants. Individual Sphagnum samples harbored distinct DNA and RNA viromes where only a small percentage (< 1%) of the total number of identified viral contigs were shared among all samples. Based on taxonomic classification, the shared viral contigs represent bacterial viruses, or phage (Caudoviricetes), as well as viruses of eukaryotes, namely nucleocytoplasmic large DNA viruses (Nucleocytoviricota) and RNA viruses (Riboviria). We linked the shared phage-like contigs to viral regions within sequenced genomes of bacterial taxa that are members of the Sphagnum core microbiome, suggesting that these contigs represent temperate phage or degraded prophage. The putative nucleocytoplasmic large DNA viruses and RNA viruses were phylogenetically diverse and showed sequence similarity to viruses associated with a broad range of hosts and environmental sources. The identification of shared viral contigs suggested that, despite the compositional heterogeneity between samples, Sphagnum mosses may harbor a core virome. Future work validating the presence of the core virome is warranted as it may aid in understanding how persistent viruses impact microbiome ecology and symbiont evolution within this climatically relevant keystone species.

Metagenomics↗