Search NASASearch

SEARCH · Search NASA

Results for “Modal”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES

A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties Prediction

Drug discovery is a complex and challenging process, requiring the optimization of candidate compounds to identify those with the potential to become safe and effective drugs. Predicting molecular properties is an indispensable step in the drug discovery pipeline. Traditionally, this process is costly and time-intensive, involving multiple rounds of experiments and clinical trials, rendering it impractical for every candidate compound. Deep learning techniques have emerged as a promising approach to drug discovery to reduce the cost and time required to identify novel drugs. However, prevalent research in deep learning models focused on predicting molecular properties has primarily fixated on single-modal models, which utilize a single modality of data, neglecting the potential benefits of combining different data modalities. To overcome this limitation, we introduce MRL-Mol: a deep \textbf{M}ultimodal \textbf{R}epresentation \textbf{L}earning framework for accurate \textbf{Mol}ecular properties prediction. MRL-Mol harnesses three data modalities: sequence, graph, and image, augmenting the depth of comprehension. Leveraging a large-scale unlabeled dataset~($\sim$1M unique molecules), we pretrain MRL-Mol to extract inter- and intra-modal information. Our study demonstrates the superior performance of MRL-Mol in predicting molecular properties across six benchmark datasets, including both classification and regression tasks. Notably, MRL-Mol outperforms other state-of-the-art molecular properties prediction models. These findings suggest that by combining information from multiple data modalities, MRL-Mol can comprehend molecules better than single-modal deep learning models and identify molecular properties with better accuracy.

Yang, Yuxin

A comparative study of multimodal data fusion strategies for planetary spectroscopy

Integrating heterogeneous data sources can improve scientific inference when different modalities capture complementary information, but doing so is challenging in high-dimensional, small-sample settings. In spectroscopy for planetary exploration, Laser-Induced Breakdown Spectroscopy (LIBS), Raman Spectroscopy (Raman), Visible Infrared Spectroscopy (VISIR), and Mid-Infrared Spectroscopy (MIR) each examine different aspects of composition and mineralogy, raising fundamental questions about when and how data fusion improves predictive performance. Using a Mars-relevant set of geologic standards with measurements from all four modalities, we present a rigorous systematic evaluation of four data fusion strategies: low-level (data) fusion, mid-level (feature) fusion, high-level (decision) fusion, and residual-boosting (sequential) fusion. We assess performance in predicting oxide composition via nested cross-validation and corrected significance testing to evaluate whether data fusion improves upon single-modality baselines. We show that data fusion does not uniformly improve accuracy, and that observed gains are modest, oxide-dependent, and sensitive to modality and model structure. To move beyond aggregate accuracy metrics, we use model coefficients, permutation importance, and residual gain analysis to examine how the fusion models weight individual modalities and to identify patterns of apparent complementarity or redundancy. Though focused on spectroscopy for planetary exploration, our framework for data fusion evaluation and interpretation extends to other scientific domains with heterogeneous and scarce data and provides a principled approach evaluating data fusion strategies, interpreting modality contributions, and understanding tradeoffs among data fusion strategies.

97 MATHEMATICS AND COMPUTING

Influence of Numerical Modeling Approaches on Damped Behavior of Flexible Beams: Preprint

Composites structures are widely used in aerospace and wind energy applications for their excellent stiffness and strength-to-weight properties. In these structures, structural damping is critical to predict vibration amplitudes, performance, and reliability. Structural damping is of particular interest for slender wings, rotorcraft blades, and wind turbine blades that can exhibit complex vibration phenomena and are frequently modeled with geometrically exact beam theory (GEBT). Standard approaches of stiffness proportional or modal damping merely assign user defined values and cannot predict damping behavior. This work compares stiffness proportional damping to two more advanced damping approaches: modal strain energy and Prony series. The modal strain energy approach uses a sectional analysis tool to calculate the beam stiffness and postprocess internal stresses from GEBT simulations. The internal stresses are then used to calculate modal damping factors. The Prony series is implemented within GEBT to directly model viscoelastic behavior of the composites. These approaches are compared by modeling the evolution of the damping factors of a realistic flexible wind turbine blade with varying rotational speed. Discrepancies between the approaches suggest areas for future modeling development, but differences in nonlinear damping values are less than current uncertainties about the magnitude of structural damping.

17 WIND ENERGY

A unified large language model–based framework for heterogeneous PV image diagnosis

With advances in imaging technologies, modern photovoltaic (PV) systems generate large volumes of heterogeneous image data, including visible, electroluminescence (EL), and infrared (IR) images. Existing PV image analysis models, particularly deep learning approaches, are typically task-specific and lack cross-modality generalization. To address this limitation, this paper proposes an open-source large language model (LLM)–based unified framework for heterogeneous PV image diagnostics. Through task-aware diagnostic prompting, the framework enables analysis of visible, EL, and IR images within a single pipeline, supporting both zero-shot and few-shot inference and binary and multiclass classification. It is compatible with state-of-the-art multimodal LLMs, including ChatGPT, Gemini, Claude, Qwen, and CLIP. The framework is evaluated on PV module condition classification (clean, soiling, snow, hail, and bird droppings) using visible images, cell crack detection using EL images, and hotspot detection using IR images. GPT-5.1 in few-shot mode achieves the best performance, with classification accuracy exceeding 97.3%. Open-source models such as Qwen and CLIP also deliver competitive results on visible images (around 90% accuracy), though their performance is more limited on EL and IR modalities. On the full ELPV dataset, the framework achieves 83.5% zero-shot accuracy, within 2.8% of the supervised CNN baseline, confirming scalability to larger benchmarks. Practical aspects such as reproducibility, response latency, and confidence estimation are systematically analyzed. The framework operates across PV image modalities without modality- or task-specific training, making it well suited as a rapid pre-screening tool to support downstream detailed diagnostics. A benchmark dataset of diverse labeled PV images is also released.

Li, Baojie

Co-orchestration of multiple instruments to uncover structure–property relationships in combinatorial libraries

The rapid growth of automated and autonomous instrumentation brings forth opportunities for the co-orchestration of multimodal tools that are equipped with multiple sequential detection methods or several characterization techniques to explore identical samples. This is exemplified by combinatorial libraries that can be explored in multiple locations via multiple tools simultaneously or downstream characterization in automated synthesis systems. In co-orchestration approaches, information gained in one modality should accelerate the discovery of other modalities. Correspondingly, an orchestrating agent should select the measurement modality based on the anticipated knowledge gain and measurement cost. Herein, we propose and implement a co-orchestration approach for conducting measurements with complex observables, such as spectra or images. The method relies on combining dimensionality reduction by variational autoencoders with representation learning for control over the latent space structure and integration into an iterative workflow via multi-task Gaussian Processes (GPs). This approach further allows for the native incorporation of the system's physics via a probabilistic model as a mean function of the GPs. We illustrate this method for different modes of piezoresponse force microscopy and micro-Raman spectroscopy on a combinatorial Sm-BiFeO3 library. However, the proposed framework is general and can be extended to multiple measurement modalities and arbitrary dimensionality of the measured signals.

47 OTHER INSTRUMENTATION

Rapid monitoring of fermentations: a feasibility study on biological 2,3-butanediol production

2,3-butanediol (2,3-BDO) is an economically important platform chemical that can be produced by the fermentation of sugars using an engineered strain of Zymomonas mobilis . These fermentations require continuous monitoring and modification of fermentation conditions to maximize 2,3-BDO yields and minimize the production of the undesired coproducts glycerol and acetoin. Because of the time required for sampling and off-line chromatographic measurement of fermentation samples, the ability of fermentation scientists to modify fermentation conditions in a timely manner is limited. The goal of this study was to test if near-infrared spectroscopy (NIRS) along with multivariate statistics could reduce the time needed for this analysis and enable real-time monitoring and control of the fermentation. In this work we developed partial least squares (PLS) calibration models to predict the concentrations of glucose, xylose, 2,3-BDO, acetoin, and glycerol in fermentations via NIRS using two different spectrometers and two different spectroscopy modalities. We first evaluated the feasibility of rapid NIRS monitoring through experiments where we measured the signals from each analyte of interest and built NIRS-based PLS models using spectra from synthetic samples containing uncorrelated concentrations of these analytes. All analytes showed unique spectral signatures, and this initial modeling showed that all analytes could be detected simultaneously. We then began work with samples from laboratory fermentation experiments and tested the feasibility of regression model development across two spectral collection modalities (at-line and on-line) and two instruments: a laboratory-grade instrument and a low-cost instrument with a more limited spectral range. All modalities showed promise in the ability to monitor Z. mobilis fermentations of glucose and xylose to 2,3-BDO. The low-cost instrument displayed a lower signal-to-noise ratio than the laboratory-grade instrument, which led to comparatively lower performance overall, but still provided sufficient accuracy to monitor fermentation trends. While the ease of use of on-line monitoring systems was favored as compared to at-line systems due to the lack of sampling required and potential for automated process control, we observed some decrease in performance due to the additional complexity of the sample matrix. We have demonstrated that NIRS combined with multivariate analysis can be used for at-line and on-line monitoring of the concentrations of glucose, xylose, 2,3-BDO, acetoin, and glycerol during Z. mobilis fermentations. The decrease in signal-to-noise ratio when using a low-cost spectrometer led to greater prediction error than the laboratory-grade spectrometer for at-line monitoring. The on-line monitoring modality showed great promise for real time process control via NIRS.

09 BIOMASS FUELS

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada

Peak2Patch: High-Fidelity Functional Group Identification through Attention-Based Fusion of Infrared and Mass Spectra

Identifying molecular structure based on spectroscopic readings is a key task in a variety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless require expert-level knowledge to decode. Machine learning has emerged as a potential solution for automating structure prediction from chemical spectra; however, current approaches generally focus on single sensor modalities, neglecting to leverage the complementary information contained within differing spectra. In this paper, we introduce Peak2Patch, a novel approach to fusion-enhanced prediction of functional groups from IR and mass spectra. First, we perform a detailed comparison of backbone networks for encoding both sparse mass spectra and dense IR spectra and demonstrate the superior performance of transformer neural networks over current state-of-the-art convolutional neural networks. Second, we evaluate three broad categories of fusion: early (raw feature), middle (deep feature), and late (decision) fusion, demonstrating the potential of a deep feature fusion-based approach. Lastly, we present Peak2Patch, our attention-based fusion scheme, which leverages cross-attention to mix features between encoded tokens of the two modalities. We validate our approach on a publicly available multimodal spectroscopic data set of 790k simulated molecules, demonstrating a large improvement in functional group prediction over both the previous state-of-the-art and our own strong single-modal baselines.

Jacobson, Philip [Sandia National Laboratories (SN

INR-TEM: Robust cavity detection in multifocus TEM images via implicit neural representations

When characterizing materials using transmission electron microscopy (TEM) images, detecting and quantifying small features in microstructures, such as cavities, pose significant challenges. Off-the-shelf object detection models, including YOLOv8, show considerable performance degradation, particularly when images vary in resolution and the objects of interest possess a low percentage of the total image region of interest. In this study, we introduce a novel detection pipeline that incorporates an implicit neural representation (INR)-based detection method, INR-TEM, and two-modality imaging (e.g., under-focused and over-focused images typically acquired during materials characterization) to improve object detection performance. The INR-TEM method incorporates a pixel-wise prediction principle inspired by pixel-wise centerness weighting. INR-TEM demonstrates superior robustness to resolution variability, maintaining high detection accuracy even at low image resolutions compared to YOLOv8. To leverage INR-TEM effectively in real-world two-modality characterization applications, we further integrate a two-stage motion correction pipeline designed explicitly for aligning multifocus TEM images. The alignment process, comprising keypoint (based on scale-invariant feature transform, SIFT) and intensity matching, significantly mitigates the adverse effects of perceived motion-induced image degradation during through-focal TEM imaging, directly enhancing INR-TEM’s detection capability over conventional single-focus images. Our integrated INR-TEM cavity detection framework notably improves performance across various cavity sizes, outperforming off-the-shelf YOLOv8 detections that rely on a single image modality.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

A case study in contrastive learning information combination: Application to technical forensics of additive manufacturing filament source identification

Combination of information from disparate data sources into a single decision is a core challenge in many fields, including the field of technical forensics. Technical forensics (TF) utilizes technical characterization of questioned samples to determine properties of that sample; these properties are then used to infer information of forensic interest, such as provenance, age, or attribution. TF is utilized in traditional forensic applications, such as the attribution of material fragments from an explosive, and in nuclear forensic applications, such as the attribution of actinides which have been interdicted out of regulatory control. The challenge of combining information from disparate sources, described alternately by many terms including “Data Fusion” and “Data Integration”, is exacerbated in the technical forensics domain due to at least two factors: the challenge of interpreting each information source singularly, and the relatively small data set sizes available. Extensive literature exists attempting to combine technical forensics information sources, both in manual and automated processes. These attempts are often bespoke to the specific information sources (such as the bi-, tri-, or quad-isotope chart (Moody, Grant, and Hutcheon 2005)), with some emerging examples of simple early- and late- fusion (, respectively). Simultaneous to the information combination efforts described in the previous paragraph, the field of natural language processing attempted (and largely succeeded) in combining information from multiple non-technical information sources. The ecosystem of “multi-modal” language models, which can take text and images as input, and generate text and images as output, became large and diverse by 2025 (Khan et al. 2025). In a generalized sense, many of these methods are trained by learning neural networks which can convert raw text or images into a vector of numbers describing the text or image, hereafter called “embeddings” and the neural networks performing the conversion are called “embedders”. By using a separate embedder for text and images, finding coincident text and images (such as images with their captions), and optimizing the parameters of the embedders such that the embeddings for the text and the image are similar, the field has found a bridge between text and images (Girdhar et al. 2023). It is the contention of the authors of this report that this insight is not limited to text and images but instead can be extended to any modality which can be found coincidently. The subject of the rest of this report is the application of this method to example multi-modal technical forensic data. Some details about the data used in this report are not appropriate for this report, and are included in a companion report (PNNL-38669).

36 MATERIALS SCIENCE

Illuminating the Material World: Autonomous Microscopy to Understand Order, Disorder, and Everything In Between

Artificial intelligence (AI) holds immense promise for revolutionizing microscopy, yet its widespread adoption has been hindered by challenges ranging from user inexperience to limited model transferability and difficulties in operationalizing machine learning. This presentation showcases our approach to developing practical autonomy for materials discovery, aiming to accelerate the integration of AI into everyday microscopy workflows. As shown in Fig. 1, I will focus on three key areas: understanding order-disorder transitions, quantifying point defects, and achieving truly device-scale microscopy. First, I will demonstrate the power of multi-modal knowledge graphs for integrating diverse microscopy data. By combining imaging, spectroscopy, and diffraction data, these graphs provide a holistic view of material behavior, capturing the intricate relationships between different modalities [1,2]. I will present a case study on how these models illuminate the structural and chemical changes associated with irradiation in oxide thin films, revealing critical insights for designing materials for extreme environments like spaceflight and nuclear energy. Specifically, I will show how multi-modal analysis clarifies the evolution of order-disorder transitions under irradiation, a key factor influencing material performance in these applications. Next, I will address the challenge of quantifying point defects in 2D materials. We demonstrate the application of computer vision and transfer learning to accurately identify and classify various defect types, such as vacancies and substitutional atoms, and to quantify their concentrations. This information is crucial for understanding and tailoring the properties of 2D materials for applications in electronics, optoelectronics, and catalysis. For example, I will show how our models can characterize the topological distribution of point defects in MXene transition metal carbides, providing valuable insights for optimizing their performance in energy storage and separation science. Finally, I will discuss our progress toward autonomous device-scale microscopy [3,4]. We are fundamentally redesigning electron microscopes around the principles of machine reasoning, enabling automation beyond basic tasks like sample navigation and data acquisition to include sophisticated experimental design. This approach paves the way for truly reproducible and massively scaled analysis campaigns. I will emphasize the importance of autonomous microscopy platforms for high-throughput materials discovery and characterization, facilitating the rapid screening of materials for a broad range of applications and accelerating the development of next-generation technologies.

36 MATERIALS SCIENCE

MCPC Characterization Round 1

Characterization data for the MCPC LDRD Agile investment is collected in multiple rounds. Contained herein is data from Round 1 in which eleven distinct samples and one control sample are characterized by five modalities: EBSD, HARDNESS, OPTICAL, SEM, and ULTRASONIC. Data is organized in subdirectories for the five modalities and each one modality containing subdirectories for the datasets and other information.

36 MATERIALS SCIENCE

MCPC Characterization Round 2

Characterization data for the MCPC LDRD Agile investment is collected in multiple rounds. Contained herein is data from Round 2 in which eleven distinct samples and one control sample are characterized by five modalities: EBSD, HARDNESS, OPTICAL, SEM, and ULTRASONIC. Data is organized in subdirectories for the five modalities and each one modality containing subdirectories for the datasets and other information.

316 stainless steel

MCPC Characterization Round 3

Characterization data for the MCPC LDRD Agile investment is collected in multiple rounds. Contained herein is data from Round 3 in which eleven distinct samples and one control sample are characterized by five modalities: EBSD, HARDNESS, OPTICAL, SEM, and ULTRASONIC. Data is organized in subdirectories for the five modalities and each one modality containing subdirectories for the datasets and other information.

316 stainless steel

Original images, ground truth annotations, and precited masks from U-NET for a fungal species x nitrogen experiment

This document describes the “MicroVision-MV003” dataset. This folder contains three subfolders: “originals”, “predicted_masks”, and “gt_masks”. The “originals” folder contains 720 original images acquired using two imaging modalities: 360 images from overhead imaging and 360 images from transmission (raw) imaging. Naming scheme:YYYY-MM-dd__plate_strain_nitrogen-level_replication_experiment_imaging-modalityYYYY-MM-dd: 2024-12-20, 2024-12-21, 2024-12-22,2024-12-23,2024-12-24,2024-12-25,2024-12-26,2024-12-27,2024-12-28,2024-12-29,2024-12-30,2024-12-31.plate: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30.strain: FG, LE.nitrogen-level: N-1, N-10, N-100.replication: 1, 2, 3, 4, 5.experiment: MV.003.imaging-modality: overhead, transmission (raw). For example, the image2024-12-20__1_FG_N-10_3_MV.003_overheadwas taken on 2024-12-20. The plate number is 1, the fungal strain is FG, the nitrogen treatment is N-10, and the replication number is 3 for experiment MV.003, acquired using the overhead imaging modality. The “predicted_masks” folder contains masks generated by a trained U-Net model for the transmission imaging modality.The “gt_masks” folder contains 326 human hand-traced masks that serve as ground truth.

09 BIOMASS FUELS

A Dynamic Hierarchical Attention Framework for Multimodal Malware Detection

The increasing use of Android in the worldwide mobile ecosystem has come along with a significant increase in advanced malware, highlighting the critical necessity for efficient, scalable, and adaptable detection systems. Despite recent advancements in machine learning improving malware detection, the majority of current solutions are limited to one, two, or three data modalities, hence neglecting the comprehensive behavioral spectrum of contemporary multi-vector threats. This thesis presents the first comprehensive multimodal framework for Android malware detection, which combines textual, time-series (temporal), graph-based (structural), and visual information using an innovative hierarchical attention mechanism and Dynamic Fusion Controller (DFC). Our methodology consistently classifies and processes modalities as either sequential or structural, facilitating content-adaptive weighting and resilient cross-modal representation learning. We advance the implementation of cutting-edge time series techniques, such as MiniRocket, for malware detection, hence creating new opportunities for temporal analysis in cybersecurity. Comprehensive experimental assessment shows that our framework performs exceptionally well, with 99.46% classification accuracy and 97.15% detection accuracy, significantly outperforming existing approaches through effective multimodal integration and hierarchical attention mechanisms.

Nazmin, Tamanna

Human perception of ionizing radiation

Here, in this work, we address the question of whether humans can perceive ionizing radiation. We conducted a thorough review of the clinical and experimental literature related to ionizing radiation, with a focus on its acute effects. Specifically, we examined the three domains of X-ray perception found in animals (abdominal, olfactory, and retinal), which led us to instances of ionizing radiation-induced hearing and taste sensory phenomena in humans thus suggesting that humans can perceive X-rays across various sensory modalities via multiple mechanisms. We also analyzed literature to understand the mechanisms associated with reported symptoms, this led us to the concept of radiomodulation, an understudied modulatory effect of sub-ablative ionizing radiation doses on neurons. Based on this review of the literature we propose the hypothesis that a significant radiomodulation mechanism is the formation of reactive oxygen species from radiolysis which activates immune and sensory signal transduction mechanisms specifically related to the redox activity in TRP and K+ channels. Additionally, we find evidence to support the previous claims of perception stemming from Cherenkov radiation and ozone production which are perceived using canonical sensory modalities. Finally, for we provide a concise summary of the applications of ionizing radiation in clinical imaging and therapy, as well as prospects for future developments of radiation technologies for biomedical and fundamental research.

Ionizing radiation