Search NASASearch

SEARCH · Search NASA

Results for “Deep Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Accurate segmentation of localized corrosion in structural alloys via deep learning

This study presents a deep learning-based approach for the automated segmentation of corrosion damage in scanning electron microscopy (SEM) images. The proposed method enables rapid and accurate segmentation of corrosion features in these SEM images, making it highly suitable for real-time applications such as automated microscopy. Specifically, a dedicated corrosion segmentation database tailored for this task is constructed. The newly constructed dataset, alongside data from two public databases, are employed to jointly train a deep learning-based model modified with a texture refinement module. Compared to the same model without the texture refinement module, the refined model substantially enhances the efficacy and efficiency of corrosion segmentation. Furthermore, the methodology developed here is extendable to segmentation tasks for other materials with similar resolution, texture, and contrast characteristics, thereby paving the way for accelerated and automated analysis in corrosion science and beyond.

Artificial Intelligence

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif

Deep learning for time series forecasting: a survey of recent advances

Time series forecasting plays a critical role in numerous real-world applications, such as finance, healthcare, transportation, and scientific computing. In recent years, deep learning has become a powerful tool for modeling complex temporal patterns and improving forecasting accuracy. This survey provides an overview of recent deep learning approaches for time series forecasting, involving various architectures including RNNs, CNNs, GNNs, transformers, large language models, MLP-based models, and diffusion models. We first identify key challenges in the field, such as temporal dependency, efficiency, and cross-variable dependency, which drive the development of forecasting techniques. Then, the general advantages and limitations of each architecture are discussed to contextualize their adaptation in time series forecasting. Furthermore, we highlight promising design trends like multi-scale modeling, decomposition, and frequency-domain techniques, which are shaping the future of the field. This paper serves as a compact reference for researchers and practitioners seeking to understand the current landscape and future trajectory of deep learning in time series forecasting.

97 MATHEMATICS AND COMPUTING

Integrating multi-modal remote sensing, deep learning, and attention mechanisms for yield prediction in plant breeding experiments

In both plant breeding and crop management, interpretability plays a crucial role in instilling trust in AI-driven approaches and enabling the provision of actionable insights. The primary objective of this research is to explore and evaluate the potential contributions of deep learning network architectures that employ stacked LSTM for end-of-season maize grain yield prediction. A secondary aim is to expand the capabilities of these networks by adapting them to better accommodate and leverage the multi-modality properties of remote sensing data. In this study, a multi-modal deep learning architecture that assimilates inputs from heterogeneous data streams, including high-resolution hyperspectral imagery, LiDAR point clouds, and environmental data, is proposed to forecast maize crop yields. The architecture includes attention mechanisms that assign varying levels of importance to different modalities and temporal features that, reflect the dynamics of plant growth and environmental interactions. The interpretability of the attention weights is investigated in multi-modal networks that seek to both improve predictions and attribute crop yield outcomes to genetic and environmental variables. This approach also contributes to increased interpretability of the model's predictions. The temporal attention weight distributions highlighted relevant factors and critical growth stages that contribute to the predictions. The results of this study affirm that the attention weights are consistent with recognized biological growth stages, thereby substantiating the network's capability to learn biologically interpretable features. Accuracies of the model's predictions of yield ranged from 0.82-0.93 R 2 ref in this genetics-focused study, further highlighting the potential of attention-based models. Further, this research facilitates understanding of how multi-modality remote sensing aligns with the physiological stages of maize. The proposed architecture shows promise in improving predictions and offering interpretable insights into the factors affecting maize crop yields, while demonstrating the impact of data collection by different modalities through the growing season. By identifying relevant factors and critical growth stages, the model's attention weights provide valuable information that can be used in both plant breeding and crop management. The consistency of attention weights with biological growth stages reinforces the potential of deep learning networks in agricultural applications, particularly in leveraging remote sensing data for yield prediction. To the best of our knowledge, this is the first study that investigates the use of hyperspectral and LiDAR UAV time series data for explaining/interpreting plant growth stages within deep learning networks and forecasting plot-level maize grain yield using late fusion modalities with attention mechanisms.

59 BASIC BIOLOGICAL SCIENCES

Deep Learning enabled spectral energy conversion for in situ exposure measurements

A detector-specific deep learning (DL) approach is presented for spectra-to-exposure conversion using large-format sodium iodide (NaI(Tl)) detectors deployed for in situ environmental radiation measurements in emergency response scenarios. Accurate determination of exposure from NaI spectra is challenging due to poor energy resolution, partial energy absorption, and the strong sensitivity of traditionally deployed analytical conversion methods to calibrated source geometry and pre-deployment assumptions. Here, to address these limitations, a multi-layer perceptron model was trained on a hybrid in situ /Monte Carlo dataset constructed to span a broad range of photon energies, spatial extents, and realistic deployment variability, representative of general in situ emergency response conditions. The DL model was evaluated against commonly fielded analytical approaches under matched simulation conditions, including a single-factor method, a G-function method, and a modeled pressurized ion chamber (PIC) baseline. This study was intentionally computational in scope to enable controlled, like-for-like comparisons between conversion techniques while minimizing confounding real-world variability. Comparison to the modeled PIC provides contextual benchmarking and is not intended as a field inter-comparison with deployed instruments. Across the evaluated 20 keV to 3 MeV energy range, the DL approach consistently exhibited higher accuracy and reduced variance relative to the analytical methods against a deterministically calculated exposure. This may indicate improved robustness to spectral complexity without reliance on source-, geometric-, or spectral region-specific optimization. While results do not represent real-world validation, the presented work demonstrates that deep learning may effectively learn the nonlinear detector response-to-exposure relationship for asymmetric NaI(Tl) detectors and offers a promising pathway for improving in situ exposure estimation using spectroscopic systems already integrated into initial real-time emergency response operations.

61 RADIATION PROTECTION AND DOSIMETRY

PickerXL, A Large Deep Learning Model to Measure Arrival Times from Noisy Seismic Signals

Precisely measuring seismic arrival times is a labor-intensive task but is critical for both earthquake monitoring and subsurface imaging. Recently published deep learning models have demonstrated superior performance compared to traditional automatic approaches for picking arrival times. Although existing deep learning models have shown promising results, further advancements are necessary as their performance is not yet satisfactory especially when applied to new regions and station networks. Increasing model size has led to improved performance in other machine learning applications. Here, we aimed to investigate whether enlarging deep learning models can increase performance on accepted benchmarks. We trained three models of varying sizes, small (1X), medium (4X), and large (16X), using globally distributed local and regional earthquake signals and background noise waveforms from a benchmark dataset, Stanford Earthquake Dataset. Our results indicate that the largest model (PickerXL) outperforms both the smaller models and Seisbench implementation of the PhaseNet model, which has the same number of parameters as our small model. The PickerXL model’s enhanced capacity to extract complex patterns from seismograms contributes to its superior arrival picking abilities compared to the smaller model.

Chai, Chengping [Oak Ridge National Laboratory (OR

Impurity gas detection for SNF canisters using probabilistic deep learning and acoustic sensing *

Abstract Monitoring impurity gases in spent nuclear fuel (SNF) canisters is a novel structural health monitoring approach for SNF in dry storage. The SNF canisters are sealed containers that do not facilitate visual access to the inside. Acoustic sensing can be deployed by taking advantage of the pathways unobstructed by internal hardware. Although the ultrasonic time-of-flight measurement can provide valuable information, it is limited in its ability to discern the concentration of only one impurity gas. As such, deep learning algorithms, particularly convolutional neural networks (CNNs), offer a promising solution. In this study, CNN-based probabilistic deep learning models were implemented to detect and quantify multiple impurity gases in helium. An experimental platform was established to simulate canister conditions, and ultrasonic test data were collected. The presence of argon and air in helium at concentrations ranging from 0% to 1.2% at increments of 0.05% was considered. The multi-layer perceptron, decision tree, and logistic regression classifiers achieved high accuracies when distinguishing pure helium from helium with impurities. CNN with dropout layers and CNN using maximum likelihood estimation showed a similar performance, indicating their ability to capture uncertainties. The ensemble CNN model exhibited improved predictions and the ability to balance individual gas concentration by integrating 1D- and 2D-CNN models. These findings contribute probabilistic deep learning solutions for impurity gas detection and analysis within SNF canisters, thus ensuring safe storage and management of SNFs.

47 OTHER INSTRUMENTATION

Gradient-based optimization of complex nanoparticle heterostructures enabled by deep learning on heterogeneous graphs

Applications of deep learning (DL) to design nanomaterials are hampered by a lack of suitable data representations and training data. Here, in this study, we report efforts to overcome these limitations and leverage DL to optimize the nonlinear optical properties of core–shell upconverting nanoparticles (UCNPs). UCNPs, which have applications in fields such as biosensing, super-resolution microscopy and three-dimensional printing, can emit visible and ultraviolet light from near-infrared excitations. We report a large-scale dataset of UCNP emission spectra based on accurate but expensive kinetic Monte Carlo simulations (N > 6,000) and use these data to train a heterogeneous graph neural network using a physically motivated representation of UCNP nanostructure. Applying gradient-based optimization on the trained graph neural network, we identify structures with 6.5× higher predicted emission under 800-nm illumination than any UCNP in our training set. Our work reveals design principles for UCNP heterostructures and presents a roadmap for DL-based inverse design of nanomaterials.

Sivonxay, Eric [Lawrence Berkeley National Laborat

Efficient Dimension Reduction of Complex Three-dimensional CO2 Saturation using Deep Learning Models

In the domain of deep learning (DL), dimension reduction is crucial for enhancing training efficiency and mitigating overfitting, particularly when managing complex data such as three-dimensional (3D) saturation data. The 3D saturation data in the context of geological carbon storage (GCS) presents unique challenges due to its inherent sparsity and the abrupt transitions at plume boundaries, known as shock fronts. To address the challenges, we proposed a novel DL framework that integrates dimension reduction with advanced 3D reconstruction techniques. Our model leveraged latent variables derived from 2D average saturation data, offering a robust and efficient solution tailored to the intricate dynamics of 3D saturation fields. The proposed framework can extract the critical features of the high-dimensional data while reducing the variable numbers, which is more tractable for DL models and enhances the model robustness and accuracy. Therefore, it provides a novel approach for modeling and analyses in complex geological scenarios, which finds great potential applications in environmental monitoring and energy storage.

Wang, Hongsheng

G2PDeep-v2: A Web-Based Deep-Learning Framework for Phenotype Prediction and Biomarker Discovery for All Organisms Using Multi-Omics Data

Multi-omics data offers rich insights into complex traits across organisms, yet integrating and analyzing these datasets for phenotype prediction and marker discovery remains challenging. Researchers need accessible tools that combine deep learning, hyperparameter optimization, visualization, and downstream analysis in a unified web platform. To address this, we developed G2PDeep-v2, a web-based platform powered by deep learning for phenotype prediction and marker discovery from multi-omics data across a wide range of organisms, including humans and plants. The server provides multiple services for researchers to create deep-learning models through an interactive interface and train these models using an automated hyperparameter tuning algorithm on high-performance computing resources. Users can visualize the results of phenotype and markers predictions and perform Gene Set Enrichment Analysis for the significant markers to provide insights into the molecular mechanisms underlying complex diseases, conditions and other biological phenotypes being studied.

59 BASIC BIOLOGICAL SCIENCES

Improved deep learning prediction of antigen–antibody interactions

Identifying antibodies that neutralize specific antigens is crucial for developing effective immunotherapies, but this task remains challenging for many target antigens. The rise of deep learning–based computational approaches presents a promising avenue to address this challenge. Here, we assess the performance of a deep learning approach through two benchmark tests aimed at predicting antibodies for the receptor-binding domain of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein. Three different strategies for constructing input sequence alignments are employed for predicting structural models of antigen–antibody complexes. In our initial testing set, which comprises known experimental structures, these strategies collectively yield a significant top-ranked prediction for 61% of cases and a success rate of 47%. Notably, one strategy that utilizes the sequences of known antigen binders outperforms the other two, achieving a precision of 90% in a subsequent test set of ~1,000 antibodies, balanced between true and control antibodies for the antigen, albeit with a lower recall of 25%. Our results underscore the potential of integrating deep learning methods with single B cell sequencing techniques to enhance the prediction accuracy of antigen–antibody interactions.

Science & Technology - Other Topics

Deep-learning based artificial intelligence tool for melt pools and defect segmentation

Accelerating fabrication of additively manufactured components with precise microstructures is important for quality and qualification of built parts, as well as for a fundamental understanding of process improvement. Accomplishing this requires fast and robust characterization of melt pool geometries and structural defects in images. This paper proposes a pragmatic approach based on implementation of deep learning models and self-consistent workflow that enable systematic segmentation of defects and melt pools in optical images. Deep learning is based on an image-to-image translation–conditional generative adversarial neural network architecture. An artificial intelligence (AI) tool based on this deep learning model enables fast and incrementally more accurate predictions of the prevalent geometric features, including melt pool boundaries and printing-induced structural defects. We present statistical analysis of geometric features that is enabled by the AI tool, showing strong spatial correlation of defects and the melt pool boundaries. The correlations of widths and heights of melt pools with dataset processing parameters show the highest sensitivity to thermal influences resulting from laser passes in adjacent and subsequent layer passes. The presented models and tools are demonstrated on the aluminum alloy and datasets produced with different sets of processing parameters. However, they have universal quality and could easily be adapted to different material compositions. The method can be easily generalized to microstructural characterizations other than optical microscopy.

additive manufacturing

Arm and shoulder muscle segmentation in axial MRI with UNet deep learning model

Quantifying individual upper-limb muscle volumes from MRI provides key insight into muscle-specific strength, deficits, and adaptations. Manual delineation is the gold standard but time‑intensive, and the performance of current deep learning approaches, particularly for small or anatomically complex muscles, remains incompletely characterized. We evaluated a state‑of‑the‑art deep learning framework across the entire upper limb and analyzed factors governing segmentation performance, with attention to the forearm. Three previously published MRI datasets (1.5 T, 3D GRE T1‑weighted; total n = 39) spanning young, middle‑aged, and older adults were curated and quality‑checked, including expert manual segmentations for 31 muscles. Following multiclass mask reconstruction, we trained three 3D nnU‑Net multiclass models matched to the muscle subsets present across datasets, using five‑fold cross‑validation and a composite Dice Similarity Coefficient (DSC) + cross entropy loss. Segmentation accuracy was assessed with DSC. Performance varied across muscles (mean DSC = 0.806 ± 0.098), ranging from 0.920 (Deltoid) to 0.461 (Extensor pollicis brevis). In uncertainty‑weighted regressions, muscle volume was positively associated with DSC (R2 = 0.36, p < 0.001), whereas training segmentation count and muscle orientation showed negligible associations (R2 ≤ 0.06). A weighted mixed‑effects model identified volume as the strongest evaluated predictor, explaining 23.9% of variance in DSC; orientation and training count each contributed <1%, leaving 61.5% unexplained. These results indicate that deep learning–based segmentation can accurately quantify muscle volume for many upper‑limb muscles but remains constrained for small, low‑contrast forearm muscles.

Gillespie, Samuel

Phase Picking Beyond Local Distances: Where Waveform Filtering Still Matters for Deep Learning Models

Waveform filtering is a standard step in traditional seismic phase picking but often receives little attention in deep learning workflows, where models are typically trained on raw or minimally processed waveforms. Although this strategy performs well for local events, we show that performance can degrade substantially at regional distances. To address this limitation, we introduce two ways to incorporate multiband-filtered waveforms into deep learning phase pickers. The stacking approach concatenates filtered inputs along the channel dimension, while the branching approach processes each frequency band through a dedicated network branch before feature fusion. Both approaches can substantially improve performance across epicentral distances of 0° to 20°, but their effectiveness depends strongly on the selected frequency bands. Tests with multiple filter banks show that filter-bank design should be treated as part of model optimization rather than as a fixed preprocessing choice. Grad-CAM analysis of the branching model indicates that band importance varies among waveform samples and across training realizations, with only a weak overall preference for the 0.25 to 0.5 Hz band. These results show that no single filter band is consistently optimal and demonstrate that explicit feature engineering remains valuable for robust deep learning-based seismic phase picking.

58 GEOSCIENCES

Large-scale deep learning for metastasis detection in pathology reports

Objectives No existing algorithm can reliably identify metastasis from pathology reports across multiple cancer types and the entire US population. In this study, we develop a deep learning model that automatically detects patients with metastatic cancer by using pathology reports from many laboratories and of multiple cancer types. Materials and Methods We use 60 471 unstructured pathology reports from 4 Surveillance, Epidemiology, and End Results (SEER) registries. The reports were coded into 1 of 3 labels: metastasis negative, metastases positive, or metastasis undetermined. We utilize a task-specific deep neural network trained from scratch and compare its performance with a widely used large language model (LLM). Results Our deep learning architecture trained on task-specific data outperforms a general-purpose LLM, with a recall of 0.894 compared to 0.824. We quantified model uncertainty and used it to defer reports for human review. We found that retaining 72.9% of reports increased recall from 0.894 to 0.969. Discussion A smaller deep learning architecture trained on task-specific data outperforms a general LLM. Equally critical to model performance is the incorporation of uncertainty quantification, achieved here through an abstention mechanism. Conclusions This study’s finding demonstrate the feasibility of developing algorithms to automatically identify metastatic cancer cases from unstructured pathology reports.

machine learning

Deep learning–based digital twins for heat pumps

Heat pumps are effective cooling and heating appliances to save energy in buildings. However, traditional heat pump models are challenging to integrate with building demands in a co-simulation environment because of the nonlinear thermodynamics of refrigerants. Developing digital twin representatives for heat pumps capable of faster calculations with good accuracy is desirable. This study aimed to establish a generic deep learning–based digital twin for heat pumps with a large amount of high-fidelity data. Two refrigerants for two different heat pumps were considered: an air source heat pump with refrigerant R-410A, an air source heat pump with refrigerant CO 2 , a water source heat pump with refrigerant R-410A, and a water source heat pump with refrigerant CO 2 . Furthermore, results showed that the deep learning (long short-term memory) models effectively represented these four heat pumps as a digital twin: (a) accuracy for training and testing showed smaller than 0.02 for heating electricity and heating demands, and (b) the digital twins showed good consistency with original data for heating electricity and heating demands (root mean square errors of less than 0.12 W and 0.19 W, respectively). Therefore, deep learning–based heat pump models can be used in the co-simulation of building mechanical systems.

Air source heat pump

Inverse design of hypoeutectoid pearlite steel microstructures using a deep learning and genetic algorithm optimization framework

Goal-oriented microstructure design in metallic materials is a challenging task due to complex structure-property relationships. Traditional experimental and computational approaches are time-intensive and economically inefficient, limiting their applicability for large-scale design space exploration. Here, in this work, we propose an end-to-end framework that integrates deep learning models with genetic optimization to design microstructures with targeted mechanical properties. Deep learning models enable accurate forward design, while their integration with genetic optimization enables efficient inverse design within a few hours, compared to days or weeks using conventional finite element simulations. The framework combines experimental characterization and finite element modeling to analyze the influence of microstructural features on the mechanical behavior of hypoeutectoid steels. Data from both experiments and simulations are used to train the deep learning models. To demonstrate its effectiveness, we apply the framework to 0.63% carbon steel with proeutectoid ferrite and pearlite phases, commonly used in industrial applications. In this study, 2D microstructures were used for modeling, selected primarily for computational efficiency and to establish proof of concept. The framework successfully optimizes microstructures for targeted yield strength, ultimate strength, and stress concentration factors while significantly reducing computational time. Beyond hypoeutectoid steels, this scalable framework can be extended to other material systems and integrated with additive manufacturing, offering an efficient approach for accelerating microstructure design for specific engineering applications.

ConvLSTM