Search NASA⌕ Search

SEARCH · Search NASA

Results for “deep learning compilers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Cross-Feature Transfer Learning for Efficient Tensor Program Generation

Tuning tensor program generation involves navigating a vast search space to find optimal program transformations and measurements for a program on the target hardware. The complexity of this process is further amplified by the exponential combinations of transformations, especially in heterogeneous environments. This research addresses these challenges by introducing a novel approach that learns the joint neural network and hardware features space, facilitating knowledge transfer to new, unseen target hardware. A comprehensive analysis is conducted on the existing state-of-the-art dataset, TenSet, including a thorough examination of test split strategies and the proposal of methodologies for dataset pruning. Leveraging an attention-inspired technique, we tailor the tuning of tensor programs to embed both neural network and hardware-specific features. Notably, our approach substantially reduces the dataset size by up to 53% compared to the baseline without compromising Pairwise Comparison Accuracy (PCA). Furthermore, our proposed methodology demonstrates competitive or improved mean inference times with only 25–40% of the baseline tuning time across various networks and target hardware. The attention-based tuner can effectively utilize schedules learned from previous hardware program measurements to optimize tensor program tuning on previously unseen hardware, achieving a top-5 accuracy exceeding 90%. This research introduces a significant advancement in autotuning tensor program generation, addressing the complexities associated with heterogeneous environments and showcasing promising results regarding efficiency and accuracy.

97 MATHEMATICS AND COMPUTING↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

FOS: Computer and information sciences↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

Schulte, Jan-Frederik [Purdue U.] (ORCID:000000034↗

Satellite Imagery of PV Site Storm Damage

"This repository contains multiple data sets focused on visible damage to photovoltaic (PV) installations following extreme weather events such as hailstorms and hurricanes. Data sets are split into two categories: the first category, the ‘manually labeled’ data, was compiled by researchers manually, and contains manually identified PV sites exposed to storms. The second data set, the ‘aggregated’ data, is a compilation of the manually labeled PV sites and deep learning-identified PV sites. The hail damage data set focuses on post-storm PV damage following a September 24, 2023 hailstorm in Austin, TX, which caused over $600 million in damages in the Austin metro area. The hurricane damage data set focuses on post-storm PV damage following Hurricanes Irma and Maria in Puerto Rico and the US Virgin Islands. Hurricanes Irma and Maria were back-to-back category 5 hurricanes, which pummeled the Caribbean and southeastern United States in September 2017, causing an estimated $115.2 billion in damages."

14 SOLAR ENERGY↗

Development and Evaluation of a General Drag Model for Gas-Solid Flows via Deep Learning

This project presents the development and evaluation of a general drag model for gas–solid multiphase flows using deep learning techniques. A comprehensive database of more than 4,000 experimental and numerical data points for spherical and non spherical particles was compiled, incorporating geometric features such as sphericity, aspect ratio, and orientation. Several predictive approaches—including traditional em pirical correlations, machine learning, and deep neural networks—were benchmarked, with the proposed Drag Coefficient Correlation-aided Deep Neural Network (DCC DNN) demonstrating superior accuracy. To account for particle–particle interactions, additional drag data were generated using CFD-based simulations of packed and flu idized beds, leading to the development of a retrained model capable of incorporat ing volume fraction effects. Integration of the trained model with the MFiX CFD solver was achieved using FTorch, enabling drag predictions during discrete element method (DEM) simulations. Validation against experimental data for single particles and fluidized beds confirmed the model’s improved predictive ability, particularly for non-spherical geometries. While the model performed strongly under fluidized con ditions, limitations remained in unfluidized regimes, suggesting a need for expanded datasets. Overall, this study demonstrates the feasibility of combining deep learning with physics-informed CFD to improve drag modeling for gas–solid flows, with promis ing implications for scaling multiphase simulations in industrial applications.

42 ENGINEERING↗

Optimizing Deep Learning Models for Climate-Related Natural Disaster Detection from UAV Images and Remote Sensing Data

This research study utilized artificial intelligence (AI) to detect natural disasters from aerial images. Flooding and desertification were two natural disasters taken into consideration. The Climate Change Dataset was created by compiling various open-access data sources. This dataset contains 6334 aerial images from UAV (unmanned aerial vehicles) images and satellite images. The Climate Change Dataset was then used to train Deep Learning (DL) models to identify natural disasters. Four different Machine Learning (ML) models were used: convolutional neural network (CNN), DenseNet201, VGG16, and ResNet50. These ML models were trained on our Climate Change Dataset so that their performance could be compared. DenseNet201 was chosen for optimization. All four ML models performed well. DenseNet201 and ResNet50 achieved the highest testing accuracies of 99.37% and 99.21%, respectively. This research project demonstrates the potential of AI to address environmental challenges, such as climate change-related natural disasters. This study’s approach is novel by creating a new dataset, optimizing an ML model, cross-validating, and presenting desertification as one of our natural disasters for DL detection. Three categories were used (Flooded, Desert, Neither). Our study relates to AI for Climate Change and Environmental Sustainability. Drone emergency response would be a practical application for our research project.

AI↗

A Full-Stack Exploration of Language-Based Parallelism in Fortran 2023

This poster explores native parallel features in Fortran 2023 through the lens of supporting applications with libraries, compilers, and parallel runtimes. The language revision informally named Fortran 2008 introduced parallelism in the form of Single Program Multiple Data (SPMD) execution with two broad feature sets: (1) loop-level parallelism via do concurrent and (2) a Partitioned Global Address Space (PGAS) comprised of distributed “coarray” data structures. Fortran’s native parallelism has demonstrated high performance [1] and reduced the burden of inserting what sometimes amounts to more directives than code. Several compilers support both feature sets, typically by translating do concurrent into serial do loops annotated by parallel directives and by translating SPMD/PGAS features into direct calls to a communication library. Our research focuses primarily on two questions: (1) can the compiler’s parallel runtime library be developed in the language being compiled (Fortran) and (2) can we define an interface to the runtime that liberates compilers from being hardwired to one runtime and vice versa. We are answering these questions by developing the Parallel Runtime Interface for Fortran (PRIF) [2] and the Co-Array Fortran Framework of Efficient Interfaces to Network Environments (Caffeine) [3]. Caffeine is initially targeting adoption by LLVM Flang, a new open-source Fortran compiler developed by a broad community in industry, academia, and government labs. We are also exploring the use of these features in Inference-Engine, a deep learning library designed to facilitate neural network training and inference for high-performance computing applications written in modern Fortran.

Rasmussen, Katherine↗

Strong Lens Discoveries in DESI Legacy Imaging Surveys DR10 with Two Deep Learning Architectures

Abstract We have conducted a search for strong gravitational lensing systems in the Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Surveys Data Release 10 (DR10). This paper is the fourth in a series of searches. This is the first catalog of lens candidates covering nearly the entirety of the extragalactic sky south of declination δ ≈ +32 ∘ , all observed by DECam, covering ∼14,000 deg 2 . We impose a z -band magnitude cut of <20 in AB magnitude. We deploy a residual neural network and EfficientNet as an ensemble trained on a compilation of known lensing systems and high-grade candidates as well as nonlenses in the same footprint. The predictions from these two base models are aggregated using a meta-learner. After applying our ensemble to the survey data, we exclude known candidates and systems, and use our own visual inspection portal to rank images in the top 0.01 percentile of all neural network recommendations. We have found 811 lens candidates, five of which are confirmed through Euclid Quick Data Release (Q1). These include 484 new candidates in the Legacy Surveys DR9 footprint, all parts of which have been searched for strong lenses at least once before, either by our group or others. Combining the discoveries from this work with those from the first three papers in this series (335, 1210, and 1512), we have discovered a total of 3868 new candidates in the DESI Legacy Surveys.

Inchausti, Jose Carlos [University of San Francisc↗

Deep learning of experimental electrochemistry for battery cathodes across diverse compositions

Artificial intelligence (AI) has emerged as a tool for discovering and optimizing novel battery materials. However, the adoption of AI in battery cathode representation and discovery is still limited due to the complexity of optimizing multiple performance properties and the scarcity of high-fidelity data. Here, we present a machine learning model (DRXNet) for battery informatics and demonstrate the application in the discovery and optimization of disordered rocksalt (DRX) cathode materials. We have compiled the electrochemistry data of DRX cathodes over the past 5 years, resulting in a dataset of more than 19,000 discharge voltage profiles on diverse chemistries spanning 14 different metal species. Learning from this extensive dataset, our DRXNet model can capture critical features in the cycling curves of DRX cathodes under various conditions. Our approach offers a data-driven solution to facilitate the rapid identification of novel cathode materials, accelerating the development of next-generation batteries for carbon neutralization.

25 ENERGY STORAGE↗

A semi-agnostic ansatz with variable structure for variational quantum algorithms

Quantum machine learning—and specifically Variational Quantum Algorithms (VQAs)—offers a powerful, flexible paradigm for programming near-term quantum computers, with applications in chemistry, metrology, materials science, data science, and mathematics. Here, one trains an ansatz, in the form of a parameterized quantum circuit, to accomplish a task of interest. However, challenges have recently emerged suggesting that deep ansatzes are difficult to train, due to flat training landscapes caused by randomness or by hardware noise. This motivates our work, where we present a variable structure approach to build ansatzes for VQAs. Our approach, called VAns (Variable Ansatz), applies a set of rules to both grow and (crucially) remove quantum gates in an informed manner during the optimization. Consequently, VAns is ideally suited to mitigate trainability and noise-related issues by keeping the ansatz shallow. We employ VAns in the variational quantum eigensolver for condensed matter and quantum chemistry applications, in the quantum autoencoder for data compression and in unitary compilation problems showing successful results in all cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Evolution of Software-Only-Simulation at NASA IV and V

Software-Only-Simulations have been an emerging but quickly developing field of study throughout NASA. The NASA Independent Verification Validation (IVV) Independent Test Capability (ITC) team has been rapidly building a collection of simulators for a wide range of NASA missions. ITC specializes in full end-to-end simulations that enable developers, VV personnel, and operators to test-as-you-fly. In four years, the team has delivered a wide variety of spacecraft simulations that have ranged from low complexity science missions such as the Global Precipitation Management (GPM) satellite and the Deep Space Climate Observatory (DSCOVR), to the extremely complex missions such as the James Webb Space Telescope (JWST) and Space Launch System (SLS).This paper describes the evolution of ITCs technologies and processes that have been utilized to design, implement, and deploy end-to-end simulation environments for various NASA missions. A comparison of mission simulators are discussed with focus on technology and lessons learned in complexity, hardware modeling, and continuous integration. The paper also describes the methods for executing the missions unmodified flight software binaries (not cross-compiled) for verification and validation activities.

Embedded↗

Automatically parallelizing batch inference on deep neural networks using Fiats and Fortran 2023 `do concurrent`

This paper introduces novel programming strategies that leverage features of the Fortran 2023 standard of the International Standards Organization (ISO) to automatically parallelize computations on deep neural networks. The paper focuses on the interplay of object-oriented, parallel, and functional programming paradigms in the Fiats deep learning library. We demonstrate how several infrequently used language features play a role in enabling efficient, parallel execution. Specifically, the ability to explicitly declare that a procedure is pure facilitates inference in the context of the language’s loop-parallelism construct `do concurrent`. Also, explicitly prohibiting the overriding of a parent type’s type-bound procedures eliminates the need for dynamic dispatch in performance-critical code. Finally, this paper uses batch inference calculations on a neural network surrogate for atmospheric aerosol dynamics to demonstrate that LLVM Flang compiler’s automatic parallelization of `do concurrent` achieves roughly the same performance and scalability as achieved by OpenMP compiler directives. We also demonstrate that double-precision inference costs 37–72% longer runtime than default-real precision with most values in the range 57-60%.

Rouson, Damian↗

Search for brown dwarfs in IC 1396 with Subaru HSC: interpreting the impact of environmental factors on substellar population

ABSTRACT Young stellar clusters are predominantly the hub of star formation and hence, ideal to perform comprehensive studies over the least explored substellar regime. Various unanswered questions like the mass distribution in brown dwarf regime and the effect of diverse cluster environment on brown dwarf formation efficiency still plague the scientific community. The nearby young cluster, IC 1396 with its feedback-driven environment, is ideal to conduct such study. In this paper, we adopt a multiwavelength approach, using deep Subaru HSC along with other data sets and machine learning techniques to identify the cluster members complete down to ∼ 0.03 M⊙ in the central 22 arcmin area of IC 1396. We identify 458 cluster members including 62 brown dwarfs which are used to determine mass distribution in the region. We obtain a star-to-brown dwarf ratio of ∼ 6 for a stellar mass range 0.03–1 M⊙ in the studied cluster. The brown dwarf fraction is observed to increase across the cluster as radial distance from the central OB-stars increases. This study also compiles 15 young stellar clusters to check the variation of star-to-brown dwarf ratio relative to stellar density and ultraviolet (UV) flux ranging within 4–2500 stars pc−2 and 0.7–7.3 G0, respectively. The brown dwarf fraction is observed to increase with stellar density but the results about the influence of incident UV flux are inconclusive within this range. This is the deepest study of IC 1396 as of yet and it will pave the way to understand various aspects of brown dwarfs using spectroscopic observations in future.

Gupta, Saumya (ORCID:0000000161843958)↗

GAHLS: an optimized graph analytics based high level synthesis framework

The urgent need for low latency, high-compute and low power on-board intelligence in autonomous systems, cyber-physical systems, robotics, edge computing, evolvable computing, and complex data science calls for determining the optimal amount and type of specialized hardware together with reconfigurability capabilities. With these goals in mind, we propose a novel comprehensive graph analytics based high level synthesis (GAHLS) framework that efficiently analyzes complex high level programs through a combined compiler-based approach and graph theoretic optimization and synthesizes them into message passing domain-specific accelerators. This GAHLS framework first constructs a compiler-assisted dependency graph (CaDG) from low level virtual machine (LLVM) intermediate representation (IR) of high level programs and converts it into a hardware friendly description representation. Next, the GAHLS framework performs a memory design space exploration while account for the identified computational properties from the CaDG and optimizing the system performance for higher bandwidth. The GAHLS framework also performs a robust optimization to identify the CaDG subgraphs with similar computational structures and aggregate them into intelligent processing clusters in order to optimize the usage of underlying hardware resources. Finally, the GAHLS framework synthesizes this compressed specialized CaDG into processing elements while optimizing the system performance and area metrics. Evaluations of the GAHLS framework on several real-life applications (e.g., deep learning, brain machine interfaces) demonstrate that it provides 14.27× performance improvements compared to state-of-the-art approaches such as LegUp 6.2.

97 MATHEMATICS AND COMPUTING↗

Disruptive Technologies and Their Putative Impacts Upon Society and Aerospace- Entering The Virtual Age

Developments in technology over the recent decades have been extraordinary. They include the IT, bio, nano, and now quantum and energetics technology arenas and their many combinatorial interactions and impacts. In the main, these are at the frontiers of the small and in a combinational, synergistic feeding frenzy with each other. They fall under the broad category of Disruptive Technologies and have greatly altered society. The outlook for the runout of these and other technology developments augers mid-term to later alterations in components of the human existence theorem, including the requirement to work for our living and our physiological makeup and longevity (Ref 1). The IT revolution began in the 1950s with the development of solid-state electronics. The biologics revolution began later in the 1960s and 1970s with DNA and genomics, and the nano revolution in the 1990s with self-forming nano systems and carbon nanotubes. Quantum technology is now developing rapidly, aided by enabling nano systems, and the energetics revolution is providing ever more efficient and less expensive renewable energy sources. The IT revolution has produced improvements of an astounding eleven orders of magnitude in computing speed since the late 1950s. As we shift from silicon to biological, optical, nano, molecular, and atomic computing, improvements of some 4 orders of magnitude are evidently possible from either optical or DNA computing [Refs 2and 3], then there are combinatorials. Then there is quantum computing, under development worldwide for an increasing number of applications and proffering phenomenal capabilities. The current fastest computers are considerably beyond human brain speed. Machine intelligence is developing well after decades of inadequate machine capability, now no longer the case, and a detour into expert systems. Researchers in machine intelligence are now pursuing deep learning approaches using neural nets, which are proving to be extremely useful. Some believe the frontier of potential human-level machine intelligence may be found in biomimetics and brain-emulation approaches. There is even a possibility of “emergence”—i.e., when the machine intelligence is complex enough that it “wakes up,” as when human intelligence emerged via evolution during the million-plus years of the hunter-gatherer epoch [ Ref 4]. In fact, some posit that human intelligence can be improved upon and is only a cul-de-sac of what is conceivable. The IT revolution has produced massive changes in human society and economics—from the Internet, enabling the rapid expansion of knowledgeability (and even what is knowable), to an increasingly pervasive trend of “tele-everything.” The extraordinary compilation, storage, and availability of truly massive amounts of information could, when combined with AI and under the mantra of “big data,” greatly improve many of our technical and commercial processes and their content including elucidating new heuristic governing laws.

Dennis M. Bushnell↗

Learning nuclear cross sections across the chart of nuclides with graph neural networks

We explore the use of deep learning techniques to learn how nuclear cross sections change as we add or remove protons and neutrons. As a proof of principle, we focus on the neutron-induced reactions in the fast energy regime. Our approach follows a two-stage learning framework. First, we apply representation learning to encode cross section data into a latent space using either variational autoencoders (VAEs) or implicit neural representations (INRs). Then, we train graph neural networks (GNNs) on the resulting embeddings to predict missing values across the nuclear chart by leveraging the topological structure of neighboring isotopes. We demonstrate accurate cross section predictions within a 9 × 9 block of missing nuclei. We also find that the optimal GNN training strategy depends on the type of latent representation used, with VAE embeddings performing best under end-to-end optimization in the original space, while INR embeddings achieve better results when the GNN is trained only in the latent space. Furthermore, using clustering algorithms, we map groups of latent vectors into regions of the nuclear chart and show that VAEs and INRs can discover some of the neutron magic numbers. These findings suggest that deep-learning models based on the representation encoding of cross sections combined with graph neural networks hold significant potential in augmenting nuclear theory models, e.g., by providing reliable estimates of covariances of cross sections, including cross-material covariances.

Machine learning↗