Search NASA⌕ Search

SEARCH · Search NASA

Results for “• Artificial intelligence (AI) / machine learning (ML), artificial intelligence, machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Status Report on Regulatory Criteria Applicable to the Use of Artificial Intelligence (AI) and Machine Learning (ML)

Although the interest in the use of artificial intelligence (AI) and machine learning (ML) in nuclear energy is increasing rapidly, at present their implementation is limited. This rapid increase in interest is not surprising considering that implementing AI and ML technology would allow for continuous monitoring, facilitate the implementation of predictive maintenance with optimized staffing plans, enable automation and autonomy opportunities that could drastically reduce fixed operation and maintenance costs, and provide training for operations and maintenance. Other industries are using AI for construction, and in the nuclear arena AI could provide great benefit in decommissioning activities. The ability of AI and ML to operate in real time vastly increases their potential impact. Before AI can be used in design, operations, or as a regulatory tool, the specifics on the regulations applicable to the use of AI for nuclear power applications need to be established. The difficulty is that the specific use cases will dictate the applicability of regulations. For example, even within the application domain associated with operations, the regulations might vary if the AI is used to create a virtual reference for plant operations or is used for training, optimization of maintenance intervals, prioritization of maintenance activities, etc. Different still is if the AI is to be used for design or setting technical specifications, which will introduce additional requirements. US Nuclear Regulatory Commission (NRC) licensing reviews are based on an applicant’s design meeting its performance assessment based on (1) safety goals and objectives, (2) deterministic and/or probabilistic analysis of accident scenarios, and (3) quantitative assessment of design alternatives against the safety goals and objectives using accepted engineering tools, methodologies, and performance criteria. The current regulatory framework does not explicitly address AI or autonomous control. However, as implementing AI technology will require the use of a digital platform, it must meet the requirements of an instrumentation and control (I&C) system. The regulatory requirements for AI, which will be incorporated into the I&C system, will be very dependent on how it is used (i.e., its functionality, safety classification, etc.). The licensing process is primarily risk-based with the identification of components and systems as nonsafety, important to safety, or safety related. A risk-informed approach allows further gradation of components and systems based on risk metrics such as core damage frequency or large early release fractions. Thus, the use cases and the risk categorization of impacted systems and components will determine the regulatory requirements. Regardless of how AI is used it presents new opportunities for risk-informing operating, maintenance, and regulatory decisions. Trustworthiness, transparency, and the ability to validate and verify the results will be paramount in showing that the systems and plant still meet their performance requirements. This report describes the results of research to identify regulatory implications of AI technologies and their uses. Specifically, this report reviews current regulatory guidance relevant to the application of AI for design (including design changes or new designs including advanced reactors), construction, operations, training, maintenance, research, testing, and as a regulatory tool. AI can be automated at different levels from purely informative purposes to autonomous controls. The focus of this review included determination of constraints on the application of AI technology, identification of any regulatory gaps or uncertainties, and clarification of anticipated technical basis information likely to be important for regulatory acceptance of these technologies. Currently, any use of AI at nuclear power plants is focused on nonsafety-related applications. The NRC and other regulatory bodies are evaluating providing guidance to address gaps rather than create new regulations to address the use of AI and ML. This approach seems to be the best to encourage AI development without adding regulatory uncertainty.

97 MATHEMATICS AND COMPUTING↗

Measuring Equality in Machine Learning Security Defenses: A Case Study in Speech Recognition

Over the past decade, the machine learning security community has developed a myriad of defenses for evasion attacks. An understudied question in that community is: for whom do these defenses defend? This work considers common approaches to defending learned systems and how security defenses result in performance inequities across different sub-populations. We outline appropriate parity metrics for analysis and begin to answer this question through empirical results of the fairness implications of machine learning security methods. We find that many methods that have been proposed can cause direct harm, like false rejection and unequal benefits from robustness training. The framework we propose for measuring defense equality can be applied to robustly trained models, preprocessing-based defenses, and rejection methods. We identify a set of datasets with a user-centered application and a reasonable computational cost suitable for case studies in measuring the equality of defenses. In our case study of speech command recognition, we show how such adversarial training and augmentation have non-equal but complex protections for social subgroups across gender, accent, and age in relation to user coverage. We present a comparison of equality between two rejection-based defenses: randomized smoothing and neural rejection, finding randomized smoothing more equitable due to the sampling mechanism for minority groups. This represents the first work examining the disparity in the adversarial robustness in the speech domain and the fairness evaluation of rejection-based defenses.

• Artificial intelligence (AI) / machine learning ↗

SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions

Instruction finetuning is a popular paradigm to align large language models (LLM) with human intent. Despite its popularity, this idea is less explored in improving the LLMs to align existing foundation models with scientific disciplines, concepts and goals. In this work, we present SciTune as a tuning framework to improve the ability of LLMs to follow scientific multimodal instructions. To test our methodology, we use a human-generated scientific instruction tuning dataset and train a large multimodal model LLaMA-SciTune that connects a vision encoder and LLM for science-focused visual and language understanding. LLaMA-SciTune significantly outperforms the state-of-the-art models in the generated figure types and captions in multiple scientific multimodal benchmarks. In comparison to the models that are fine-tuned with machine generated data only, LLaMA-SciTune surpasses human performance on average and in many sub-categories on the ScienceQA benchmark.

• Artificial intelligence (AI) / machine learning ↗

Using Artificial Intelligence (AI) and Machine Learning (ML) to conduct Space Missions Solid Waste Management Survey

The National Aeronautics and Space Administration (NASA) Solid Waste Management team has been focusing on technologies that can operate in microgravity. NASA aims to conduct both short and long-term transit and planetary missions on the lunar and Mars surfaces. Therefore, an updated waste survey is needed to explore technologies for operation in microgravity for transit missions and partial gravity for planetary missions. This paper will utilize Artificial Intelligence and Machine Learning techniques to conduct the survey and generate knowledge graphs for Spacecraft Waste Management.

Artificial Intelligence↗

Artificial Intelligence in Aviation Safety Applications - Exploring Myths and Truths of AI and ML

Artificial Intelligence (AI) and machine learning (ML) are gaining increased attention as ways to leverage the world's data to solve problems. Although AI and ML offer much potential, there are often misconceptions about the application of such techniques.Panel speakers will present machine learning approaches they have developed on a variety of aviation data, including digital flight data, safety reporting data, and voice communications data. They will discuss the purpose of the application, the data used, and the lessons learned in the development and deployment of their solutions. The panel will also discuss common pitfalls in developing an AI solution, the dangers of the current hype around AI, tips for gaining value from a ML solution, how to determine whether a ML approach is appropriate for a problem, and more.

Reeves, Scott (Capt.)↗

Multi-Source Machine Learning and Thermoplastics Enhanced Aerostructure Manufacturing (mTEAM)

RTX Technology Research Center (RTRC), together with Collins Aerospace (Collins) and Oak Ridge National Laboratory (ORNL) has developed an Artificial Intelligence (AI) / Machine Learning (ML) guided solution to advance the manufacturing and assembly of high performance and lightweight thermoplastic composite (TPC) aerospace products. The solution aims to lower risk, cost and lead time for induction heating based welding and consolidation processes for TPC structure. The cost and lead time of part and material specific process development for induction welding (IW) and induction consolidation will be reduced by replacing traditional empirical methods with optimization methods that merge AI/ML and physics-based process simulations and process experiments with sensing and controls. TPC-IW process development is empirical in nature, and uncertainties in material & process behavior exist near & far from the induction coil. Physics-based simulations can be leveraged directly for process optimization but can be too computationally expensive to run in high fidelity and real time to do robust process optimization. The key impact of successful TPC induction consolidation and welding is cost & lead time reduction for part & material specific consolidation and welding recipes. This is an enabler for more rapid deployment of TPC structures via joining assembly, which can reduce energy & cost intensive usage of autoclaves & ovens. The solution aimed to advance the U.S. Department of Energy’s interests in using thermoplastics and automation in composite manufacturing for improvement of products for existing markets via increased production speeds, reduced costs, and lowered use of energy. Welded TPC structures can offer significant weight & energy savings for high-value commercial aerospace & industrial applications compared to metal & thermoset composite structures assembled by mechanical fastening and/or adhesive bonding. The project was organized into two Budget Periods. Budget Period 1 (BP1) was 15 months and its goal was to perform ML process optimization framework development & deployment on lab-coupon aerostructure components. A Go/No-Go Review was performed at the end of BP1 to verify fulfilment of key tasks & milestones to justify a Go Decision to move into the next Budget Period. Budget Period 2 (BP2) was 12 months and its goal was the deployment of the ML framework for ML process optimization of pilot industrial scale aerostructure components. The overall project aim was to develop & demonstrate ML-enhanced modeling framework that learns process-property mapping from multiple data sources at different fidelities. During BP1, the team accomplished key tasks & milestones to demonstrate the concept of multi-source ML for TPC aerostructure consolidation and assembly. First, the team completed documentation of induction based TPC heating requirements including baseline metrics to compare measured results against. Next the team completed demonstration of data generation from physics-based simulations for ML surrogate model generation and demonstrated the integration of physics-based simulation data into multi-source AI/ML algorithms. In parallel, the team established the lab-coupon scale induction welding system and completed a process to label and reduce generated data from physics-based simulation and experiments for ML surrogate models to enable multi-source ML model training & testing. To complete BP1, the team integrated physics-based simulation data and experimental data into multi-source ML algorithms. This was based on the team completing ML deployment of the induction welding on a lab system at RTRC and AI/ML deployment on existing induction welding line at Collins. ORNL visited both Collins and RTRC sites to witness the TPC induction welding process. Then, ORNL designed and constructed a new version of their vision-based sensing system better adapted to acquire process signals of the TPC induction welding process for process anomaly and defect detection. In BP2, the team accomplished key tasks & milestones to scale up multi-source ML for TPC aerostructure consolidation and assembly from the lab-coupon scale to the pilot-industrial scale. In BP2, the team demonstrated real time anomaly & defect detection via experiments performed by ORNL & RTRC. The team completed ML-optimization heating trials for TPC induction consolidation at Collins, and the team confirmed pilot industrial scale experimental data from Collins was compatible with the developed ML pipeline from RTRC. The team completed sub-element scale ML process optimization demonstration at RTRC, where the team leveraged RTRC’s robotic TPC welding setup to de-risk the ML process optimization by performing ML analysis of recorded temperatures to account for complex part features. Then, the team applied its ML-derived control strategies and ML process optimization framework at Collins to the pilot-industrial scale on a demo skin-stiffener part representative of a nacelle aerostructure fan cowl section. The key innovation is the AI/ML framework enabling effective process development of high performance, lightweight, energy efficient TPCs for composite aircraft structures.

36 MATERIALS SCIENCE↗

Dynamic Channel Assignments for Efficient Use of Aviation Spectrum Allocations

The demand for voice and data communications continues to rise with the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations throughout the National Airspace System (NAS). Recent studies have shown that the anticipated growing demand for spectrum resources will exceed the capacity of existing aviation spectrum allocations. Further, airspace configurations, via assignment of fixed channel allocations within standard service volumes, do not allow for the dynamic and efficient distribution of spectrum resources based on airspace demand; as a result, a new approach to aviation spectrum management is needed to support the forecasted needs of new airspace users. The National Aeronautics and Space Administration (NASA) is investigating applications of artificial intelligence (AI), machine learning (ML), and other advanced concepts to solve a dynamic constraint satisfaction problem which is analogous to the frequency assignment problem faced by aviation. Procedures and strategies for dynamic channel allocation can be borrowed from other large-scale mobile services (i.e., 4G/5G applications) and can provide a novel spectrum management approach that allows for the intelligent utilization of aviation spectrum throughout the airspace while maintaining the strict quality of service prescribed by aeronautical standards.

communications↗

Dynamic Channel Assignments for Efficient Use of Aviation Spectrum Allocations

The demand for voice and data communications continues to rise with the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations throughout the National Airspace System (NAS). Recent studies have shown that the anticipated growing demand for spectrum resources will exceed the capacity of existing aviation spectrum allocations. Further, airspace configurations, via assignment of fixed channel allocations within standard service volumes, do not allow for the dynamic and efficient distribution of spectrum resources based on airspace demand; as a result, a new approach to aviation spectrum management is needed to support the forecasted needs of new airspace users. The National Aeronautics and Space Administration (NASA) is investigating applications of artificial intelligence (AI), machine learning (ML), and other advanced concepts to solve a dynamic constraint satisfaction problem which is analogous to the frequency assignment problem faced by aviation. Procedures and strategies for dynamic channel allocation can be borrowed from other large-scale mobile services (i.e., 4G/5G applications) and can provide a novel spectrum management approach that allows for the intelligent utilization of aviation spectrum throughout the airspace while maintaining the strict quality of service prescribed by aeronautical standards.

Communications↗

Integrated Research Infrastructure Architecture Blueprint Activity (Final Report 2023)

The complexity of scientific pursuits is increasing rapidly with aspects that require dynamic integration of experiment, observation, theory, modeling, simulation, visualization, machine learning (ML), artificial intelligence (AI), and analysis. Research projects across the Department of Energy (DOE) are increasingly data and compute intensive. Innovative research teams are accelerating the pace of discovery by using high-performance computational and data tools in their research workflows and leveraging multiple research infrastructures. Additionally, several recent high-level U.S. government reports underscore the necessity of a new advanced computing ecosystem for international competitiveness and national security. International competitors are moving forward with major research infrastructure integration efforts that seek to capture a competitive advantage in the global innovation race. Owing to its unparalleled constellation of world-class experimental and observational facilities and high-performance and extreme-scale computational, data, and networking infrastructure, DOE is positioned to be a global leader in this new era of integrated science. However, this new integration paradigm will demand continuing evolution to ensure the U.S. remains a global leader in research and innovation. The DOE Office of Science (SC) has seized on the strategic importance of integration and has adopted a vision for Integrated Research Infrastructure (IRI): To empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation. To respond to the evolving computational requirements of research and the competitive international innovation landscape, experimental facilities could be connected with high performance computing resources for near real-time analysis, and resources should be provided for merging enormous and diverse data for AI/ML techniques and analysis.

97 MATHEMATICS AND COMPUTING↗

Environmentally Assisted Fatigue in Light Water Reactor Environment

This report summarizes the Environmentally Assisted Fatigue (EAF) research conducted at ANL under the US DOE Light Water Reactor Sustainability (LWRS) program. Starting from a rich background in theoretical and experimental EAF, ANL previously developed an approach to evaluate fatigue performance of reactor materials in light water reactor environments with the correction factor F en . The approach was based on a large body of experimental work performed at ANL and elsewhere, and was consistent with American Society of Mechanical Engineers (ASME)’s methodology governing the design and construction of reactor components. In recent years, the program was focused on component fatigue prediction and made several major and fundamental contributions in this area. These accomplishments help meet the needs identified by the industry concerning component level fatigue predictions in complex, transient conditions. The main contribution of the ANL program involved the development of a system-level model for estimating residual strain and life of nuclear reactor coolant system components under connected-system-thermal-mechanical boundary conditions. The goal was to predict the stress hotspots, strain residuals, strain amplitudes and the resulting fatigue lives. Thermal-mechanical stress analysis was performed considering thermal stratification and a design-basis reactor loading cycle. Based on the finite element (FE) model results, the strain residuals, strain amplitudes and resulting fatigue lives of reactor coolant system (RCS) components were predicted. The results show that some of the RCS components can have significantly different strain amplitudes, residual strain, and fatigue lives, despite having similar geometry and material. In addition, the simulated component-level strain profile can guide the selection of appropriate test inputs for conducting laboratory-scale EAF tests. Building upon the system-level model, ANL developed a digital twin (DT) framework to predict the structural states and associated fatigue life of components in real-time. This framework is a comprehensive system designed to predict the structural states and fatigue lives of reactor components. It includes multiple models and integrates artificial intelligence (AI), machine learning (ML), and FE based modeling tools to evaluate the structural states and fatigue lives.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Smart Pixels: In-pixel AI for on-sensor data filtering

We present a smart pixel prototype readout integrated circuit (ROIC) designed in CMOS 28 nm bulk process, with in-pixel implementation of an artificial intelligence (AI) / machine learning (ML) based data filtering algorithm designed as proof-of-principle for a Phase III upgrade at the Large Hadron Collider (LHC) pixel detector. The first version of the ROIC consists of two matrices of 256 smart pixels, each 25$\times$25 µm\textsuperscript{2} in size. Each pixel consists of a charge-sensitive preamplifier with leakage current compensation and three auto-zero comparators for a 2-bit flash-type ADC. The frontend is capable of synchronously digitizing the sensor charge within 25 ns. Measurement results show an equivalent noise charge (ENC) of $\sim$30e\textsuperscript{-} and a total dispersion of $\sim$100e\textsuperscript{-} The second version of the ROIC uses a fully connected two-layer neural network (NN) to process information from a cluster of 256 pixels to determine if the pattern corresponds to highly desirable high-momentum particle tracks for selection and readout. The digital NN is embedded in-between analog signal processing regions of the 256 pixels without increasing the pixel size and is implemented as fully combinatorial digital logic to minimize power consumption and eliminate clock distribution, and is active only in the presence of an input signal. The total power consumption of the neural network is $\sim$ 300 $\mu$W. The NN performs momentum classification based on the generated cluster patterns and even with a modest momentum threshold, it is capable of 54.4\% – 75.4\% total data rejection, opening the possibility of using the pixel information at 40MHz for the trigger. The total power consumption of analog and digital functions per pixel is $\sim$ 6 $\mu$W per pixel, which corresponds to $\sim$ 1 W/cm\textsuperscript{2} staying within the experimental constraints.

Parpillon, Benjamin↗

2025 Workshop on Envisioning Frontiers in AI and Computing for Biological Research: Position Papers

This workshop aims to identify key research directions for transforming biology using artificial intelligence (AI), machine learning (ML) and computational methods to facilitate the discovery of new behaviors, mechanisms, and designs of biological processes relevant to DOE missions, underpinning a broader U.S. bioeconomy. By developing novel AI/ML technologies to analyze and interpret complex biological data, researchers can organize and simulate biological processes at various scales as well as advance predictive understanding and manipulation of biological systems. This integration of computation, experimentation, and next-generation experimental technologies can lead to discoveries in new biological behaviors and mechanisms relevant to DOE missions. The focus is on how advanced computational and mathematical methods can impact this mission by exploring digital twins, foundation models, automated laboratory experiments, modeling of complex living systems, and data-driven approaches for the biodesign of plants and microbial systems. While data management is important, it is not the primary focus of this workshop, which will assess the current state, trends, and AI/ML challenges at the interface between biology and computational science to identify opportunities for high-impact research at their intersection. The goal is to define research needs and opportunities that align with biological sciences, computational sciences, and applied mathematics research.

59 BASIC BIOLOGICAL SCIENCES↗

Report for the DOE Office of Science Workshop on Envisioning Frontiers in AI and Computing for Biological Research

Artificial intelligence (AI), machine learning (ML), and high-performance computing (HPC) are poised to transform biological research, spurring innovation in biotechnology and biosystems design. "is transformation will bring an explosion of new capabilities to control the expression of genomic information in living organisms and harness that information to invent new biobased technologies (Jinek et al. 2012; NASEM 2025).

59 BASIC BIOLOGICAL SCIENCES↗

Power and Propulsion: Small Core Advanced Thermal Management Project Overview

This presentation is intended to provide an overview of the various efforts under the Advanced Air Transport Technology (AATT) project related to engine thermal management. Sustainable flight is the key motivator for all our efforts, and innovative thermal technologies play a crucial role in achieving this goal. The technologies described in this presentation can be applied to both traditional gas turbines and advanced cycles which utilize alternative fuels. Within our project, we utilize the capabilities of artificial intelligence (AI), machine learning (ML), and additive manufacturing. This ranges from using AI to generate heat exchanger fin topologies, to using ML for a reduction in computational cost which allows for a more thorough design exploration. Many times, the resulting topologies can only be realized through additive manufacturing techniques. Most of what’s presented is currently low TRL, but the intent is to achieve TRL 4 by the end of the project.

Propulsion↗

Smart Pixels: In-pixel AI for on-sensor data filtering

We present a smart pixel prototype readout integrated circuit (ROIC) designed in CMOS 28 nm bulk process, with in-pixel implementation of an artificial intelligence (AI) / machine learning (ML) based data filtering algorithm designed as proof-of-principle for a Phase III upgrade at the Large Hadron Collider (LHC) pixel detector. The first version of the ROIC consists of two matrices of 256 smart pixels, each 25$\times$25 $\mu$m$^2$ in size. Each pixel consists of a charge-sensitive preamplifier with leakage current compensation and three auto-zero comparators for a 2-bit flash-type ADC. The frontend is capable of synchronously digitizing the sensor charge within 25 ns. Measurement results show an equivalent noise charge (ENC) of $\sim$30e$^-$ and a total dispersion of $\sim$100e$^-$ The second version of the ROIC uses a fully connected two-layer neural network (NN) to process information from a cluster of 256 pixels to determine if the pattern corresponds to highly desirable high-momentum particle tracks for selection and readout. The digital NN is embedded in-between analog signal processing regions of the 256 pixels without increasing the pixel size and is implemented as fully combinatorial digital logic to minimize power consumption and eliminate clock distribution, and is active only in the presence of an input signal. The total power consumption of the neural network is $\sim$ 300 $\mu$W. The NN performs momentum classification based on the generated cluster patterns and even with a modest momentum threshold, it is capable of 54.4% - 75.4% total data rejection, opening the possibility of using the pixel information at 40MHz for the trigger. The total power consumption of analog and digital functions per pixel is $\sim$ 6 $\mu$W per pixel, which corresponds to $\sim$ 1 W/cm$^2$ staying within the experimental constraints.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Social Bias in AI and its Implications

Previous studies have documented many different types of biases that exist in artificial intelligence (AI) and machine learning (ML) systems. We reviewed the literature on AI and ML bias with a focus on social implications and found that bias in AI and ML can potentially have harmful social impacts on individuals and/or groups of people. By affecting people differently according to characteristics such as race, gender, or sexual orientation, AI and ML systems may lead to harm by exacerbating social inequities. We recount examples of issues that have occurred in systems that use technology that might be used at NASA and elsewhere so that similar issues might be identified and mitigated in future systems. We also provide interested parties with a gateway into existing work on social bias in AI and ML systems.

artificial intelligence (AI)↗

Artificial Intelligence and Machine Learning Applications in Modern Power Systems

Machine learning (ML) and artificial intelligence (AI) algorithms offer valuable tools for the analysis and interpretation of large datasets. These tools have the capability to uncover insights that may not be readily apparent within these datasets. In recent years, the integration of ML and AI has become increasingly prevalent in various applications within the power system domain. One of the earliest instances of machine learning in power systems can be traced back to demand forecasting, where artificial neural networks were employed for short-term load forecasting. In contemporary power systems, an abundance of high-resolution geospatial and temporal data is generated at various time intervals, ranging from sub-seconds (Phasor Measurement Units or PMUs) to seconds (Supervisory Control and Data Acquisition or SCADA), minutes (Process Information or PI), and extending to days, months, and years. These datasets contain valuable information concerning system reliability and performance. This information holds the potential to offer critical insights into system operations, as well as solutions for predicting and mitigating contingencies to prevent cascading outages. Despite the immense power of machine learning tools, system operators, planners, and utilities often exhibit hesitancy in fully embracing AI-enabled system operations and planning. This cautious approach persists, even as numerous diverse applications of machine learning continue to emerge in the realm of power systems. In this chapter, our focus will delve deep into ML and AI applications tailored for power systems. These applications aim to furnish system operators with enhanced situational awareness and augment their decision-making capabilities, especially during challenging operating conditions. Specific areas of interest encompass root cause analyses of electricity market datasets and the strategic selection of representative samples from vast power system databases for training ML/AI models. Finally, the chapter will conclude with a short discussion on the future of ML/AI in power systems and possible directions that the industry is moving towards.

power system applications, machine learning (ML), ↗

Mixed-Precision S/DGEMM Using the TF32 and TF64 Frameworks on Low-Precision AI Tensor Cores

Using NVIDIA graphics processing units (GPUs) equipped with Tensor Cores has enabled the significant acceleration of general matrix multiplication (GEMM) for applications in machine learning (ML) and artificial intelligence (AI) and in high-performance computing (HPC) generally. The use of such power-efficient, specialized accelerators can provide a performance increase between 8 × and 20 ×, albeit with a loss in precision. However, a high level of precision is required in many large scientific and HPC applications, and computing in single or double precision is still necessary for many of these applications to maintain accuracy. Fortunately, mixed-precision methods can be employed to maintain a higher level of numerical precision while also taking advantage of the performance increases from computing with lower-precision AI cores. With this in mind, we extend the state of the art by using NVIDIA’s new TF32 framework. This new framework not only burdens some constraints of the previous frameworks, such as costly 32 16-bit castings but also provides an equivalent precision and performance by using a much simpler approach. We also propose a new framework called TF64 that attempts double-precision arithmetic with low-precision Tensor Cores. Although this framework does not exist yet, we validated the correctness of this idea and achieved an equivalent of 64-bit precision on 32-bit hardware.

Valero Lara, Pedro↗