Search NASASearch

SEARCH · Search NASA

Results for “vision transformers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Vision Foundation Models in Remote Sensing: A survey

Artificial intelligence (AI) technologies have profoundly transformed the field of remote sensing (RS), revolutionizing data collection, processing, and analysis. Traditionally reliant on manual interpretation and task-specific models, RS research has been significantly enhanced by the advent of foundation models (FMs)—large-scale pretrained AI models capable of performing a wide array of tasks with unprecedented accuracy and efficiency. This article provides a comprehensive survey of FMs in the RS domain. We categorize these models based on their architectures, pretraining datasets, and methodologies. Through detailed performance comparisons, we highlight emerging trends and the significant advancements achieved by those FMs. Additionally, we discuss technical challenges, practical implications, and future research directions, addressing the need for high-quality data, computational resources, and improved model generalization. Our research also finds that pretraining methods, particularly self-supervised learning (SSL) techniques like contrastive learning (CL) and masked autoencoders (MAEs), remarkably enhance the performance and robustness of FMs. This survey aims to serve as a resource for researchers and practitioners by providing a panorama of advances and promising pathways for the continued development and application of FMs in RS.

data models

Foundation models for atomistic simulation of chemistry and materials

Conventional computational methods for modeling chemical and materials systems are limited by system size and timescale, forcing a trade-off between quantum-mechanical accuracy and the sampling needed for realistic observables. Large language and vision foundation models — pre-trained on massive datasets using transformer architectures — have revolutionized many fields. It is thus interesting to ask whether a foundation model — subject to suitable data, parameter scaling and training — could enable learned simulations of chemistry and materials. Here, in this study, we review the field of machine-learned interatomic potentials (MLIPs) and posit that scaling up large and diverse chemical and materials datasets and highly expressive architectures using advanced training strategies should result in models that are: more efficient, transferable, robust to out-of-distribution scenarios, and easier to fine-tune to a variety of downstream physical observables than models trained from scratch on small datasets corresponding to specific, targeted atomistic simulation tasks. We provide specific criteria for creating such large-scale MLIP foundation models, coordinated strategies for their development, evaluation and deployment, and highlight potential emergent capabilities that could transform predictive simulations in chemistry and materials science and accelerate discovery across multiple technological domains.

Yuan, Eric C.-Y. [University of California, Berkel

Dual use of image based tracking techniques: Laser eye surgery and low vision prosthesis

With a concentration on Fourier optics pattern recognition, we have developed several methods of tracking objects in dynamic imagery to automate certain space applications such as orbital rendezvous and spacecraft capture, or planetary landing. We are developing two of these techniques for Earth applications in real-time medical image processing. The first is warping of a video image, developed to evoke shift invariance to scale and rotation in correlation pattern recognition. The technology is being applied to compensation for certain field defects in low vision humans. The second is using the optical joint Fourier transform to track the translation of unmodeled scenes. Developed as an image fixation tool to assist in calculating shape from motion, it is being applied to tracking motions of the eyeball quickly enough to keep a laser photocoagulation spot fixed on the retina, thus avoiding collateral damage.

Juday, Richard D.

Dual Use of Image Based Tracking Techniques: Laser Eye Surgery and Low Vision Prosthesis

With a concentration on Fourier optics pattern recognition, we have developed several methods of tracking objects in dynamic imagery to automate certain space applications such as orbital rendezvous and spacecraft capture, or planetary landing. We are developing two of these techniques for Earth applications in real-time medical image processing. The first is warping of a video image, developed to evoke shift invariance to scale and rotation in correlation pattern recognition. The technology is being applied to compensation for certain field defects in low vision humans. The second is using the optical joint Fourier transform to track the translation of unmodeled scenes. Developed as an image fixation tool to assist in calculating shape from motion, it is being applied to tracking motions of the eyeball quickly enough to keep a laser photocoagulation spot fixed on the retina, thus avoiding collateral damage.

Juday, Richard D.

Aero-Propulsion Control Research in Support of NASA Aeronautics Research Strategic Thrusts

In the past few years, NASA (National Aeronautics and Space Administration) Aeronautics Research Mission Directorate (ARMD) has introduced and updated a “New Blueprint for Transforming Global Aviation†. This blueprint consists of six NASA Aeronautics Research Strategic Thrusts â€" “The updated vision is designed to ensure that through NASA's aeronautical research the United States will maintain its leadership in the sky and sustain aviation so that it remains a key economic driver and cultural touchstone for the nation.†In mid-2016, technology development roadmaps were developed by ARMD for each of the strategic research thrusts and these roadmaps are continually being updated based on feedback from the broader aeronautics research community. The NASA Aeronautics research vision is implemented through a set of 4 programs â€" Advanced Air Vehicles Program (AAVP), Airspace Operations and Safety Program (AOSP), Integrated Aviation Systems Program (IASP), and Transformative Aeronautics Concepts Program (TACP). The Intelligent Control and Autonomy Branch (ICAB) at NASA Glenn Research Center (GRC) in Cleveland, Ohio, is leading and participating in various projects in partnership with other organizations within GRC and across NASA, the U.S. aerospace industry, and academia to develop advanced controls and health management technologies for aero-propulsion systems that will help meet the goals of the ARMD programs. These efforts are primarily under the various projects under AAVP, AOSP, and TACP. The ICAB current research tasks in support of ARMD program are described in this paper. The paper provides motivation, background, technical approach and recent accomplishments for these tasks, as well as a couple of tasks completed in the previous fiscal year.

Propulsion Control

The Adaptable and Resilient Safety System: The Human Factor in Future In-Time Aviation Safety Management Systems

In-time integrated safety management will be paramount for safely enabling the envisioned transformations of the future National Airspace System (NAS). The path for realizing the vision includes addressing the increasing need for advanced data analytics and fusion of aviation safety data, managed by human decision-makers. The paper describes safety management systems and its’ challenges, and how the concept of In-time Aviation Safety Management Systems addresses the need to ensure an adaptable and resilient future safety system in the envisioned transformed NAS. Finally, it discusses potential human factors challenges, including new human roles and responsibilities, new information and cognitive requirements, new intelligent technologies that change human-system interaction and coordination, and new design paradigms for human system integration and teaming.

L. Prinzel

The Adaptable and Resilient Safety System: The Human Factor in Future In-Time Aviation Safety Management Systems

In-time integrated safety management will be paramount for safely enabling the envisioned transformations of the future National Airspace System (NAS). The path for realizing the vision includes addressing the increasing need for advanced data analytics and fusion of aviation safety data, managed by human decision-makers. The paper describes safety management systems and its’ challenges, and how the concept of In-time Aviation Safety Management Systems addresses the need to ensure an adaptable and resilient future safety system in the envisioned transformed NAS. Finally, it discusses potential human factors challenges, including new human roles and responsibilities, new information and cognitive requirements, new intelligent technologies that change human-system interaction and coordination, and new design paradigms for human system integration and teaming.

Aviation Safety

Optical pattern recognition; Proceedings of the Meeting, Los Angeles, CA, Jan. 17, 18, 1989

Papers on optical pattern recognition are presented, covering topics such as the estimation of satellite pose and motion parameters using a neural net tracker, associative memory, optical implmentation of programmable neural networks, optoelectronic neural networks, dynamic autoassociative neural memory, heteroassociative memory, bilinear pattern recognition processors, optical processing of optical correlation plane data, and a synthetic discriminant function-based nonlinear optical correlator. Other topics include an interactive optical-digital image processor, geometric transformations for video compression and human teleoperator display, quasiconformal remapping for compensation of human visual field defects, hybrid vision for automated spacecraft landing, advanced symbolic and inference optical correlation filters, and a rotationally invariant holographic tracking system. Additional topics include the detection of rotational and scale-varying objects with a programmable joint transform correlator, a single spatial light modulator binary nonlinear optical correlator, optical joint transform correlation, linear phase coefficient composite filters, and binary phase-only filters.

Liu, Hua-Kuang

Digital Engineering Strategy Overview

Systems are changing and engineering practices must mind the balance between evolutionary and revolutionary change as we move towards increasingly agile processes, enabled by interconnected tools, to best provide for partnered collaboration. This is the first of many evolutions of the Goddard Digital Engineering strategy, in preparation for the NASA 2040 vision.

Digital Engineering

Digital Engineering at Goddard: Exploring the Digital Thread

Systems are changing and engineering practices must mind the balance between evolutionary and revolutionary change as we move towards increasingly agile processes, enabled by interconnected tools, to best provide for partnered collaboration. The Digital Thread is a foundational capability for a functional Digital Engineering Ecosystem, which is Phase 1 of Goddard's Digital Engineering strategy in an effort to achieve Goddard 2040 vision.

Digital Engineering

Improving the Concrete Crack Detection Process via a Hybrid Visual Transformer Algorithm

Inspections of concrete bridges across the United States represent a significant commitment of resources, given their biannual mandate for many structures. With a notable number of aging bridges, there is an imperative need to enhance the efficiency of these inspections. This study harnessed the power of computer vision to streamline the inspection process. Our experiment examined the efficacy of a state-of-the-art Visual Transformer (ViT) model combined with distinct image enhancement detector algorithms. We benchmarked against a deep learning Convolutional Neural Network (CNN) model. These models were applied to over 20,000 high-quality images from the Concrete Images for Classification dataset. Traditional crack detection methods often fall short due to their heavy reliance on time and resources. This research pioneers bridge inspection by integrating ViT with diverse image enhancement detectors, significantly improving concrete crack detection accuracy. Notably, a custom-built CNN achieves over 99% accuracy with substantially lower training time than ViT, making it an efficient solution for enhancing safety and resource conservation in infrastructure management. These advancements enhance safety by enabling reliable detection and timely maintenance, but they also align with Industry 4.0 objectives, automating manual inspections, reducing costs, and advancing technological integration in public infrastructure management.

42 ENGINEERING

3-D sensing with polar exponential sensor arrays

The present computations for three-dimensional vision involve, in such cases as those of scaling for perspective and optic flow, their reduction to additive operations by the implicit logarithmic transformation of image coordinates. Expressions for such computations are derived and applied to illustrative examples of sensor design. The advantages of polar exponential arrays over X-Y rasters for binocular vision are noted to encompass the inference of range and three-dimensional position from local image velocity without knowledge of pixel location, provided that the relative velocity of the target and sensor are known by some other means.

Weiman, Carl F. R.

Optical calculation of correlation filters for a robotic vision system

A method is presented for designing optical correlation filters based on measuring three intensity patterns: the Fourier transform of a filter object, a reference wave and the interference pattern produced by the sum of the object transform and the reference. The method can produce a filter that is well matched to both the object, its transforming optical system and the spatial light modulator used in the correlator input plane. A computer simulation was presented to demonstrate the approach for the special case of a conventional binary phase-only filter. The simulation produced a workable filter with a sharp correlation peak.

Knopp, Jerome

A vision-based end-point control for a two-link flexible manipulator

The measurement and control of the end-effector position of a large two-link flexible manipulator are investigated. The system implementation is described and an initial algorithm for static end-point positioning is discussed. Most existing robots are controlled through independent joint controllers, while the end-effector position is estimated from the joint positions using a kinematic relation. End-point position feedback can be used to compensate for uncertainty and structural deflections. Such feedback is especially important for flexible robots. Computer vision is utilized to obtain end-point position measurements. A look-and-move control structure alleviates the disadvantages of the slow and variable computer vision sampling frequency. This control structure consists of an inner joint-based loop and an outer vision-based loop. A static positioning algorithm was implemented and experimentally verified. This algorithm utilizes the manipulator Jacobian to transform a tip position error to a joint error. The joint error is then used to give a new reference input to the joint controller. The convergence of the algorithm is demonstrated experimentally under payload variation. A Landmark Tracking System (Dickerson, et al 1990) is used for vision-based end-point measurements. This system was modified and tested. A real-time control system was implemented on a PC and interfaced with the vision system and the robot.

Obergfell, Klaus

Pictorial communication in virtual and real environments

Papers about the communication between human users and machines in real and synthetic environments are presented. Individual topics addressed include: pictorial communication, distortions in memory for visual displays, cartography and map displays, efficiency of graphical perception, volumetric visualization of 3D data, spatial displays to increase pilot situational awareness, teleoperation of land vehicles, computer graphics system for visualizing spacecraft in orbit, visual display aid for orbital maneuvering, multiaxis control in telemanipulation and vehicle guidance, visual enhancements in pick-and-place tasks, target axis effects under transformed visual-motor mappings, adapting to variable prismatic displacement. Also discussed are: spatial vision within egocentric and exocentric frames of reference, sensory conflict in motion sickness, interactions of form and orientation, perception of geometrical structure from congruence, prediction of three-dimensionality across continuous surfaces, effects of viewpoint in the virtual space of pictures, visual slant underestimation, spatial constraints of stereopsis in video displays, stereoscopic stance perception, paradoxical monocular stereopsis and perspective vergence. (No individual items are abstracted in this volume)

Ellis, Stephen R.

Convergent Aeronautics Solutions Project

NASA is committed to transforming our aviation system to best meet demands and opportunities of the future. With a vision of safe, efficient, flexible, and environmentally sustainable air transportation, the NASA Aeronautics Research Mission Directorate is conducting research and development to address future needs of the aviation community, the Nation, and the world. While our NASA Aeronautics vision and strategy reaches into the next 25 years and beyond, we recognize that our vision and strategy must be responsive to new discoveries and emerging markets. For this reason, we are empowering our research community to redefine the future of aviation by dreaming up convergent/transformative ideas and studying if those ideas are possible. By modeling the new NASA Aeronautics' Convergent Aeronautics Solutions (CAS) Project after the venture capital community, we created opportunities for teams of intrapreneurs to mature their ideas into concepts through rapid feasibility studies. The CAS Project expects teams to consider the complexities and potential benefits of multi-disciplinary solutions and to leverage technology advances from outside the field of aeronautics. We also expect teams to explore their concepts in a rapid, iterative manner that allows them to learn and adjust their research approach. Within a year or two, teams are responsible for reporting on the feasibility of their concept. The findings inform NASA Aeronautics strategic planning and further investment. This presentation will give an overview of CAS.

Transformative

Localization Using Visual Odometry and a Single Downward-Pointing Camera

Stereo imaging is a technique commonly employed for vision-based navigation. For such applications, two images are acquired from different vantage points and then compared using transformations to extract depth information. The technique is commonly used in robotics for obstacle avoidance or for Simultaneous Localization And Mapping, (SLAM). Yet, the process requires a number of image processing steps and therefore tends to be CPU-intensive, which limits the real-time data rate and use in power-limited applications. Evaluated here is a technique where a monocular camera is used for vision-based odometry. In this work, an optical flow technique with feature recognition is performed to generate odometry measurements. The visual odometry sensor measurements are intended to be used as control inputs or measurements in a sensor fusion algorithm using low-cost MEMS based inertial sensors to provide improved localization information. Presented here are visual odometry results which demonstrate the challenges associated with using ground-pointing cameras for visual odometry. The focus is for rover-based robotic applications for localization within GPS-denied environments.

Swank, Aaron J.