Search NASA⌕ Search

SEARCH · Search NASA

Results for “Video vision transformer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ThermoPore: Predicting part porosity based on thermal images using deep learning

Part qualification is often a critical and labor-intensive process in additive manufacturing, particularly in the detection of defects such as porosity, which stands to benefit significantly from advancements in machine learning. We present a deep learning approach for quantifying and localizing ex-situ porosity within Laser Powder Bed Fusion fabricated samples utilizing in-situ thermal image monitoring data. Our goal is to build the real time porosity map of parts based on thermal images acquired during the build. The quantification task builds upon the established Convolutional Neural Network model architecture to predict pore count and the localization task leverages the spatial and temporal attention mechanisms of the novel Video Vision Transformer model to indicate areas of expected porosity. Our model for porosity quantification achieved a R 2 score of 0.57 and our model for porosity localization produced an average Intersection over Union (IoU) score of 0.32 and a maximum of 1.0. This work is setting the foundations of part porosity “Digital Twins” based on additive manufacturing monitoring data and can be applied downstream to reduce time-intensive post-inspection and testing activities during part qualification and certification. In addition, we seek to accelerate the acquisition of crucial insights normally only available through ex-situ part evaluation by means of machine learning analysis of in-situ process monitoring data.

Deep learning↗

Hybrid vision activities at NASA Johnson Space Center

NASA's Johnson Space Center in Houston, Texas, is active in several aspects of hybrid image processing. (The term hybrid image processing refers to a system that combines digital and photonic processing). The major thrusts are autonomous space operations such as planetary landing, servicing, and rendezvous and docking. By processing images in non-Cartesian geometries to achieve shift invariance to canonical distortions, researchers use certain aspects of the human visual system for machine vision. That technology flow is bidirectional; researchers are investigating the possible utility of video-rate coordinate transformations for human low-vision patients. Man-in-the-loop teleoperations are also supported by the use of video-rate image-coordinate transformations, as researchers plan to use bandwidth compression tailored to the varying spatial acuity of the human operator. Technological elements being developed in the program include upgraded spatial light modulators, real-time coordinate transformations in video imagery, synthetic filters that robustly allow estimation of object pose parameters, convolutionally blurred filters that have continuously selectable invariance to such image changes as magnification and rotation, and optimization of optical correlation done with spatial light modulators that have limited range and couple both phase and amplitude in their response.

Juday, Richard D.↗

Optical pattern recognition; Proceedings of the Meeting, Los Angeles, CA, Jan. 17, 18, 1989

Papers on optical pattern recognition are presented, covering topics such as the estimation of satellite pose and motion parameters using a neural net tracker, associative memory, optical implmentation of programmable neural networks, optoelectronic neural networks, dynamic autoassociative neural memory, heteroassociative memory, bilinear pattern recognition processors, optical processing of optical correlation plane data, and a synthetic discriminant function-based nonlinear optical correlator. Other topics include an interactive optical-digital image processor, geometric transformations for video compression and human teleoperator display, quasiconformal remapping for compensation of human visual field defects, hybrid vision for automated spacecraft landing, advanced symbolic and inference optical correlation filters, and a rotationally invariant holographic tracking system. Additional topics include the detection of rotational and scale-varying objects with a programmable joint transform correlator, a single spatial light modulator binary nonlinear optical correlator, optical joint transform correlation, linear phase coefficient composite filters, and binary phase-only filters.

Liu, Hua-Kuang↗

The “Gravity” of Combustion, Fluid and Soft Matter Research

Over the next year, the National Academies of Science, Engineering and Medicine (NASEM) will be developing the report for the next Decadal Survey on Life and Physical Sciences Research in space 2023-2032. This document will be used by the Science Mission Directorate in the National Aeronautics and Space Administration (NASA) to provide the framework for the vision, priorities, and strategic plan and budget for NASA’s research efforts in the area of biological and physical sciences in space, especially with regards to effects of gravity or the lack thereof. Gravity can affect fluid motion, shapes interfacial boundaries, squeezes compressible volumes by their surroundings, and ultimately affects heat and mass transfer as well as chemical reactions. NASEM will be requesting input from the research community via white papers in order to generate a comprehensive vision and strategy for a decade of transformative science at the frontiers of science space. Note: Accompanying videos are included in this submission and are listed based on there corresponding slide. For viewing you will need to download each mp4 formatted video, total run time 4 mins 2 secs. w/ color/sound.

combustion↗

ThermalTracker 3D

ThermalTracker-3D is a stereo-vision solution for evaluating flight tracks of birds and bats around offshore wind turbines. Using a pair of thermal video cameras, the technology remotely senses movement of animals and objects, day and night, near critical assets. It generates motion tracks by collapsing a sequence of video frames from each camera into a single image that contains an entire flight track and then applies stereo-vision processing to transform the flight track into three dimensions. The approach allows tracking in near real time and automatically identifies moving objects based on features from the motion track and object size.

Matzner, Shari↗

Prototype Optical Correlator For Robotic Vision System

Known and unknown images fed in electronically at high speed. Optical correlator and associated electronic circuitry developed for vision system of robotic vehicle. System recognizes features of landscape by optical correlation between input image of scene viewed by video camera on robot and stored reference image. Optical configuration is Vander Lugt correlator, in which Fourier transform of scene formed in coherent light and spatially modulated by hologram of reference image to obtain correlation.

Scholl, Marija S.↗

Dual use of image based tracking techniques: Laser eye surgery and low vision prosthesis

With a concentration on Fourier optics pattern recognition, we have developed several methods of tracking objects in dynamic imagery to automate certain space applications such as orbital rendezvous and spacecraft capture, or planetary landing. We are developing two of these techniques for Earth applications in real-time medical image processing. The first is warping of a video image, developed to evoke shift invariance to scale and rotation in correlation pattern recognition. The technology is being applied to compensation for certain field defects in low vision humans. The second is using the optical joint Fourier transform to track the translation of unmodeled scenes. Developed as an image fixation tool to assist in calculating shape from motion, it is being applied to tracking motions of the eyeball quickly enough to keep a laser photocoagulation spot fixed on the retina, thus avoiding collateral damage.

Juday, Richard D.↗

Dual Use of Image Based Tracking Techniques: Laser Eye Surgery and Low Vision Prosthesis

With a concentration on Fourier optics pattern recognition, we have developed several methods of tracking objects in dynamic imagery to automate certain space applications such as orbital rendezvous and spacecraft capture, or planetary landing. We are developing two of these techniques for Earth applications in real-time medical image processing. The first is warping of a video image, developed to evoke shift invariance to scale and rotation in correlation pattern recognition. The technology is being applied to compensation for certain field defects in low vision humans. The second is using the optical joint Fourier transform to track the translation of unmodeled scenes. Developed as an image fixation tool to assist in calculating shape from motion, it is being applied to tracking motions of the eyeball quickly enough to keep a laser photocoagulation spot fixed on the retina, thus avoiding collateral damage.

Juday, Richard D.↗

Discrete Gabor Filters For Binocular Disparity Measurement

Discrete Gabor filters proposed for use in determining binocular disparity - difference between positions of same feature or object depicted in stereoscopic images produced by two side-by-side cameras aimed in parallel. Magnitude of binocular disparity used to estimate distance from cameras to feature or object. In one potential application, cameras charge-coupled-device video cameras in robotic vision system, and binocular disparities and distance estimates used as control inputs - for example, to control approaches to objects manipulated or to maintain safe distances from obstacles. Binocular disparities determined from phases of discretized Gabor transforms.

Weiman, Carl F. R.↗

Infrared sensors and systems for enhanced vision/autonomous landing applications

There exists a large body of data spanning more than two decades, regarding the ability of infrared imagers to 'see' through fog, i.e., in Category III weather conditions. Much of this data is anecdotal, highly specialized, and/or proprietary. In order to determine the efficacy and cost effectiveness of these sensors under a variety of climatic/weather conditions, there is a need for systematic data spanning a significant range of slant-path scenarios. These data should include simultaneous video recordings at visible, midwave (3-5 microns), and longwave (8-12 microns) wavelengths, with airborne weather pods that include the capability of determining the fog droplet size distributions. Existing data tend to show that infrared is more effective than would be expected from analysis and modeling. It is particularly more effective for inland (radiation) fog as compared to coastal (advection) fog, although both of these archetypes are oversimplifications. In addition, as would be expected from droplet size vs wavelength considerations, longwave outperforms midwave, in many cases by very substantial margins. Longwave also benefits from the higher level of available thermal energy at ambient temperatures. The principal attraction of midwave sensors is that staring focal plane technology is available at attractive cost-performance levels. However, longwave technology such as that developed at FLIR Systems, Inc. (FSI), has achieved high performance in small, economical, reliable imagers utilizing serial-parallel scanning techniques. In addition, FSI has developed dual-waveband systems particularly suited for enhanced vision flight testing. These systems include a substantial, embedded processing capability which can perform video-rate image enhancement and multisensor fusion. This is achieved with proprietary algorithms and includes such operations as real-time histograms, convolutions, and fast Fourier transforms.

Kerr, J. Richard↗

Controlling telerobots with video data and compensating for time-delayed video using Omniview

Remote viewing is critical for teleoperations, but the inherent limitations of standard video reduce the operator's effectiveness. These limitations have been compensated for in many ways, from using the operator's adaptability, to augmenting his capability with feedback from a variety of sensors and simulations. Omniview can overcome some of these limitations and improve the operator's efficiency without adding additional sensors or computational burden. It can minimize the potential collisions with facility equipment, provide peripheral vision, and display multiple images simultaneously from a single input device. The Omniview technology provides electronic pan, tilt, magnify, and rotational orientation within a hemispherical field-of-view without any moving parts. Image sizes, viewing directions, scale, offset, etc., may be adjusted to fit operator needs. This paper discusses the derivation of the image transformation, the design of the electronics, and two applications to telepresence that are under development. These are Video Emulated Tweening (VET), and Manipulator Guidance and Positioning (ManGAP). The VET effort uses Omniview to compensate for time-delayed video in teleoperation of remote vehicles. In ManGAP two Omniview systems are used to provide two sets of orientation vectors to points in the field-of-view (FOV). These vectors then provide absolute position information to both control the position of the telerobot, and to avoid collisions with the work sight equipment.

Kuban, Dan↗

Safe2Ditch Steer-To-Clear Development and Flight Testing

This paper describes a series of small unmanned aerial system (sUAS) flights performed at NASA Langley Research Center in April and May of 2019 to test a newly added Steer-to-Clear feature for the Safe2Ditch (S2D) prototype system. S2D is an autonomous crash management system for sUAS. Its function is to detect the onset of an emergency for an autonomous vehicle, and to enable that vehicle in distress to execute safe landings to avoid injuring people on the ground or damaging property. Flight tests were conducted at the City Environment Range for Testing Autonomous Integrated Navigation (CERTAIN) range at NASA Langley. Prior testing of S2D focused on rerouting to an alternate ditch site when an occupant was detected in the primary ditch site. For Steer-to-Clear testing, S2D was limited to a single ditch site option to force engagement of the Steer-to-Clear mode. The implementation of Steer-to-Clear for the flight prototype used a simple method to divide the target ditch site into four quadrants. An RC car was driven in circles in one quadrant to simulate an occupant in that ditch site. A simple implementation of Steer-to- Clear was programmed to land in the opposite quadrant to maximize distance to the occupant’s quadrant. A successful mission was tallied when this occurred. Out of nineteen flights, thirteen resulted in successful missions. Data logs from the flight vehicle and the RC car indicated that unsuccessful missions were due to geolocation error between the actual location of the RC car and the derived location of it by the Vision Assisted Landing component of S2D on the flight vehicle. Video data indicated that while the Vision Assisted Landing component reliably identified the location of the ditch site occupant in the image frame, the conversion of the occupant’s location to earth coordinates was sometimes adversely impacted by errors in sensor data needed to perform the transformation. Logged sensor data was analyzed to attempt to identify the primary error sources and their impact on the geolocation accuracy. Three trends were observed in the data evaluation phase. In one trend, errors in geolocation were relatively large at the flight vehicle’s cruise altitude, but reduced as the vehicle descended. This was the expected behavior and was attributed to sensor errors of the inertial measurement unit (IMU). The second trend showed distinct sinusoidal error for the entire descent that did not always reduce with altitude. The third trend showed high scatter in the data, which did not correlate well with altitude. Possible sources of observed error and compensation techniques are discussed.

Petty, Bryan J.↗

Advances in image compression and automatic target recognition; Proceedings of the Meeting, Orlando, FL, Mar. 30, 31, 1989

Various papers on image compression and automatic target recognition are presented. Individual topics addressed include: target cluster detection in cluttered SAR imagery, model-based target recognition using laser radar imagery, Smart Sensor front-end processor for feature extraction of images, object attitude estimation and tracking from a single video sensor, symmetry detection in human vision, analysis of high resolution aerial images for object detection, obscured object recognition for an ATR application, neural networks for adaptive shape tracking, statistical mechanics and pattern recognition, detection of cylinders in aerial range images, moving object tracking using local windows, new transform method for image data compression, quad-tree product vector quantization of images, predictive trellis encoding of imagery, reduced generalized chain code for contour description, compact architecture for a real-time vision system, use of human visibility functions in segmentation coding, color texture analysis and synthesis using Gibbs random fields.

Tescher, Andrew G.↗

Hybrid vision for automated spacecraft landing

A hyrbid real-time vision system concept is proposed for Mars lander guidance and control for the Mars Rover/Sample Return mission. The system includes digital and optical processing methods with high speed digital image warping to preprocess a video image for optical correlation. The system also includes position estimation by synthetic estimation filtering, image stabilization by joint transform optical correlation, and hazard identification from measurements of optical flow by sequential image subtraction.

Juday, Richard D.↗

Image remapping strategies applied as protheses for the visually impaired

Maculopathy and retinitis pigmentosa (rp) are two vision defects which render the afflicted person with impaired ability to read and recognize visual patterns. For some time there has been interest and work on the use of image remapping techniques to provide a visual aid for individuals with these impairments. The basic concept is to remap an image according to some mathematical transformation such that the image is warped around a maculopathic defect (scotoma) or within the rp foveal region of retinal sensitivity. NASA/JSC has been pursuing this research using angle invariant transformations with testing of the resulting remapping using subjects and facilities of the University of Houston, College of Optometry. Testing is facilitated by use of a hardware device, the Programmable Remapper, to provide the remapping of video images. This report presents the results of studies of alternative remapping transformations with the objective of improving subject reading rates and pattern recognition. In particular a form of conformal transformation was developed which provides for a smooth warping of an image around a scotoma. In such a case it is shown that distortion of characters and lines of characters is minimized which should lead to enhanced character recognition. In addition studies were made of alternative transformations which, although not conformal, provide for similar low character distortion remapping. A second, non-conformal transformation was studied for remapping of images to aid rp impairments. In this case a transformation was investigated which allows remapping of a vision field into a circular area representing the foveal retina region. The size and spatial representation of the image are selectable. It is shown that parametric adjustments allow for a wide variation of how a visual field is presented to the sensitive retina. This study also presents some preliminary considerations of how a prosthetic device could be implemented in a practical sense, vis-a-vis, size, weight and portability.

Johnson, Curtis D.↗

Pictorial communication in virtual and real environments

Papers about the communication between human users and machines in real and synthetic environments are presented. Individual topics addressed include: pictorial communication, distortions in memory for visual displays, cartography and map displays, efficiency of graphical perception, volumetric visualization of 3D data, spatial displays to increase pilot situational awareness, teleoperation of land vehicles, computer graphics system for visualizing spacecraft in orbit, visual display aid for orbital maneuvering, multiaxis control in telemanipulation and vehicle guidance, visual enhancements in pick-and-place tasks, target axis effects under transformed visual-motor mappings, adapting to variable prismatic displacement. Also discussed are: spatial vision within egocentric and exocentric frames of reference, sensory conflict in motion sickness, interactions of form and orientation, perception of geometrical structure from congruence, prediction of three-dimensionality across continuous surfaces, effects of viewpoint in the virtual space of pictures, visual slant underestimation, spatial constraints of stereopsis in video displays, stereoscopic stance perception, paradoxical monocular stereopsis and perspective vergence. (No individual items are abstracted in this volume)

Ellis, Stephen R.↗

Programmable Remapper

Input image remapped rapidly and accurately onto different coordinate grid. Analog/digital electronic image-processing system developed to warp input images onto arbitrary coordinate grids at video rates. Advantages of system include antialiasing effect of many-to-one data path and speed of lookup-table operation. Lookup tables reprogrammed easily with help of computer that generates table values from mathematical description of desired transformation. Applications include real-time corrections of distortions in input optics of image sensors, corrections for repeatable nonlinear scanning, and aiding persons of impaired vision by deliberately distorting images onto remaining functional portions of retinas.

Juday, Richard D.↗

Bottleneck Detection in Modular Construction Factories Using Computer Vision

The construction industry is increasingly adopting off-site and modular construction methods due to the advantages offered in terms of safety, quality, and productivity for construction projects. Despite the advantages promised by this method of construction, modular construction factories still rely on manually-intensive work, which can lead to highly variable cycle times. As a result, these factories experience bottlenecks in production that can reduce productivity and cause delays to modular integrated construction projects. To remedy this effect, computer vision-based methods have been proposed to monitor the progress of work in modular construction factories. However, these methods fail to account for changes in the appearance of the modular units during production, they are difficult to adapt to other stations and factories, and they require a significant amount of annotation effort. Due to these drawbacks, this paper proposes a computer vision-based progress monitoring method that is easy to adapt to different stations and factories and relies only on two image annotations per station. In doing so, the Scale-invariant feature transform (SIFT) method is used to identify the presence of modular units at workstations, and the Mask R-CNN deep learning-based method is used to identify active workstations. This information was synthesized using a near real-time data-driven bottleneck identification method suited for assembly lines in modular construction factories. This framework was successfully validated using 420 h of surveillance videos of a production line in a modular construction factory in the U.S., providing 96% accuracy in identifying the occupancy of the workstations and an F-1 Score of 89% in identifying the state of each station on the production line. The extracted active and inactive durations were successfully used via a data-driven bottleneck detection method to detect bottleneck stations inside a modular construction factory. The implementation of this method in factories can lead to continuous and comprehensive monitoring of the production line and prevent delays by timely identification of bottlenecks.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗