Search NASA⌕ Search

SEARCH · Search NASA

Results for “Video vision transformer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

ThermoPore: Predicting part porosity based on thermal images using deep learning

Part qualification is often a critical and labor-intensive process in additive manufacturing, particularly in the detection of defects such as porosity, which stands to benefit significantly from advancements in machine learning. We present a deep learning approach for quantifying and localizing ex-situ porosity within Laser Powder Bed Fusion fabricated samples utilizing in-situ thermal image monitoring data. Our goal is to build the real time porosity map of parts based on thermal images acquired during the build. The quantification task builds upon the established Convolutional Neural Network model architecture to predict pore count and the localization task leverages the spatial and temporal attention mechanisms of the novel Video Vision Transformer model to indicate areas of expected porosity. Our model for porosity quantification achieved a R 2 score of 0.57 and our model for porosity localization produced an average Intersection over Union (IoU) score of 0.32 and a maximum of 1.0. This work is setting the foundations of part porosity “Digital Twins” based on additive manufacturing monitoring data and can be applied downstream to reduce time-intensive post-inspection and testing activities during part qualification and certification. In addition, we seek to accelerate the acquisition of crucial insights normally only available through ex-situ part evaluation by means of machine learning analysis of in-situ process monitoring data.

Deep learning↗

ThermalTracker 3D

ThermalTracker-3D is a stereo-vision solution for evaluating flight tracks of birds and bats around offshore wind turbines. Using a pair of thermal video cameras, the technology remotely senses movement of animals and objects, day and night, near critical assets. It generates motion tracks by collapsing a sequence of video frames from each camera into a single image that contains an entire flight track and then applies stereo-vision processing to transform the flight track into three dimensions. The approach allows tracking in near real time and automatically identifies moving objects based on features from the motion track and object size.

Matzner, Shari↗

Bottleneck Detection in Modular Construction Factories Using Computer Vision

The construction industry is increasingly adopting off-site and modular construction methods due to the advantages offered in terms of safety, quality, and productivity for construction projects. Despite the advantages promised by this method of construction, modular construction factories still rely on manually-intensive work, which can lead to highly variable cycle times. As a result, these factories experience bottlenecks in production that can reduce productivity and cause delays to modular integrated construction projects. To remedy this effect, computer vision-based methods have been proposed to monitor the progress of work in modular construction factories. However, these methods fail to account for changes in the appearance of the modular units during production, they are difficult to adapt to other stations and factories, and they require a significant amount of annotation effort. Due to these drawbacks, this paper proposes a computer vision-based progress monitoring method that is easy to adapt to different stations and factories and relies only on two image annotations per station. In doing so, the Scale-invariant feature transform (SIFT) method is used to identify the presence of modular units at workstations, and the Mask R-CNN deep learning-based method is used to identify active workstations. This information was synthesized using a near real-time data-driven bottleneck identification method suited for assembly lines in modular construction factories. This framework was successfully validated using 420 h of surveillance videos of a production line in a modular construction factory in the U.S., providing 96% accuracy in identifying the occupancy of the workstations and an F-1 Score of 89% in identifying the state of each station on the production line. The extracted active and inactive durations were successfully used via a data-driven bottleneck detection method to detect bottleneck stations inside a modular construction factory. The implementation of this method in factories can lead to continuous and comprehensive monitoring of the production line and prevent delays by timely identification of bottlenecks.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Automated Assembly Progress Monitoring in Modular Construction Factories Using Computer Vision-Based Instance Segmentation

Modular construction has recently gained interest as a transformative construction method. In this method, a large portion of the construction is performed inside factories, where processes are fast-paced and interdependent; therefore, any deviation from the schedule can delay the production. Such deviations are frequent in modular factories due to the labor-intensive nature of the tasks. This propagation of delays can be mitigated by continuously monitoring each process; however, current manual monitoring methods are laborious, and recently proposed contact sensor-based methods are intrusive to the work. In addition, recent computer vision-based monitoring methods inside factories are limited to detection algorithms that fail to provide the pixel-level accuracy required for assembly progress monitoring in highly occluded factory scenes, and they require a large number of manual annotations. Therefore, this paper proposes a method to monitor the installation of subassemblies in modular construction factories using mask R-CNN instance segmentation and improves the data efficiency of the model using a copy-paste augmentation method. This method was validated on the CCTV videos captured from a modular construction factory in the US, resulting in a 9% mAP improvement in segmentation.

computer vision↗

Generalist multimodal AI: A review of architectures, challenges and opportunities

Multimodal models are expected to be a critical component to future advances in artificial intelligence. Here, this field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

Artificial intelligence (AI)↗