Vision-Language Models for Vision Tasks: A Survey Most visual recognition studies rely heavily on crowd-labelled data in deep neural networks (DNNs) training, and they usually train a DNN for each single visual recognition task, leading to a laborious and time-consuming visual recognition paradigm. To address the two challenges, Vision-Language Models (VLMs) have been intensively investigated recently, which learns rich vision-language correlation from web-scale image-text pairs that are almost...
Cited 79074 times
Cited 25433 times
Cited 5411 times
Cited 5352 times
Cited 4550 times
Cited 3626 times