Back to knowledge graph
science
machine-learning
Confidence 75%

Vision-Language Models for Vision Tasks: A Survey

Vision-Language Models for Vision Tasks: A Survey Most visual recognition studies rely heavily on crowd-labelled data in deep neural networks (DNNs) training, and they usually train a DNN for each single visual recognition task, leading to a laborious and time-consuming visual recognition paradigm. To address the two challenges, Vision-Language Models (VLMs) have been intensively investigated recently, which learns rich vision-language correlation from web-scale image-text pairs that are almost...

Anonymous preview shows an excerpt only. Sign in to read the full item.

Cited 859 times
Vision-Language Models for Vision Tasks: A Survey | Awareness Public Knowledge