Back to knowledge graph
science
computer-science
Confidence 75%

Abstract

PVT v2: Improved baselines with pyramid vision transformer Transformers have recently lead to encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs: (i) a linear complexity attention layer, (ii) an overlapping patch embedding, and (iii) a convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linearity and provides signif...

Source:
Cited 2351 times
undefined | Awareness Public Knowledge