عودة إلى خريطة المعرفة
science
computer-science
الضمانة 75%

الخلاصة

PVT v2: Improved baselines with pyramid vision transformer Transformers have recently lead to encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs: (i) a linear complexity attention layer, (ii) an overlapping patch embedding, and (iii) a convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linearity and provides signif...

المصدر:
تم الإشارة إليه 2351 مرة