DualToken-ViT: Position-aware Efficient Vision Transformer with Dual Token Fusion
Zhenzhen Chu, Jiayu Chen, Cen Chen +4 authors
DualToken-ViT combines the strengths of CNNs and ViTs to achieve efficient vision processing with enhanced global and position-aware features.