TensorX
返回文献探索

Paper · arXiv 2311.10642

Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers

Vukasin Bozic, Danilo Dordervic, Daniele Coppola, Joseph Thommes

25 upvotesNovember 17, 2023arXiv 预印本
AI 摘要

Shallow feed-forward networks can emulate the performance of the attention mechanism in Transformers, as demonstrated by their competitive results on sequence-to-sequence tasks using knowledge distillation.

Transformerattention mechanismshallow feed-forward networksknowledge distillationIWSLT2017ablation studies

Abstract

This work presents an analysis of the effectiveness of using standard shallow feed-forward networks to mimic the behavior of the attention mechanism in the original Transformer model, a state-of-the-art architecture for sequence-to-sequence tasks. We substitute key elements of the attention mechanism in the Transformer with simple feed-forward networks, trained using the original components via knowledge distillation. Our experiments, conducted on the IWSLT2017 dataset, reveal the capacity of these "attentionless Transformers" to rival the performance of the original architecture. Through rigorous ablation studies, and experimenting with various replacement network types and sizes, we offer insights that support the viability of our approach. This not only sheds light on the adaptability of shallow feed-forward networks in emulating attention mechanisms but also underscores their potential to streamline complex architectures for sequence-to-sequence tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号