Fugu-MT 論文翻訳(概要): Setting the Record Straight on Transformer Oversmoothing

論文の概要: Setting the Record Straight on Transformer Oversmoothing

arxiv url: http://arxiv.org/abs/2401.04301v2
Date: Tue, 13 Feb 2024 15:48:46 GMT
ステータス: 翻訳完了
システム内更新日: 2024-02-14 18:41:49.425494
Title: Setting the Record Straight on Transformer Oversmoothing
Title（参考訳）: Transformer Oversmoothing における記録線の設定
Authors: Gb\`etondji J-S Dovonon, Michael M. Bronstein, Matt J. Kusner
Abstract要約: トランスフォーマーベースのモデルは、最近、さまざまなドメインセットで大成功を収めています。近年の研究では、トランスフォーマーは本質的に低域通過フィルタであり、徐々に入力を過度に過度に処理すると主張している。実際、トランスフォーマーは本質的にローパスフィルタではない。代わりに、トランスフォーマーが過度に滑らかであるか否かは、更新方程式の固有スペクトルに依存する。
参考スコア（独自算出の注目度）: 39.478055825375
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Transformer-based models have recently become wildly successful across a diverse set of domains. At the same time, recent work has argued that Transformers are inherently low-pass filters that gradually oversmooth the inputs. This is worrisome as it limits generalization, especially as model depth increases. A natural question is: How can Transformers achieve these successes given this shortcoming? In this work we show that in fact Transformers are not inherently low-pass filters. Instead, whether Transformers oversmooth or not depends on the eigenspectrum of their update equations. Further, depending on the task, smoothing does not harm generalization as model depth increases. Our analysis extends prior work in oversmoothing and in the closely-related phenomenon of rank collapse. Based on this analysis, we derive a simple way to parameterize the weights of the Transformer update equations that allows for control over its filtering behavior. For image classification tasks we show that smoothing, instead of sharpening, can improve generalization. Whereas for text generation tasks Transformers that are forced to either smooth or sharpen have worse generalization. We hope that this work gives ML researchers and practitioners additional insight and leverage when developing future Transformer models.
Abstract（参考訳）: トランスフォーマーベースのモデルは最近、さまざまなドメインでかなり成功しています。同時に、最近の研究はトランスフォーマーは本質的に低域通過フィルタであり、徐々に入力を過度に過大評価すると主張している。これは一般化を制限し、特にモデル深度が増加すると心配である。この欠点を考えると、トランスフォーマーはこれらの成功をどうやって達成できるのか? 本研究では、トランスフォーマーは本質的に低域通過フィルタではないことを示す。代わりに、トランスフォーマーがオーバームースかどうかは、更新方程式の固有スペクトルに依存する。さらに, モデル深度の増加に伴い, 平滑化が一般化を損なうことはない。我々の分析は、過密化や階級崩壊の密接な関係の現象における先行研究を延長する。この解析に基づいて,フィルタ動作の制御を可能にする変圧器更新方程式の重みをパラメータ化するための簡易な手法を導出する。画像分類タスクでは、シャープ化の代わりにスムース化が一般化を改善することが示される。テキスト生成タスクでは、スムーズかシャープに強制されるトランスフォーマーは、より悪い一般化をもたらす。この研究によって、ML研究者や実践者が将来のTransformerモデルを開発する際に、さらなる洞察と活用を得られることを期待しています。

論文の概要: Setting the Record Straight on Transformer Oversmoothing

関連論文リスト