論文の概要: EPIR: An Efficient Patch Tokenization, Integration and Representation Framework for Micro-expression Recognition
- arxiv url: http://arxiv.org/abs/2604.08106v1
- Date: Thu, 09 Apr 2026 11:24:17 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-10 18:34:05.880774
- Title: EPIR: An Efficient Patch Tokenization, Integration and Representation Framework for Micro-expression Recognition
- Title(参考訳): EPIR:マイクロ圧縮認識のための効率的なパッチトークン化・統合・表現フレームワーク
- Authors: Junbo Wang, Liangyu Fu, Yuke Li, Yining Zhu, Xuecheng Wu, Kun Hu,
- Abstract要約: 我々は、EPIR(EPIR)の効率的なパッチトークン化、統合、表現フレームワークを提案する。
EPIRは高い認識性能と低い計算複雑性のバランスをとることができる。
4つの人気のある公開データセットについて広範な実験を行う。
- 参考スコア(独自算出の注目度): 31.059157960656478
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Micro-expression recognition can obtain the real emotion of the individual at the current moment. Although deep learning-based methods, especially Transformer-based methods, have achieved impressive results, these methods have high computational complexity due to the large number of tokens in the multi-head self-attention. In addition, the existing micro-expression datasets are small-scale, which makes it difficult for Transformer-based models to learn effective micro-expression representations. Therefore, we propose a novel Efficient Patch tokenization, Integration and Representation framework (EPIR), which can balance high recognition performance and low computational complexity. Specifically, we first propose a dual norm shifted tokenization (DNSPT) module to learn the spatial relationship between neighboring pixels in the face region, which is implemented by a refined spatial transformation and dual norm projection. Then, we propose a token integration module to integrate partial tokens among multiple cascaded Transformer blocks, thereby reducing the number of tokens without information loss. Furthermore, we design a discriminative token extractor, which first improves the attention in the Transformer block to reduce the unnecessary focus of the attention calculation on self-tokens, and uses the dynamic token selection module (DTSM) to select key tokens, thereby capturing more discriminative micro-expression representations. We conduct extensive experiments on four popular public datasets (i.e., CASME II, SAMM, SMIC, and CAS(ME)3. The experimental results show that our method achieves significant performance gains over the state-of-the-art methods, such as 9.6% improvement on the CAS(ME)$^3$ dataset in terms of UF1 and 4.58% improvement on the SMIC dataset in terms of UAR metric.
- Abstract(参考訳): マイクロ圧縮認識は、現在、個人の実際の感情を得ることができる。
深層学習に基づく手法、特にトランスフォーマーに基づく手法は目覚ましい結果を得たが、これらの手法は多頭部自己注意におけるトークンの多さから計算の複雑さが高い。
さらに、既存のマイクロ圧縮データセットは小規模であるため、Transformerベースのモデルでは、効率的なマイクロ圧縮表現を学習することが困難である。
そこで本研究では,高認識性能と低計算複雑性のバランスをとることができるEPIR(Efficient Patch tokenization, Integration and Representation framework)を提案する。
具体的には,2つの標準シフトトークン化(DNSPT)モジュールを提案し,隣接する顔領域の画素間の空間的関係を学習する。
次に,複数のカスケードトランスフォーマーブロック間の部分トークンを統合するトークン統合モジュールを提案する。
さらに、まずトランスフォーマーブロックの注目度を向上し、自己トークンに対する注意計算の不要な焦点を減らすための識別トークン抽出器を設計し、動的トークン選択モジュール(DTSM)を用いて鍵トークンを選択し、より識別的なマイクロ表現表現をキャプチャする。
我々は,4つのパブリックデータセット(CASME II,SAMM,SMIC,CAS(ME)3)について広範な実験を行った。
その結果,CAS(ME)$^3$データセットが9.6%,SMICデータセットが4.58%向上した。
関連論文リスト
- Polynomial Mixing for Efficient Self-supervised Speech Encoders [50.58463928808225]
Polynomial Mixer (PoM) はマルチヘッド自己注意の代替品である。
PoMは下流音声認識タスクでその性能を達成する。
論文 参考訳(メタデータ) (2026-02-28T14:45:55Z) - Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit [45.18582668677648]
大規模言語モデルにおいて,トークン化剤を移植するためのトレーニング不要な手法を提案する。
それぞれの語彙外トークンを,共有トークンの疎線形結合として近似する。
我々は,OMPがベースモデルの性能を最良にゼロショット保存できることを示す。
論文 参考訳(メタデータ) (2025-06-07T00:51:27Z) - Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles [23.134664392314264]
トークン化は、言語モデル(LM)における多くの未理解の欠点と関連している。
本研究は, トークン化がモデルとバイトレベルのモデルを比較し比較することによって, モデル性能に与える影響について検討する。
本稿では,学習トークン分布と等価バイトレベル分布とのマッピングを確立するフレームワークであるByte-Token Representation Lemmaを紹介する。
論文 参考訳(メタデータ) (2024-10-11T23:30:42Z) - Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and Localization [52.87635234206178]
本稿では,多面的操作検出と局所化に適したMoNFAPという新しいフレームワークを提案する。
このフレームワークには2つの新しいモジュールが含まれている: Forgery-aware Unified Predictor (FUP) Module と Mixture-of-Noises Module (MNM)。
論文 参考訳(メタデータ) (2024-08-05T08:35:59Z) - ClusTR: Exploring Efficient Self-attention via Clustering for Vision
Transformers [70.76313507550684]
本稿では,密集自己注意の代替として,コンテンツに基づくスパースアテンション手法を提案する。
具体的には、合計トークン数を減少させるコンテンツベースの方法として、キーとバリュートークンをクラスタ化し、集約する。
結果として得られたクラスタ化されたTokenシーケンスは、元の信号のセマンティックな多様性を保持するが、より少ない計算コストで処理できる。
論文 参考訳(メタデータ) (2022-08-28T04:18:27Z) - Transformer-based Context Condensation for Boosting Feature Pyramids in
Object Detection [77.50110439560152]
現在の物体検出器は、通常マルチレベル特徴融合(MFF)のための特徴ピラミッド(FP)モジュールを持つ。
我々は,既存のFPがより優れたMFF結果を提供するのに役立つ,新しい,効率的なコンテキストモデリング機構を提案する。
特に,包括的文脈を2種類の表現に分解・凝縮して高効率化を図っている。
論文 参考訳(メタデータ) (2022-07-14T01:45:03Z) - Squeezeformer: An Efficient Transformer for Automatic Speech Recognition [99.349598600887]
Conformerは、そのハイブリッドアテンション・コンボリューションアーキテクチャに基づいて、様々な下流音声タスクの事実上のバックボーンモデルである。
Squeezeformerモデルを提案する。これは、同じトレーニングスキームの下で、最先端のASRモデルよりも一貫して優れている。
論文 参考訳(メタデータ) (2022-06-02T06:06:29Z) - Adaptive Fourier Neural Operators: Efficient Token Mixers for
Transformers [55.90468016961356]
本稿では,Fourierドメインのミキシングを学習する効率的なトークンミキサーを提案する。
AFNOは、演算子学習の原則的基礎に基づいている。
65kのシーケンスサイズを処理でき、他の効率的な自己認識機構より優れている。
論文 参考訳(メタデータ) (2021-11-24T05:44:31Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。