論文の概要: RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar
- arxiv url: http://arxiv.org/abs/2607.13189v1
- Date: Tue, 14 Jul 2026 18:39:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-16 16:39:12.566961
- Title: RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar
- Title(参考訳): RAGthoven at SemEval-2026 Task 1: Multi-Stage Pipeline Walks into a Benchmark and Barely Clears the Bar
- Abstract要約: RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chinese)。
RAGthovenは、創造的なテキスト生成を多段階の大規模言語モデル(LLM)パイプラインに分解する。
ReAct型シーケンシャルツールコール(Exp09)と自律型マルチブランチオーケストレーション(Exp10)の2つのエージェント変異を評価した。
- 参考スコア(独自算出の注目度): 0.1899342874107254
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chinese). RAGthoven decomposes creative text generation into a multi-stage large language model (LLM) pipeline (Planner, Best-of-N Writer, Reflector for self-critique, LLM-as-a-judge Judge) grounded in computational humor theory (Benign Violation Theory, Script-based Semantic Theory of Humor) and refined across ten experiments. In our final configuration, we augment the Planner with retrieval-augmented generation (RAG) from a curated joke corpus, seeding generation with diverse joke mechanisms. We also evaluate two agentic variants -- ReAct-style sequential tool-calling (Exp09) and autonomous multi-branch orchestration (Exp10) -- that expose the same four stages with a deterministic ConstraintAudit checker. Across four frontier models on a held-out 12-instance English sample, neither agentic variant produced outputs we judged superior to the non-agentic pipeline despite substantially higher tool-call budgets. RAGthoven shares Rank 1 with the Gemini 2.5 Flash baseline in all three languages, with overlapping organizer-reported confidence intervals. In Spanish, it leads the baseline by 42 raw Elo points (1182 vs. 1140), while in English (1045 vs. 1081) and Chinese (1045 vs. 1053) the baseline holds the higher raw rating within the same statistical tie. Together, these results suggest language-dependent diminishing returns from elaborate multi-stage prompt engineering and agentic scaffolding once a strong frontier model is in the loop.
- Abstract(参考訳): 本稿では,SemEval-2026 Task 1(MWAHAHA),Subtask A(英語,スペイン語,中国語)のためのシステムであるRAGthovenを紹介する。
RAGthovenは、創造的なテキスト生成を多段階の大規模言語モデル(LLM)パイプライン(Planner, Best-of-N Writer, Reflector for Self-critique, LLM-as-a-judge Judge)に分解し、計算的ユーモア理論(Benign Violation Theory, Script-based Semantic Theory of Humor)に基づいて10の実験で精査する。
最終構成では、最適化されたジョークコーパスからの検索拡張生成(RAG)によりプランナーを増強し、多様なジョーク機構を備えたシード生成を行う。
また、ReActスタイルのシーケンシャルツールコール(Exp09)と自律型マルチブランチオーケストレーション(Exp10)という2つのエージェント変数を評価し、同じ4つのステージを決定論的ConstraintAuditチェッカーで公開する。
12-instance の英語サンプルの4つのフロンティアモデルに対して,ツールコールの予算が著しく高いにもかかわらず,エージェントが生成した出力は非エージェントパイプラインより優れていると判断した。
RAGthovenは3つの言語すべてで、Gemini 2.5 Flashベースラインとランク1を共有しており、オーガナイザが報告した信頼区間が重複している。
スペイン語ではエロ点が42点(1182点=1140点)、英語(1045点=1081点)、中国語(1045点=1053点)である。
これらの結果から,強いフロンティアモデルがループ内にあると,言語に依存した高度な多段階のプロンプトエンジニアリングとエージェント足場からのリターンが低下することが示唆された。
関連論文リスト
- UOL@IDEM at BEA 2026 Shared Task 1: Neural Fusion and Feature-Rich Modeling for L1-Aware Vocabulary Difficulty Prediction [1.9746060146273674]
本稿では, BEA 2026 の L1-aware vocabulary difficulty 予測タスクに対する UOL@IDEM のクローズトトラック提案について述べる。
我々は、このタスクを回帰としてモデル化し、スペイン語、ドイツ語、マンダリン中国語の別系統を訓練する。
本システムは,多言語文脈表現と,周波数,表面形状,検索証拠,セマンティックアライメント,コグネート類似性,マスク付き言語モデル予測可能性などの特徴を結合する。
論文 参考訳(メタデータ) (2026-06-23T12:30:48Z) - Benchmarking Speech-to-Speech Translation Models [55.00303727199927]
音声音声翻訳(S2ST)は急速に進歩しているが、オフライン評価には統一されたプロトコルが欠けている。
8次元にわたる46のメトリクスを統合するベンチマークフレームワークを導入する。
FLEURSとCVSSから1,248のモデル言語構成でデプロイする。
論文 参考訳(メタデータ) (2026-06-02T07:01:33Z) - MIPIAD: Multilingual Indirect Prompt Injection Attack Defense with Qwen -- TF-IDF Hybrid and Meta-Ensemble Learning [0.7161783472741748]
MIPIADは英語とバングラ語で評価された防衛フレームワークである。
これは、Qwen2.5-1.5BからLoRA(XLPID)、TF-IDFレキシカル特徴、検証調整アンサンブルを通じて微調整されたシーケンスを組み合わせたものである。
論文 参考訳(メタデータ) (2026-05-08T05:34:28Z) - From prompting to evidence-based translation: A RAG+prompt system for Japanese-Chinese translation and its pedagogical potential [0.0]
本研究では,言語解析,埋め込み型検索,迅速な構築,LLM生成をベースモデルを変更することなく統合した検索強化型RAG+Prompt翻訳システムについて検討した。
我々は、RAG+Prompt翻訳システムにより、NMCCを含む文のJa-Zh翻訳を解釈可能で監査可能な方法で改善することが結論付けられた。
論文 参考訳(メタデータ) (2026-05-05T05:48:08Z) - PolyFrame at MWE-2026 AdMIRe 2: When Words Are Not Enough: Multimodal Idiom Disambiguation [0.533024001730262]
PolyFrameは、画像+テキストランキング(Subtask A)とテキストのみのキャプションランキング(Subtask B)の両方のための統合パイプラインである。
全てのモデルでは、凍ったCLIPスタイルの視覚言語エンコーダと、軽量モジュールのみを訓練する多言語BGE M3エンコーダが保持されている。
マルチリンガルブラインドテストでは,Subtask Aは0.35/0.73,Subtask Bは0.32/0.71であった。
論文 参考訳(メタデータ) (2026-02-20T23:07:55Z) - "Don't Teach Minerva": Guiding LLMs Through Complex Syntax for Faithful Latin Translation with RAG [0.5076419064097734]
本稿では,オープンソースのLarge Language Modelsを上位レベルのプロプライエタリシステムに統計的に匹敵する性能レベルに引き上げる,再現可能なドラフトベース改良パイプラインを提案する。
標準的なドメイン内テストセット(Rosenthal, 2023)と12世紀のラテン文字(2025)からなる新しいドメイン外テストセット(OOD)である。
論文 参考訳(メタデータ) (2025-11-03T11:11:27Z) - Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations [59.056367787688146]
本稿では, マルチリンガル数学推論 (xMR) LLM の探索と学習の先駆者である。
我々は10の異なる言語を含む最初の多言語数学推論命令データセットMGSM8KInstructを構築した。
翻訳を利用して、10個の異なる言語を含む最初の多言語数学推論命令データセットMGSM8KInstructを構築した。
論文 参考訳(メタデータ) (2023-10-31T08:09:20Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。