論文の概要: Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matching
- arxiv url: http://arxiv.org/abs/2608.23920v1
- Date: Mon, 24 Aug 2026 23:54:08 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-26 14:09:34.665118
- Title: Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matching
- Title(参考訳): Wontopos Tablet 2: 語彙マッチングのない多言語・多モーダルメモリ検索
- Abstract要約: Tablet-2は、言語モデルのための長期記憶エンジンである。
検索パスには語彙マッチングもキーワードスコアリングも言語モデルも含まない。
クロスモーダル3600写真300枚を14ヶ国語で公開し、密度が言語独立を許さないことを示した。
- 参考スコア(独自算出の注目度): 6.056779285861064
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We measure tablet-2, a production long-term memory engine for language models, on the text benchmarks the field already uses and on cross-lingual retrieval of photographs stored with no text at all. Its retrieval path contains no lexical matching, no keyword scoring, and no language model of its own. On LongMemEval-S (500 questions) it scores 95.7% [93.4, 97.1]; on BEAM-1M (700 questions, 2.21M stored memories) 67.5% [64.8, 70.2]. Those are question-sampling intervals, not the run-to-run spread, which is an order of magnitude narrower. Most of the paper is about how little they mean alone. Holding engine, corpus, settings and judge fixed, changing only the reader moves LongMemEval-S by 2.0 points; changing only the re-ask budget moves BEAM-1M by 8.9. Neither is stated in the reports we compare against, and the second exceeds most gaps there, so we give that table as a placement and not a ranking. For the multimodal axis we run two controls. Against BM25, configured as strongly as we could, we reach 95.2% mean recall@5 over 70 store-and-query language cells where BM25 reaches 19.0% and is exactly zero in 54. On captionless photographs a lexical method has no document to score at all. Open dense baselines on 300 Crossmodal-3600 photographs in 14 languages show that density confers no language independence: one scores 91.0% on English and 4.7% on Russian from identical image vectors, and a multilingual variant collapses on Telugu and Swahili. Our spread across languages is 14.0 against their 27.5 and 27.7. Three results run against us and are reported at equal weight: low-resource languages degrade sharply (Swahili 53.0%, Telugu 64.0%), attaching captions lowers cross-lingual retrieval by 11.4 points, and one setting omitted into one stage of our own retrieval cost 37 points of Korean top-1 accuracy while leaving nine languages untouched.
- Abstract(参考訳): 言語モデル用の長期記憶エンジンであるTable-2を、すでに使用しているテキストベンチマークと、テキストを全く持たない画像の言語間検索に基づいて測定する。
検索パスには語彙マッチングもキーワードスコアリングも言語モデルも含まない。
LongMemEval-S (500の質問)では95.7% [93.4, 97.1]、BEAM-1M (700の質問、2.21Mの記憶)では67.5% [64.8, 70.2]である。
これらは質問サンプリング間隔であり、ラン・トゥ・ラン・スプレッドではなく、桁違いに狭くなる。
論文の大半は、彼らが単独でどれだけ少ないかに関するものです。
エンジン、コーパス、セッティング、そしてジャッジを固定し、リーダのみがLongMemEval-Sを2.0ポイント変更し、再割り当て予算のみをBEAM-1Mを8.9ポイント変更した。
比較したレポートにはどちらも記載されていないが、第2の表は、ほとんどのギャップを越えているので、その表を順位ではなく配置として与える。
マルチモーダル軸では、2つのコントロールを実行します。
出来る限り強く設定されたBM25に対して、95.2%の平均リコール@5が70以上のストア・アンド・クエリ言語セルに到達し、BM25は19.0%に達し、54では正確にゼロである。
キャプションのない写真では、語彙的な方法には、スコアする文書がまったくありません。
クロスモーダル3600の14言語で撮影された300枚の写真では、密度が言語独立性(英語版)を示さないことが示されており、同じ画像ベクトルから91.0%、ロシア語で4.7%、テルグ語とスワヒリ語で多言語変種が崩壊している。
言語間の拡散は27.5と27.7に対して14.0である。
低リソース言語は急激な劣化(スワヒリ53.0%、テルグ64.0%)、字幕を付けると言語横断検索が11.4ポイント低下し、1つの設定は韓国語のトップ1精度の37ポイントの1段階に省略され、9つの言語は未使用のままである。
関連論文リスト
- The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages [0.0]
大規模言語モデル(LLM)は、人々が印象を形成する方法をますます仲介する。
ほとんどの監視は英語で行われ、英語のクエリが代表画像を返すと仮定する。
11の北、バルト、中央ヨーロッパ市場から約66のブランドを問い合わせました。
論文 参考訳(メタデータ) (2026-06-22T11:05:43Z) - Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data [2.3424047967193826]
我々は23言語にまたがる100万以上の多ラベルサンプルからなる大規模合成学習コーパスを構築した。
DistilBERT (135Mパラメータ) から XLM-R-Large (560Mパラメータ) までの6つの多言語トランスフォーマーエンコーダを訓練・比較する。
GoEmotions (英語) とSemEval-2018 Task 1 E-c (英語,アラビア語,スペイン語) を用いて, ゼロショットの全モデルを評価する。
論文 参考訳(メタデータ) (2026-04-14T12:04:17Z) - Do Multilingual VLMs Reason Equally? A Cross-Lingual Visual Reasoning Audit for Indian Languages [0.0]
インド語における最初の言語間視覚的推論監査について紹介する。
MathVista、ScienceQA、MMMUの980の質問はヒンディー語、タミル語、テルグ語、ベンガル語、カンナダ語、マラタイ語に翻訳される。
英語からインド語に切り替えた場合、精度は9.8~25ポイント低下する。
論文 参考訳(メタデータ) (2026-03-23T05:56:02Z) - Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech [61.759910921200834]
言語間の文エンコーダは通常、数百の言語をカバーしている。
我々はOmniSONARを紹介した。OmniSONARは全言語、言語横断、言語横断の文埋め込みモデルである。
論文 参考訳(メタデータ) (2026-03-17T14:47:35Z) - A Multi-Language Object-Oriented Programming Benchmark for Large Language Models [61.267115598083315]
35の既存ベンチマークの調査では、3つの大きな不均衡が明らかになった。
85.7%は単一のプログラミング言語に重点を置いている。
94.3%は関数レベルまたはステートメントレベルのタスクのみを対象としている。
80%以上は平均10件未満のテストケースを含む。
論文 参考訳(メタデータ) (2025-09-30T11:30:08Z) - Scaling Speech Technology to 1,000+ Languages [66.31120979098483]
MMS(Massively Multilingual Speech)プロジェクトは、タスクに応じてサポート言語を10~40倍増やす。
主な材料は、一般に公開されている宗教文書の読解に基づく新しいデータセットである。
我々は,1,406言語,1,107言語用1つの多言語自動音声認識モデル,同一言語用音声合成モデル,4,017言語用言語識別モデルについて,事前学習したwav2vec 2.0モデルを構築した。
論文 参考訳(メタデータ) (2023-05-22T22:09:41Z) - On the Off-Target Problem of Zero-Shot Multilingual Neural Machine
Translation [104.85258654917297]
識別対象言語信号の符号化に失敗すると、オフターゲットとなり、語彙距離が近くなることが判明した。
多言語語彙構築のための言語認識語彙共有(LAVS)を提案する。
我々は11言語で多言語機械翻訳ベンチマーク実験を行った。
論文 参考訳(メタデータ) (2023-05-18T12:43:31Z) - No Language Left Behind: Scaling Human-Centered Machine Translation [69.28110770760506]
低レベルの言語と高レベルの言語のパフォーマンスギャップを狭めるためのデータセットとモデルを作成します。
何千ものタスクをトレーニングしながらオーバーフィッティングに対処するために,複数のアーキテクチャとトレーニングの改善を提案する。
本モデルでは,従来の最先端技術と比較して,BLEUの44%の改善を実現している。
論文 参考訳(メタデータ) (2022-07-11T07:33:36Z) - Few-shot Learning with Multilingual Language Models [66.49496434282564]
多様な言語群をカバーするバランスの取れたコーパス上で,多言語の自動回帰言語モデルを訓練する。
私たちの最大のモデルは、20以上の代表言語で数ショットの学習において、新しい最先端の技術を定めています。
本稿では,モデルがどこで成功し,失敗するかを詳細に分析し,特に言語間の文脈内学習を可能にすることを示す。
論文 参考訳(メタデータ) (2021-12-20T16:52:35Z) - Language Detection Engine for Multilingual Texting on Mobile Devices [0.415623340386296]
全世界で20億人以上のモバイルユーザーがソフトキーボードで複数の言語を入力している。
単言語キーボードでは、誤訂正された単語の38%が別の言語で有効である。
多言語タイピングのための高速で軽量で正確な言語検出エンジン(LDE)を提案する。
論文 参考訳(メタデータ) (2021-01-07T16:49:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。