論文の概要: Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
- arxiv url: http://arxiv.org/abs/2610.09227v1
- Date: Tue, 06 Oct 2026 23:44:32 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 21:58:22.64821
- Title: Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
- Title(参考訳): ドリフトによる量子化:テキスト埋め込みのためのラベルなし混合精度後処理量子化
- Abstract要約: 混合精度のポストトレーニング量子化には、モジュールごとの感度信号が必要である。
ラベルフリーの代用として、量子化誘起表現のドリフトを測定する。
- 参考スコア(独自算出の注目度): 2.8029321044571773
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.
- Abstract(参考訳): 混合精度のポストトレーニング量子化にはモジュール単位の感度信号が必要である; テキスト埋め込みでは、明らかなもの -- 量子化時にモジュールがコストのかかる検索品質 -- は、デプロイがほとんど持たない関連ラベルを必要とする。
1つのモジュールを量子化し、コーパスを再エンコードし、出力埋め込みが完全に精度の高い位置からどこまで移動したかを記録する。
デプロイされた出力表現は、高密度のレトリバーをランク付けします。
5つの開発埋め込み装置、構成レベルのドリフトオーダー、マクロのホールトアウト検索品質に対する混合精度プランのサンプリング、0.911のスピアマン、キャリブレーションコーパスと使用可能なレジーム内の検索ドメイン間の感度トランスポート、モジュールドリフトはランク一貫性を持つが数値的にはそうではない。
この方法は、ラベルなし、検索なしのハードパックバイトの予算の下で1つの加算割当である。
従来のLieQ基準に対する事前登録方向仮説(3/3は主予算、崩壊なし)とドリフトスコア(2/3は両面のLieQ鋼材より2/3は低いが、主予算ドリフトは3点すべて(-0.99, -0.85, -1.01点)で同じ予算の均一精度よりも数値的に低く、モジュールと全体モデルドリフトが設計されている。
出力ドリフトは、普遍的に最適な割り当て目標ではなく、頑丈な粗い感度信号であり、転送されたサイン-ジオメトリー適応の破滅的な失敗を回避し、均一に崩壊するストレスのある予算で使用することができるが、強い均一な操作点周辺の微細な再分配は未解決のままである。
関連論文リスト
- Transferable Low-Rank Convolutional Bases for Onboarding Unseen Medical Imaging Modalities [0.0]
本研究では,厳密なドメインアウトプロトコルの下でのアンフォボーディング問題について検討する。
畳み込みバックボーンは、ソースモダリティに基づいて事前訓練され、永久に凍結され、その後、目に見えないモダリティを許容する必要がある。
完全な微調整は、ソースのモダリティを破滅的に劣化させることによってのみ、最も高い目標精度に達する。
論文 参考訳(メタデータ) (2026-07-18T17:09:05Z) - What Accuracy and Gradient Cosine Miss: Evaluating Feedback Alignment via Scale Stability, Reference Validity, and Depth Utility [48.10132234701036]
本稿では,3つのチェック(スケール安定性,参照妥当性,深度ユーティリティ)に基づく診断評価プロトコルを提案する。
複数のアーキテクチャや手法にまたがって、我々のプロトコルは広い校正マージンを持つ全ての障害を識別する。
論文 参考訳(メタデータ) (2026-06-19T06:04:17Z) - StippleDiffusion: Capacity-Constrained Stippling using Controlled Diffusion [41.23458880886283]
ストイップルパターンは、局所密度がターゲット画像を追跡する点集合であり、伝統的に密度ごとの反復によって生成される。
本稿では, 学習した局所的な点分布と, 推論時の連続的, 画像定義容量制約を同時に満足する最初の拡散型サンプリング器を提案する。
単一のトレーニングされたチェックポイントは、推論時に任意のターゲット密度を受け入れ、トレーニング中に見られなかったポイント予算に一般化し、出力ポイント数からほぼ独立してスタイップルを生成する。
論文 参考訳(メタデータ) (2026-05-15T10:12:42Z) - FLARE: Task-agnostic embedding model evaluation through a normalization process [11.999388314465191]
フローベースラベルレス表現埋め込み評価(FLARE)
11のデータセットと8の埋め込み器で、FLAREは監督ベンチマークでSpearmanの0.90ドルに達した。
論文 参考訳(メタデータ) (2026-04-19T09:31:52Z) - Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration [76.08899010904652]
CapCalは、ランキング決定から位置バイアスを機械的に分離する、トレーニング不要のフレームワークである。
シングルパス効率を保ちながら、トレーニング不要の手法で優れた性能を発揮する。
論文 参考訳(メタデータ) (2026-04-11T10:47:22Z) - Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation [3.018583625592182]
ほとんどの推定器は、全ての不確実性モードを単一の信頼スコアに分解し、いつより多くの計算を割り当てるか、あるいは推論を調整するべきかについての信頼性の高い推論を防ぐ。
非確実性誘導推論時間選択(Uncertainty-Guided Inference-Time Selection)は,データ駆動型(データ駆動型)とモデル駆動型不確実性を,深い特徴空間で直接的に解消する軽量な推論時間フレームワークである。
論文 参考訳(メタデータ) (2025-11-15T23:47:30Z) - Unsupervised Conformal Inference: Bootstrapping and Alignment to Control LLM Uncertainty [49.19257648205146]
生成のための教師なし共形推論フレームワークを提案する。
我々のゲートは、分断されたUPPよりも厳密で安定した閾値を提供する。
その結果は、ラベルのない、API互換の、テスト時間フィルタリングのゲートになる。
論文 参考訳(メタデータ) (2025-09-26T23:40:47Z) - Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time Shifts [80.32933059529135]
TTA(Test-Time Adaptation)メソッドが出現し、推論中にターゲット分布に適応する。
我々は、堅牢なM3ODの両不確実性を共同で最小化するために設計された、最初のTTAフレームワークであるDual Uncertainity Optimization (DUO)を提案する。
並列に,明瞭な意味的手がかりを持つ領域における幾何学的コヒーレンスを保存する意味認識型正規場制約を設計する。
論文 参考訳(メタデータ) (2025-08-28T07:09:21Z) - Ambiguity-aware Point Cloud Segmentation by Adaptive Margin Contrastive Learning [65.94127546086156]
本稿では,ポイントクラウド上のセマンティックセマンティックセグメンテーションのための適応的マージン比較学習法を提案する。
まず,両立度推定フレームワークにコントラスト学習を組み込んだAMContrast3Dを設計する。
共同トレーニングの洞察に触発されて、並列にトレーニングされた2つのブランチとAMContrast3D++を統合することを提案する。
論文 参考訳(メタデータ) (2025-07-09T07:00:32Z) - Distribution-free binary classification: prediction sets, confidence
intervals and calibration [106.50279469344937]
分布自由条件における二項分類のための不確実性定量化(キャリブレーション、信頼区間、予測セット)の3つの概念について検討する。
固定幅と一様質量の両双対の双対確率に対する信頼区間を導出する。
我々の「三脚」定理の結果として、双有理確率に対するこれらの信頼区間は分布自由キャリブレーションに繋がる。
論文 参考訳(メタデータ) (2020-06-18T14:17:29Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。