論文の概要: idSCD: Identifying Training Datasets through Semantic Correlation Descriptors
- arxiv url: http://arxiv.org/abs/2605.30462v1
- Date: Thu, 28 May 2026 18:38:15 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-01 20:56:50.175291
- Title: idSCD: Identifying Training Datasets through Semantic Correlation Descriptors
- Title(参考訳): idSCD:意味的相関記述子によるトレーニングデータセットの同定
- Authors: Andrada Gobeaja, Ionut Hodoroaga, Elena Burceanu, Marius Leordeanu,
- Abstract要約: データセットは、モデルの学習されたセマンティック相関構造にデータセット固有のトレースを残している、と我々は主張する。
意味相関記述子(SCD)に基づくホワイトボックスのセマンティックフィンガープリント手法を提案する。
本稿では,目標データセットがモデルのトレーニングミックスに含まれるか否かを判定する,実用的なSCDベースの会員スコアを提案する。
- 参考スコア(独自算出の注目度): 5.756453494739546
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Can a dataset be recognized from the spurious correlations it induces during training? We argue that datasets leave dataset-specific traces in a model's learned semantic correlation structure: incidental regularities that are predictive within a dataset, but not causal for the underlying task, can be internalized during training. We use this insight to study dataset-level membership inference, moving beyond existing methods that rely on behavioral or distributional evidence such as confidence scores, losses, margins, generated samples, or query responses. We introduce a white-box semantic fingerprinting approach based on semantic correlation descriptors (SCDs), which capture the semantic correlation structure learned by a model and make it comparable across dataset mixtures. In a controlled leave-one-dataset-out diagnostic, SCDs recover dataset-specific changes and perfectly separate matching from non-matching dataset pairs. We then propose a practical SCD-based membership score that tests whether a target dataset is part of a model's training mixture using only the model's SCD and the target dataset's standalone SCD, without requiring leave-one-dataset-out models. Across three diverse experimental settings, with dataset groups for natural language inference, emotion classification, and medical text classification, we test both the advantages and limitations of SCD-based membership inference with different degrees of semantic separation and keyword support between dataset splits. On average, the classifier based on this score achieves the highest performance and the lowest std, outperforming black-box baselines RMIA, Attack-P, and LiRA, as well as the white-box SIF baseline. These results show that dataset membership can be traced through internal semantic correlations, with the largest relative gain exceeding 60% in ROC-AUC when dataset groups expose distinct semantic particularities.
- Abstract(参考訳): データセットは、トレーニング中に引き起こされる刺激的な相関から認識できますか?
我々は、データセットが学習したセマンティック相関構造にデータセット固有のトレースを残していると主張している。
この洞察を用いてデータセットレベルのメンバシップ推定を研究し、信頼度スコアや損失、マージン、生成されたサンプル、クエリ応答といった行動的あるいは分散的な証拠に依存する既存の方法を超えていきます。
本研究では,セマンティック相関記述子(SCD)に基づくセマンティックフィンガープリント手法を提案する。
コントロールされたLeft-one-dataset-out診断では、SCDはデータセット固有の変更を回復し、非マッチングデータセットペアと完全に分離する。
そこで,本研究では,モデルSCDとモデルSCDのみを用いて,対象データセットがモデルのトレーニングミックスの一部であるかどうかを,一括データセットアウトモデルを必要としない,実用的なSCDベースの会員スコアを提案する。
自然言語推論,感情分類,医用テキスト分類のためのデータセット群を用いた3つの実験的設定において,SCDに基づくメンバシップ推論の利点と限界を,セマンティック分離の程度の違いとデータセット分割間のキーワードサポートで検証した。
このスコアに基づく分類器は、平均して最高性能と最低stdを達成し、White-box SIFベースラインと同様にRMIA、Attack-P、LiRAよりも優れる。
これらの結果から,データセット群が個々の意味的特異性を明らかにすると,ROC-AUCの相対的増加率が60%を超えることが示唆された。
関連論文リスト
- Statistical Embeddings for Similarity, Retrieval, and Interpretable Alignment of Numeric Tabular Datasets [0.0]
提案手法は,構造化探索データ解析記述子による数値データセットの特徴付けを行う。
カノニカル相関解析(CCA)のペナル化された定式化は、スパースで解釈可能な可変レベル対応を復元するために用いられる。
この手法は、汎用ベンチマーク、材料情報学、核グレードのグラファイトのキャラクタリゼーションにまたがる15のデータセットで評価される。
論文 参考訳(メタデータ) (2026-05-28T17:40:42Z) - SAS: Semantic-aware Sampling for Generative Dataset Distillation [55.27114962330541]
本稿では,コントラスト言語-画像事前学習(CLIP)をポストサンプリングのセマンティクスとして活用することで,データセット蒸留のセマンティック・アウェア・パースペクティブを導入する。
我々のゴールは、コンパクトであるだけでなく、意味的にクラス差別的で多様である蒸留データセットを得ることです。
論文 参考訳(メタデータ) (2026-05-18T08:05:46Z) - Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models [64.58262227709842]
ARISE(Attention-weighted Representation with Integrated Semantic Embeddings)が紹介される。
正確なクラスタリングのためにカテゴリデータのメトリック空間を補完するセマンティックアウェア表現を構築する。
8つのベンチマークデータセットの実験では、7つの代表的なデータセットよりも一貫した改善が示されている。
論文 参考訳(メタデータ) (2026-01-03T11:37:46Z) - AssayMatch: Learning to Select Data for Molecular Activity Models [21.367469322467418]
AssayMatchはデータ選択のためのフレームワークで、テストセットの関心に合わせたより小さく、より均質なトレーニングセットを構築する。
AssayMatchによって選択されたデータに基づいてトレーニングされたモデルが、完全なデータセットでトレーニングされたモデルの性能を上回ることができることを示す。
論文 参考訳(メタデータ) (2025-11-20T06:25:51Z) - CBW: Towards Dataset Ownership Verification for Speaker Verification via Clustering-based Backdoor Watermarking [85.68235482145091]
大規模音声データセットは貴重な知的財産となった。
本稿では,新しいデータセットのオーナシップ検証手法を提案する。
我々のアプローチはクラスタリングに基づくバックドア透かし(CBW)を導入している。
我々は,ベンチマークデータセットに対する広範な実験を行い,本手法の有効性とロバスト性を検証した。
論文 参考訳(メタデータ) (2025-03-02T02:02:57Z) - Measuring Bias of Web-filtered Text Datasets and Bias Propagation Through Training [22.53813258871828]
大規模言語モデル(LLM)の事前学習データセットのバイアスについて,データセット分類実験により検討した。
ニューラルネットワークは、単一のテキストシーケンスが属するデータセットを驚くほどよく分類することができる。
論文 参考訳(メタデータ) (2024-12-03T21:43:58Z) - Dataset Distillation-based Hybrid Federated Learning on Non-IID Data [18.226454363903446]
本稿では,データセット蒸留を統合して,独立および等分散(IID)データを生成するハイブリッド・フェデレーション学習フレームワークHFLDDを提案する。
特に、クライアントを異種クラスタに分割し、クラスタ内の異なるクライアント間でのデータラベルがバランスが取れないようにします。
このトレーニングプロセスは、従来のIDデータに対するフェデレーション学習に似ているため、非IIDデータがモデルトレーニングに与える影響を効果的に軽減する。
論文 参考訳(メタデータ) (2024-09-26T03:52:41Z) - Affinity Clustering Framework for Data Debiasing Using Pairwise
Distribution Discrepancy [10.184056098238765]
グループ不均衡(グループ不均衡)は、データセットにおける表現バイアスの主要な原因である。
本稿では、アフィニティクラスタリングを利用して、ターゲットデータセットの非保護および保護されたグループの表現のバランスをとるデータ拡張手法であるMASCを提案する。
論文 参考訳(メタデータ) (2023-06-02T17:18:20Z) - Data-SUITE: Data-centric identification of in-distribution incongruous
examples [81.21462458089142]
Data-SUITEは、ID(In-distriion)データの不連続領域を特定するためのデータ中心のフレームワークである。
我々は,Data-SUITEの性能保証とカバレッジ保証を実証的に検証する。
論文 参考訳(メタデータ) (2022-02-17T18:58:31Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。