論文の概要: A Corpus of Real Scam- and Spam-Call Conversations from an Active Voice-Agent Honeypot
- arxiv url: http://arxiv.org/abs/2609.29528v1
- Date: Tue, 25 Aug 2026 06:29:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-28 05:09:56.057214
- Title: A Corpus of Real Scam- and Spam-Call Conversations from an Active Voice-Agent Honeypot
- Title(参考訳): アクティブ音声応答型ハニーポットによる実写とスパムコール会話のコーパス
- Abstract要約: アクティブ音声エージェントハニーポットによって収集された実際のスカムコール会話のデータセットを提示する。
最初の53日間のウィンドウで10,015件のインバウンド詐欺とスパム電話を捉えた。
各コールには、ターンレベルの書き起こし、3チャンネルオーディオ、ターン毎のレイテンシテレメトリ、自動ラベルのレイヤが格納されている。
- 参考スコア(独自算出の注目度): 11.07975536046373
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Real conversations between fraudsters and their targets are among the most informative artifacts for studying telephone scams, yet also the scarcest: passive honeypots overwhelmingly capture automated messages and hang-ups, large-scale studies characterize call metadata rather than dialogue, and manual scam-baiting does not scale. We present a dataset of real scam-call conversations collected by an active voice-agent honeypot. Dedicated numbers are seeded into the lead-generation channels fraud operations harvest; inbound callers are answered by a low-latency conversational agent that adopts a plausible target persona and sustains the interaction while every call is recorded, transcribed, and automatically labeled. Over an initial 53-day window we captured 10,015 inbound scam and spam calls (6,601 with two or more turns): roughly 895 hours of audio and 328,869 transcribed turns from 5,665 distinct originating numbers. Under a holistic classifier the substantive calls are predominantly predatory-but-legal lead generation ("spam", about three in five), while about one in seven is an outright "scam" (949 in this snapshot). Each call carries a turn-level transcript, three-channel audio, per-turn latency telemetry, and layers of automatic labels, including a holistic scam/spam/legitimate judgment corroborated by independent human review (75% agreement on the binary decision). We describe the collection system, the record structure, and technical validation of the corpus's realism and label quality, including that the agent is recognized as non-human in only about 5% of engaged calls. We also benchmark established scam-detection methods, where detectors trained on published synthetic dialogue collapse in precision on real traffic.
- Abstract(参考訳): 受動的ハニーポットは、自動化されたメッセージやハングアップを圧倒的に捉え、大規模な研究は対話ではなく、メタデータを特徴付け、手動の詐欺行為はスケールしない。
アクティブ音声エージェントハニーポットによって収集された実際のスカムコール会話のデータセットを提示する。
有意なターゲットペルソナを採用し、すべての呼び出しが記録され、転写され、自動的にラベル付けされる間、対話を継続する低遅延の会話エージェントによって、インバウンドコールが応答される。
最初の53日間のウィンドウで、10,015件のインバウンド詐欺とスパム電話(6,601件、2回以上)を捉えました。
全体分類法の下では、実体的呼び出しは主に捕食的だが法的なリード生成("spam", three in 5")であり、一方7分の1は完全な"詐欺"(このスナップショットでは949)である。
それぞれの呼び出しには、ターンレベルの書き起こし、3チャンネルのオーディオ、ターン毎のレイテンシーテレメトリ、および独立した人間レビュー(バイナリ決定に関する75%の合意)によって裏付けられた全体論的詐欺/スパム/嫡出判断を含む自動ラベル層が格納されている。
本稿では,コーパスの現実性とラベル品質の収集システム,記録構造,技術的検証について述べる。
また,実トラフィックの高精度な合成対話の崩壊を訓練した検出器を,確立されたスカム検出手法のベンチマークを行った。
関連論文リスト
- TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection [55.35119005976305]
テレコム詐欺スクリプトは急速に進化し、日常的なサービス会話に似たように設計されている。
ベンチマークは、以前に確立されたテストセットを上書きすることなく、新しく観察された詐欺パターンを組み込まなければならない。
本稿では,Mixed-Tree Anti-Fraud Generation Pipelineで構築したTeleAntiFraud 2.0について述べる。
論文 参考訳(メタデータ) (2026-09-16T14:41:11Z) - The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls [8.497127814945147]
2024年、FCC(連邦通信委員会)は、電話消費者保護法(TCPA)に基づくAI生成音声を提出した。
ピアレビューされた測定では、不要なコールトラフィックがマシンによってどれだけ置かれているかは示されていない。
10,987件の通話が66日間にわたって録音された。
論文 参考訳(メタデータ) (2026-09-10T06:28:50Z) - Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate [6.015147919395612]
我々は10,211件のインバウンド詐欺とスパム通話の完全コーパスを分析した。
我々は、捕食的だが合法的なリード生成の大きな流れから、機密情報を誘惑するアウトライト詐欺を分離する。
営業時間(週末日より平日あたり6.6倍)
論文 参考訳(メタデータ) (2026-08-25T06:41:12Z) - Real-Time Voicemail Detection in Telephony Audio Using Temporal Speech Activity Features [1.066048003460524]
本稿では,事前学習したニューラル音声活動検出器の音声活動パターンから15の時間的特徴を抽出する軽量なアプローチを提案する。
2つの評価セットで合計764の電話録音を行い、96.1%の精度で合成する。
論文 参考訳(メタデータ) (2026-04-02T17:43:24Z) - Cloning a Conversational Voice AI Agent from Call\,Recording Datasets for Telesales [0.0]
通話記録のコーパスから会話音声AIエージェントをクローンする手法を提案する。
我々のシステムは電話で顧客に耳を傾け、合成音声で応答し、トップパフォーマンスの人間エージェントから学んだ構造化されたプレイブックに従う。
本発明のクローン化剤は、導入、製品コミュニケーション、販売ドライブ、異物処理、閉店を含む22の基準で、人為的エージェントに対して評価される。
論文 参考訳(メタデータ) (2025-09-05T07:36:12Z) - Characterizing Robocalls with Multiple Vantage Points [50.423738777985314]
苦情や通話量はまだ高いものの、無言通話は緩やかに減少傾向にある。
ロボコールがSTIR/SHAKENに適応していることがわかりました。
以上の結果から,電話スパムの特徴化と防止に向けた今後の取り組みの最も有望な方向性が浮かび上がっている。
論文 参考訳(メタデータ) (2024-10-22T18:54:12Z) - Jäger: Automated Telephone Call Traceback [45.67265362470739]
分散セキュアなコールトレースバックシステムであるJ"agerを紹介します。
J"agerは、部分的なデプロイであっても、数秒で呼び出しをトレースできる。
論文 参考訳(メタデータ) (2024-09-04T16:09:28Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。