論文の概要: MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- arxiv url: http://arxiv.org/abs/2608.04205v1
- Date: Tue, 04 Aug 2026 20:04:49 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 17:51:08.560675
- Title: MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- Title(参考訳): MatrAIx:830億のペルソナエージェントで世界をシミュレート
- Authors: Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song,
- Abstract要約: MatrAIxは、AIシステムとデジタル製品を異種ユーザでテストするための、人口規模でシミュレートされたユーザ評価基盤である。
第一に、ペルソナ8Bは1,290のカテゴリー次元で表される830億のペルソナレコードを含んでいる。
第2に、MatrAIx Playgroundは、ユーザがデジタル製品を評価し、対話する4つの環境を提供する。
第3に、MatrAIxは、コマース、ソフトウェア、ファイナンス、ヘルスケアを含む25のドメインにまたがる1010のアプリケーションタスクを提供する。
- 参考スコア(独自算出の注目度): 175.06699039914625
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
- Abstract(参考訳): AIシステムとデジタル製品の人間による評価は、コストがかかり、遅く、スケールが難しい。
オフライン評価はよりスケーラブルだが、人間の多様性と対話的な振る舞いを抽象化することが多い。
そこで我々は、異種ユーザとAIシステムおよびデジタル製品をテストするための、人口規模でシミュレーションされたユーザ評価基盤であるMatrAIxを紹介した。
第一に、ペルソナ8Bは1,290のカテゴリー次元で表される830億のペルソナレコードを含んでいる。
レコードは、相関属性を保存する依存グラフからサンプリングされるか、あるいは、人間が承認したプロファイルから抽出される。
我々は,599,847人の人体と40,000人の合成記録からなる,約100万人体からなる品質フィルタコアセットをリリースする。
第2に、MatrAIx Playgroundは、さまざまなユーザがデジタル製品を評価し、対話する4つの環境を提供する。
第3に、MatrAIxは、コマース、ソフトウェア、ファイナンス、ヘルスケアを含む25のドメインにまたがる1010のアプリケーションタスクを提供する。
8つのタスクで18,189回評価試験を行った。
ペルソナ・エージェントはクロード・オプス 4.8、GPT 5.5、クロード・ハイク 4.5の3つのLLMを動力とした。
結果として得られたフィードバックは、価格が上がった後のためらみ、AIアシスタントが失敗した後に継続する意思、レイテンシ寛容など、ペルソナの背景によって意思決定や嗜好がどのように異なるかを把握する。
まず,10つの行動特性と4つの環境にまたがるペルソナ付着性の評価を行った。
宣言された行動は366の治験(91.5%)で表されたか、正しく抑制された。
第2に,人間とLLMの審査員は,ヒトのグラウンド・ペルソナの抽出品質を評価した。
MatrAIxは、AIシステムとデジタル製品を評価するためのエンドツーエンドのインフラストラクチャを提供する。
関連論文リスト
- On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists [113.03797263688519]
多くの科学者は、AIレビュアーを研究を評価する専門知識のない確率的システムと見なしている。
既存のAIレビュアーの評価では、評決が人間の評決に合致するかどうかに焦点が当てられている。
論文 参考訳(メタデータ) (2026-05-20T03:33:55Z) - DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas [13.83414782465312]
物語完全合成ペルソナを合成するためのスケーラブルな生成エンジンであるDEEPPERSONAを紹介する。
まず、アルゴリズムによって、数百以上の階層的に組織化された属性からなる、最も大きな人間帰属分類を構築できる。
我々は、平均して数百の構造化属性と約1MBの物語テキストを持つコヒーレントで現実的なペルソナを条件付きで生成する。
論文 参考訳(メタデータ) (2025-11-10T17:37:56Z) - Everyone prefers human writers, including AI [0.0]
我々は,Raymond Queneaus Exercises Style (1947) を用いて帰属バイアスを測定する実験を行った。
人間は+13.7ポイント(pp)バイアス(コーエンのh = 0.28, 95%CI: 0.21-0.34)を示し、AIモデルは+34.3ポイントバイアス(h = 0.70, 95%CI: 0.65-0.76)を示した。
論文 参考訳(メタデータ) (2025-10-09T21:33:30Z) - SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants? [61.07963107032645]
大規模言語モデル(LLM)は、対話型アプリケーションでますます使われている。
人間の評価は、マルチターン会話におけるパフォーマンスを評価するためのゴールドスタンダードのままである。
我々は、909の注釈付き人間とLLMの会話を2つの対話タスクで行うベンチマークであるSimulatorArenaを紹介した。
論文 参考訳(メタデータ) (2025-10-06T23:17:44Z) - Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models [118.44328586173556]
MLLM(Multimodal Large Language Models)は視覚的理解タスクにおいて大きな進歩を見せている。
Human-MMEは、人間中心のシーン理解におけるMLLMのより総合的な評価を提供するために設計された、キュレートされたベンチマークである。
我々のベンチマークは、単一対象の理解を多対多の相互理解に拡張する。
論文 参考訳(メタデータ) (2025-09-30T12:20:57Z) - Adaptive Monitoring and Real-World Evaluation of Agentic AI Systems [3.215065407261898]
大規模言語モデルと外部ツールを組み合わせたマルチエージェントシステムは、研究機関からハイテイクドメインへと急速に移行している。
この「先進的な」続編は、アルゴリズムのインスタンス化や経験的な証拠を提供することで、そのギャップを埋める。
AMDMは擬似ゴールドリフトで異常検出遅延を12.3秒から5.6秒に減らし、偽陽性率を4.5%から0.9%に下げる。
論文 参考訳(メタデータ) (2025-08-28T15:52:49Z) - Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions [11.751234495886674]
LLMベースのデジタルツインシミュレーションは、AI、社会科学、デジタル実験の研究に大いに貢献する。
我々は、米国におけるN = 2,058$参加者(平均2.42時間)の代表サンプルを、合計500の質問を含む4つの波で調査した。
最初の分析では、データは高品質であることが示唆され、個人と集合レベルでの人間の振る舞いを良く予測するデジタルツインの構築が約束されている。
論文 参考訳(メタデータ) (2025-05-23T05:05:11Z) - Can Machines Imitate Humans? Integrative Turing-like tests for Language and Vision Demonstrate a Narrowing Gap [56.611702960809644]
3つの言語タスクと3つの視覚タスクで人間を模倣するAIの能力をベンチマークする。
次に,人間1,916名,AI10名を対象に,72,191名のチューリング様試験を行った。
模倣能力は従来のAIパフォーマンス指標と最小限の相関を示した。
論文 参考訳(メタデータ) (2022-11-23T16:16:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。