論文の概要: Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling
- arxiv url: http://arxiv.org/abs/2608.11444v2
- Date: Mon, 17 Aug 2026 20:10:15 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-19 15:55:59.075659
- Title: Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling
- Title(参考訳): 抗がん剤反応モデリングのための大規模AI-Readyデータ
- Authors: Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor, Thomas Brettin, Rick Stevens,
- Abstract要約: 我々は、IMPROVEベンチマークを、製薬ゲノムデータの大規模な統合により大幅に拡張する。
拡張されたリソースには、何百万もの薬物反応の測定、より広範なマルチオミクスのカバレッジ、化学多様性の大幅な増加が含まれている。
- 参考スコア(独自算出の注目度): 2.323664982546693
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. However, their predictive performance is often constrained by limited dataset scale and insufficient coverages of cancer and chemical spaces. In addition, inconsistent benchmarking practices hinder reliable comparison across models. Standardized frameworks, such as the Innovative Methodologies and New Data for Predictive Oncology Model Evaluation (IMPROVE) project, provide unified data schemas and evaluation protocols for consistent benchmarking, but improving model generalizability requires larger and more diverse training data. In this work, we substantially expand the IMPROVE benchmark through large-scale integration of pharmacogenomic data, primarily from PharmacoDB, together with additional smaller data sources. The expanded resource includes millions of drug response measurements, broader multi-omics coverage, and a major increase in chemical diversity, adding more than 50,000 compounds. To evaluate the impact of the new dataset compared to the original IMPROVE benchmark dataset, we trained DRP models using the two datasets and assess their prediction performance using a common test set and several evaluation strategies, including drug-blind, cancer-blind, and disjoint data splits. While cancer-blind performance remained comparable to the original benchmark, models trained on the expanded dataset showed consistent improvements in drug-blind and disjoint settings, indicating enhanced generalization to previously unseen compounds. These results position the expanded dataset as a community resource that provides a richer foundation for developing DRP models intended to aid in the discovery of novel anticancer drugs.
- Abstract(参考訳): 薬物反応予測(DRP)モデルは、薬理ゲノミクス研究の活発な領域であり、有効な抗がん剤の同定を加速する可能性がある。
しかしながら、それらの予測性能は、限られたデータセットスケールと、がんや化学空間のカバー不足によって制約されることが多い。
さらに、一貫性のないベンチマークプラクティスは、モデル間で信頼できる比較を妨げます。
IMPROVE(Innovative Methodologies and New Data for Predictive Oncology Model Evaluation)プロジェクトのような標準化されたフレームワークは、一貫したベンチマークのために統一されたデータスキーマと評価プロトコルを提供するが、モデルの一般化性を改善するには、より大きな、より多様なトレーニングデータが必要である。
本研究では,主にPharmacoDBからの薬理ゲノミクスデータの大規模統合によるIMPROVEベンチマークを,さらに小さなデータソースとともに大幅に拡張する。
拡張されたリソースには、数百万の薬物反応の測定、より広範なマルチオミクスのカバレッジ、化学多様性の大幅な増加、50,000以上の化合物の追加が含まれる。
従来のIMPROVEベンチマークデータセットと比較して,新しいデータセットの影響を評価するため,2つのデータセットを用いてDRPモデルを訓練し,共通のテストセットと薬物盲検,がん盲検,解離性データ分割を含むいくつかの評価戦略を用いて予測性能を評価した。
がん盲検のパフォーマンスは元々のベンチマークに匹敵するものの、拡張データセットでトレーニングされたモデルでは、薬物盲検と解離性の設定が一貫した改善が見られ、以前は見つからなかった化合物への一般化が促進された。
これらの結果は、拡張データセットを、新しい抗がん剤の発見を支援することを目的としたDRPモデルを開発するための、より豊かな基盤を提供するコミュニティリソースとして位置づけている。
関連論文リスト
- Investigating the Impact of Histopathological Foundation Models on Regressive Prediction of Homologous Recombination Deficiency [52.50039435394964]
回帰に基づくタスクの基礎モデルを体系的に評価する。
我々は5つの最先端基礎モデルを用いて、スライド画像全体(WSI)からパッチレベルの特徴を抽出する。
乳房、子宮内膜、肺がんコホートにまたがるこれらの抽出された特徴に基づいて、連続したRDDスコアを予測するモデルが訓練されている。
論文 参考訳(メタデータ) (2026-01-29T14:06:50Z) - VECT-GAN: A variationally encoded generative model for overcoming data scarcity in pharmaceutical science [32.92218213317144]
既存のデータセットは小さく、ノイズが多いため、有効性は制限されることが多い。
我々は、小型でノイズの多いデータセットを増強するために特別に設計された生成モデルを開発する。
我々は,ChEMBL 上で事前学習した VECT-GAN を pip パッケージとして利用できるようにした。
論文 参考訳(メタデータ) (2025-01-15T18:23:33Z) - MedDiffusion: Boosting Health Risk Prediction via Diffusion-based Data
Augmentation [58.93221876843639]
本稿では,MedDiffusion という,エンドツーエンドの拡散に基づくリスク予測モデルを提案する。
トレーニング中に合成患者データを作成し、サンプルスペースを拡大することで、リスク予測性能を向上させる。
ステップワイズ・アテンション・メカニズムを用いて患者の来訪者間の隠れた関係を識別し、高品質なデータを生成する上で最も重要な情報をモデルが自動的に保持することを可能にする。
論文 参考訳(メタデータ) (2023-10-04T01:36:30Z) - Drug Synergistic Combinations Predictions via Large-Scale Pre-Training
and Graph Structure Learning [82.93806087715507]
薬物併用療法は、より有効で安全性の低い疾患治療のための確立された戦略である。
ディープラーニングモデルは、シナジスティックな組み合わせを発見する効率的な方法として登場した。
我々のフレームワークは、他のディープラーニングベースの手法と比較して最先端の結果を達成する。
論文 参考訳(メタデータ) (2023-01-14T15:07:43Z) - Targeted-BEHRT: Deep learning for observational causal inference on
longitudinal electronic health records [1.3192560874022086]
RCTが確立したNull causal associationの因果モデリングについて検討した。
本研究では,観測研究用データセットと変換器ベースモデルであるTargeted BEHRTと2倍のロバストな推定手法を開発した。
本モデルでは,高次元EHRにおけるリスク比推定のベンチマークと比較し,RRの精度の高い推定結果が得られた。
論文 参考訳(メタデータ) (2022-02-07T20:05:05Z) - Bootstrapping Your Own Positive Sample: Contrastive Learning With
Electronic Health Record Data [62.29031007761901]
本稿では,新しいコントラスト型正規化臨床分類モデルを提案する。
EHRデータに特化した2つのユニークなポジティブサンプリング戦略を紹介します。
私たちのフレームワークは、現実世界のCOVID-19 EHRデータの死亡リスクを予測するために、競争の激しい実験結果をもたらします。
論文 参考訳(メタデータ) (2021-04-07T06:02:04Z) - Ensemble Transfer Learning for the Prediction of Anti-Cancer Drug
Response [49.86828302591469]
本稿では,抗がん剤感受性の予測にトランスファーラーニングを適用した。
我々は、ソースデータセット上で予測モデルをトレーニングし、ターゲットデータセット上でそれを洗練する古典的な転送学習フレームワークを適用した。
アンサンブル転送学習パイプラインは、LightGBMと異なるアーキテクチャを持つ2つのディープニューラルネットワーク(DNN)モデルを使用して実装されている。
論文 参考訳(メタデータ) (2020-05-13T20:29:48Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。