論文の概要: Ligand-Conditioned Discrete Diffusion for Protein Sequence-Structure Co-Design
- arxiv url: http://arxiv.org/abs/2605.27413v1
- Date: Fri, 15 May 2026 04:45:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-01 02:55:43.02808
- Title: Ligand-Conditioned Discrete Diffusion for Protein Sequence-Structure Co-Design
- Title(参考訳): タンパク質配列構造共設計のためのLigand-Conditioned Discrete Diffusion
- Authors: Chen Wei, Fanding Xu, Minghao Sun, Zhiyuan Liu, Lin Wang, Tianrui Jia, Yihang Zhou, Yang Zhang,
- Abstract要約: textbfProtLiD$2$, a textbfProtein textbfLigand-conditioned textbfDiscrete textbfDiffusion model for protein sequence-structure co-design。
- 参考スコア(独自算出の注目度): 20.54191936517173
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Proteins perform their biological functions through three-dimensional structures encoded by amino acid sequences, and ligand-binding protein co-design requires models that generate sequence-structure compatible proteins under explicit ligand constraints. Although continuous diffusion and flow-based models support ligand-aware design in coordinate or latent spaces, existing discrete diffusion protein language models mainly operate over sequence or structure tokens without direct small-molecule conditioning. We introduce \textbf{ProtLiD$^2$}, a \textbf{Prot}ein \textbf{L}igand-conditioned \textbf{D}iscrete \textbf{D}iffusion model for protein sequence-structure co-design. ProtLiD$^2$ jointly generates amino-acid sequence and discrete structure tokens while incorporating ligand chemical and geometric information through geometry-aware cross-attention. Trained on over one million ligand-protein complexes, ProtLiD$^2$ extends masked discrete diffusion to ligand-aware functional protein design. We further propose maximum confidence-margin guided ReMask decoding, an inference-time self-correction strategy that retains confident predictions and remasks uncertain tokens. ProtLiD$^2$ improves global fold confidence over Complexa in whole-protein design, increasing TM-score from 0.672 to 0.802 and pLDDT from 64.55 to 73.00. In pocket co-design, ProtLiD$^2$ reduces active-site BB-RMSD from 3.46/3.40Å for FAIR/PocketGen to 1.97Å, and improves ligand-aware pass rates over PocketGen from 14.86% to 59.73% and from 6.08% to 23.49% under stricter docking thresholds. These results support ligand-conditioned discrete diffusion as an effective token-space framework for functional protein co-design. Code will be available at https://github.com/auroua/ProtLiD.
- Abstract(参考訳): タンパク質はアミノ酸配列によってコードされる3次元構造を通して生物学的機能を実行するが、リガンド結合タンパク質の共設計は、明示的なリガンド制約の下で配列構造互換タンパク質を生成するモデルを必要とする。
連続拡散およびフローベースモデルは座標空間や潜在空間におけるリガンド認識設計をサポートするが、既存の離散拡散タンパク質言語モデルは、直接小分子条件のないシーケンスや構造トークン上で主に機能する。
本稿では, タンパク質配列構造共設計のための {textbf{ProtLiD$^2$}, a \textbf{Prot}ein \textbf{L}igand- Conditioned \textbf{D}iscrete \textbf{D}iffusion modelを紹介する。
ProtLiD$^2$でアミノ酸配列と離散構造トークンを共同生成し、幾何認識のクロスアテンションを通じて配位子化学的および幾何学的情報を組み込む。
100万以上のリガンド-タンパク質複合体で訓練されたProtLiD$^2$は、リガンド-認識機能タンパク質設計への個別拡散を隠蔽する。
さらに,信頼度予測と不確実なトークンの再マスクを保持する推論時自己補正戦略である,最大信頼度マージン誘導ReMask復号を提案する。
ProtLiD$^2$は、タンパク質全体の設計におけるコンプレックスに対するグローバルな折りたたみ信頼性を改善し、TMスコアを0.672から0.802に、pLDDTを64.55から73.00に向上させる。
ポケット共同設計では、ProtLiD$^2$は、アクティブサイトBB-RMSDをFAIR/PocketGenの3.46/3.40ドルから1.97ドルに減らし、より厳密なドッキング閾値の下で14.86%から59.73%、および6.08%から23.49%に改善した。
これらの結果は、機能タンパク質共設計のための効果的なトークン空間フレームワークとして、配位子条件付き離散拡散をサポートする。
コードはhttps://github.com/auroua/ProtLiD.comで入手できる。
関連論文リスト
- Torsion-Space Diffusion for Protein Backbone Generation with Geometric Refinement [0.0]
新しいモデルでは、ねじれ角を識別してタンパク質のバックボーンを生成し、構築による完全な局所幾何学を保証する。
標準PDBタンパク質の実験は、100%結合長の精度を示し、構造的コンパクト性を大幅に改善した。
論文 参考訳(メタデータ) (2025-11-24T14:51:29Z) - SiDGen: Structure-informed Diffusion for Generative modeling of Ligands for Proteins [0.0]
マスク付きSMILES生成とポケット認識のための軽量な折りたたみ機能を統合したタンパク質条件拡散フレームワークSiDGenを提案する。
SiDGenは、タンパク質の埋め込みから粗い構造信号をプールする合理化モードと、より強い結合のために局所化された対のバイアスを注入するフルモードの2つの条件付けパスをサポートしている。
自動ベンチマークでは、SiDGenは高い妥当性、一意性、新規性を生み出し、ドッキングベースの評価において競合性能を達成し、適切な分子特性を維持する。
論文 参考訳(メタデータ) (2025-11-12T18:25:51Z) - ProteinAE: Protein Diffusion Autoencoders for Structure Encoding [64.77182442408254]
本稿では,新規かつ合理化されたタンパク質拡散オートエンコーダであるProteinAEを紹介する。
プロテインAEは、タンパク質のバックボーン座標を直接E(3)から連続的でコンパクトな潜在空間にマッピングする。
本研究では,既存のオートエンコーダよりも優れた,最先端の再構築品質を実現することを実証する。
論文 参考訳(メタデータ) (2025-10-12T14:30:32Z) - PPDiff: Diffusing in Hybrid Sequence-Structure Space for Protein-Protein Complex Design [20.033392739225658]
PPDiffは、任意のタンパク質標的に対するバインダーの配列と構造を共同で設計する拡散モデルである。
PPDiffは、我々の開発したシークエンス構造間ネットワーク上に、因果的注意層を持つ。
このモデルはPPBenchで事前訓練され、2つの現実世界のアプリケーションで微調整される。
論文 参考訳(メタデータ) (2025-06-13T02:39:14Z) - ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning [49.2607661375311]
本稿では,逆折り畳みモデルの計算的拡張性,自動化,継続的な自己改善を可能にする新しいフレームワークであるProteinZeroを提案する。
ProteinZeroは、タンパク質設計のすべての主要な指標において、既存の手法を大幅に上回っている。
特に、CATH-4.3上で実行されるRL全体は、報酬を含む3日以内に1つの8X GPUノードで実行できる。
論文 参考訳(メタデータ) (2025-06-09T06:08:59Z) - A Latent Diffusion Model for Protein Structure Generation [50.74232632854264]
本稿では,タンパク質モデリングの複雑さを低減できる潜在拡散モデルを提案する。
提案手法は, 高い設計性と効率性を有する新規なタンパク質のバックボーン構造を効果的に生成できることを示す。
論文 参考訳(メタデータ) (2023-05-06T19:10:19Z) - Structure-informed Language Models Are Protein Designers [69.70134899296912]
配列ベースタンパク質言語モデル(pLM)の汎用的手法であるLM-Designを提案する。
pLMに軽量な構造アダプターを埋め込んだ構造手術を行い,構造意識を付加した構造手術を行った。
実験の結果,我々の手法は最先端の手法よりも大きなマージンで優れていることがわかった。
論文 参考訳(メタデータ) (2023-02-03T10:49:52Z) - State-specific protein-ligand complex structure prediction with a
multi-scale deep generative model [68.28309982199902]
タンパク質-リガンド複合体構造を直接予測できる計算手法であるNeuralPLexerを提案する。
我々の研究は、データ駆動型アプローチがタンパク質と小分子の構造的協調性を捉え、酵素や薬物分子などの設計を加速させる可能性を示唆している。
論文 参考訳(メタデータ) (2022-09-30T01:46:38Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。