論文の概要: An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility
- arxiv url: http://arxiv.org/abs/2607.02212v2
- Date: Fri, 03 Jul 2026 23:05:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 17:33:48.976152
- Title: An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility
- Title(参考訳): 水溶性に対する化学および構造的寄与を特徴付ける付加的MLP-GNNフレームワーク
- Authors: Sampreeti Bhattacharya, Arkaprava Roy,
- Abstract要約: 水溶性は、早期の薬物発見の鍵となる性質である。
ほとんどの予測モデルは、物理化学的記述子と分子グラフ情報を単一の表現に統合する。
この2つの情報ソースをトレーニングを通じて分離する付加的なディープラーニングフレームワークを提案する。
- 参考スコア(独自算出の注目度): 0.40105987447353786
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Aqueous solubility is a key property in early-stage drug discovery, but most predictive models merge physicochemical descriptors and molecular graph information into a single representation, obscuring whether a prediction is driven by global chemistry, molecular structure, or both. We present an additive deep-learning framework that keeps these two sources of information separate throughout training: physicochemical descriptors are encoded by a multilayer perceptron (the chemical branch) and molecular graph topology by a graph neural network (the structural branch), with the two outputs combined only at the prediction stage through an additive model with an optional multiplicative interaction. This design provides a direct decomposition of chemical and structural components that can be examined separately after training. Furthermore, pretraining on the larger AqSolDB dataset and fine-tuning on the smaller BigSolDB2 dataset substantially improve accuracy and reduce run-to-run variations, indicating generalizability of the learned features from the data-rich settings. We further interpret the fitted model using best linear projections of the branch outputs, molecule-level embedding summaries across solubility classes, and atom-level GNNExplainer masks aggregated over functional groups. These analyses show that the chemical branch aligns with familiar physicochemical descriptors, while the structural branch captures graph-topological and functional-group patterns associated with solubility. Across both datasets, the framework attains competitive predictive performance while making the distinct roles of chemical and structural information more transparent.
- Abstract(参考訳): 水溶性は初期の薬物発見において重要な性質であるが、ほとんどの予測モデルは物理化学的記述子と分子グラフ情報を単一の表現に融合させ、予測が地球化学、分子構造、あるいはその両方によって駆動されているかどうかを判断する。
物理化学的記述子は多層パーセプトロン (化学分岐) と分子グラフトポロジー (分子グラフトポロジー) によって、グラフニューラルネットワーク (構造分岐) によって符号化される。
この設計は、化学成分と構造成分を直接分解し、訓練後に個別に検査することができる。
さらに、より大きなAqSolDBデータセットの事前トレーニングと、より小さなBigSolDB2データセットの微調整により、精度が大幅に向上し、実行時のバリエーションが減少し、データリッチな設定から学習した機能の一般化可能性を示している。
さらに、分岐出力の最良の線形射影、溶解度クラスにまたがる分子レベルの埋め込みサマリー、官能基に集約された原子レベルのGNNExplainerマスクを用いて、適合モデルを解釈する。
これらの分析により, 化学分岐はよく知られた物理化学的記述と一致し, 構造分岐は溶解度に関連するグラフトポロジカルおよび官能基パターンを捉えた。
両方のデータセットにわたって、このフレームワークは、化学的および構造的情報の明確な役割をより透明にしながら、競争力のある予測性能を達成する。
関連論文リスト
- Combining Graph Neural Networks and Mixed Integer Linear Programming for Molecular Inference under the Two-Layered Model [6.107266553770076]
我々は、GNNを学習手法として利用するmol-infer-GNN(mol-infer-GNN)に基づく分子推論フレームワークを開発する。
提案したGNNモデルでは,単純な構造であるにもかかわらず,いくつかの特性に対して満足度の高い学習性能が得られる。
論文 参考訳(メタデータ) (2025-07-05T06:57:37Z) - Structure-Aware Compound-Protein Affinity Prediction via Graph Neural Network with Group Lasso Regularization [11.87029706744257]
化合物タンパク質親和性を予測するためにグラフニューラルネットワーク(GNN)を実装したフレームワークを提案する。
我々は、グループラッソとスパースグループラッソカラー化正規化を用いて、構造認識損失関数を持つGNNを訓練する。
提案手法は,共通ノード情報と非共通ノード情報をスパースグループラッソに統合することにより,特性予測を改善した。
論文 参考訳(メタデータ) (2025-07-04T06:12:18Z) - Pre-trained Molecular Language Models with Random Functional Group Masking [54.900360309677794]
SMILESをベースとしたアンダーリネム分子アンダーリネム言語アンダーリネムモデルを提案し,特定の分子原子に対応するSMILESサブシーケンスをランダムにマスキングする。
この技術は、モデルに分子構造や特性をよりよく推測させ、予測能力を高めることを目的としている。
論文 参考訳(メタデータ) (2024-11-03T01:56:15Z) - Bi-level Contrastive Learning for Knowledge-Enhanced Molecule Representations [68.32093648671496]
分子に固有の二重レベル構造を考慮に入れたGODEを導入する。
分子は固有のグラフ構造を持ち、より広い分子知識グラフ内のノードとして機能する。
異なるグラフ構造上の2つのGNNを事前学習することにより、GODEは対応する知識グラフサブ構造と分子構造を効果的に融合させる。
論文 参考訳(メタデータ) (2023-06-02T15:49:45Z) - Atomic and Subgraph-aware Bilateral Aggregation for Molecular
Representation Learning [57.670845619155195]
我々は、原子とサブグラフを意識したバイラテラルアグリゲーション(ASBA)と呼ばれる分子表現学習の新しいモデルを導入する。
ASBAは、両方の種類の情報を統合することで、以前の原子単位とサブグラフ単位のモデルの限界に対処する。
本手法は,分子特性予測のための表現をより包括的に学習する方法を提供する。
論文 参考訳(メタデータ) (2023-05-22T00:56:00Z) - Do Large Scale Molecular Language Representations Capture Important
Structural Information? [31.76876206167457]
本稿では,MoLFormerと呼ばれる効率的なトランスフォーマーエンコーダモデルのトレーニングにより得られた分子埋め込みについて述べる。
実験の結果,グラフベースおよび指紋ベースによる教師付き学習ベースラインと比較して,学習された分子表現が競合的に機能することが確認された。
論文 参考訳(メタデータ) (2021-06-17T14:33:55Z) - Flexible dual-branched message passing neural network for quantum
mechanical property prediction with molecular conformation [16.08677447593939]
メッセージパッシングフレームワークに基づく分子特性予測のための二重分岐ニューラルネットワークを提案する。
本モデルでは,様々なスケールで異種分子の特徴を学習し,予測対象に応じて柔軟に学習する。
論文 参考訳(メタデータ) (2021-06-14T10:00:39Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。