論文の概要: Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction
- arxiv url: http://arxiv.org/abs/2608.25548v1
- Date: Wed, 26 Aug 2026 09:01:43 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-27 14:15:15.693704
- Title: Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction
- Title(参考訳): タンパク質適合度予測のための直交投影によるタンパク質言語モデル埋め込みの解釈
- Authors: Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal, Henrike O. Heyne, Nadja Klein, Bernhard Y. Renard,
- Abstract要約: PLM埋め込みは, 生体化学的特性と相関するパターンをコードし, タンパク質の適合性予測への寄与を定量的に示す。
この計算効率の良いアプローチは、ここで考慮された特徴や埋め込みに限らず、タンパク質の適合性予測を超えた問題設定に容易に移行可能である。
- 参考スコア(独自算出の注目度): 2.365650942022388
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on several downstream tasks including protein fitness prediction. However, PLM embeddings are not directly interpretable and, thereby, it remains unclear what features they encode. To gain insight into which biochemical properties of the protein are driving the prediction, we leverage an orthogonal projection technique that removes linear effects of known tabular features from embeddings and extend it to high-order and interaction effects. In this way, we remove the effects of interpretable biochemical features from PLM embeddings. In an ablation study, we show that this leads to a decrease in performance for a downstream classifier trained only on the embeddings to predict protein fitness. In an additional evaluation, we find that these biochemical features explain a substantial part of the variance in the predictions of this classifier. Hence, we can show that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness. This computationally efficient approach is not limited to the features or embeddings considered here and is readily transferable to problem settings beyond protein fitness prediction.
- Abstract(参考訳): 近年,生物医学におけるタンパク質言語モデル(PLM)の普及が進んでいる。
それらの埋め込みは、タンパク質配列の豊富な数値表現を提供し、タンパク質の適合性予測を含む下流のいくつかのタスクで最先端のパフォーマンスを達成する。
しかし、PLMの埋め込みは直接解釈できないため、どの機能をコード化しているのかは不明だ。
タンパク質の生化学的性質がどのような予測を導いているかを理解するために,我々は,既知の表状特徴の線形効果を埋め込みから取り除き,高次および相互作用効果に拡張する直交射影法を利用する。
このようにして, PLM埋め込みから解釈可能な生化学的特徴の影響を除去する。
アブレーション研究では, タンパク質の適合性を予測するために, 組込みのみを訓練した下流分類器の性能低下が示唆された。
さらに,これらの生化学的特徴が,この分類器の予測のばらつきのかなりの部分について説明できることが示唆された。
したがって, PLM の埋め込みは, 生体化学的特性と相関するパターンを符号化し, タンパク質の適合性予測への寄与を定量的に示すことができる。
この計算効率の良いアプローチは、ここで考慮された特徴や埋め込みに限らず、タンパク質の適合性予測を超えた問題設定に容易に移行可能である。
関連論文リスト
- Sparse Autoencoders for Low-$N$ Protein Function Prediction and Design [0.0]
アミノ酸配列からのタンパク質機能の予測は、データスカース機構における中心的な課題である。
タンパク質言語モデル(pLM)は進化的インフォームド埋め込みとスパースオートエンコーダ(SAE)を提供することによって分野を進歩させた。
SAEは、24のシーケンスしか持たないが、フィットネス予測において、ESM2ベースラインよりも一貫して優れているか、競争している。
論文 参考訳(メタデータ) (2025-08-25T23:56:39Z) - Long-context Protein Language Modeling Using Bidirectional Mamba with Shared Projection Layers [76.95505296417866]
言語モデル(LM)の自己教師による訓練は、有意義な表現の学習や創薬設計において、タンパク質配列に大きな成功を収めている。
ほとんどのタンパク質LMは、短い文脈長を持つ個々のタンパク質に基づいて訓練されたトランスフォーマーアーキテクチャに基づいている。
そこで本研究では,選択的構造化状態空間モデルに基づく代替タンパク質であるBiMamba-Sに基づくLC-PLMを提案する。
論文 参考訳(メタデータ) (2024-10-29T16:43:28Z) - ProLLM: Protein Chain-of-Thoughts Enhanced LLM for Protein-Protein Interaction Prediction [54.132290875513405]
タンパク質-タンパク質相互作用(PPI)の予測は、生物学的機能や疾患を理解する上で重要である。
PPI予測に対する従来の機械学習アプローチは、主に直接的な物理的相互作用に焦点を当てていた。
PPIに適したLLMを用いた新しいフレームワークProLLMを提案する。
論文 参考訳(メタデータ) (2024-03-30T05:32:42Z) - NaNa and MiGu: Semantic Data Augmentation Techniques to Enhance Protein Classification in Graph Neural Networks [60.48306899271866]
本稿では,背骨化学および側鎖生物物理情報をタンパク質分類タスクに組み込む新しい意味データ拡張手法を提案する。
具体的には, 分子生物学的, 二次構造, 化学結合, およびタンパク質のイオン特性を活用し, 分類作業を容易にする。
論文 参考訳(メタデータ) (2024-03-21T13:27:57Z) - Multi-level Protein Representation Learning for Blind Mutational Effect
Prediction [5.207307163958806]
本稿では,タンパク質構造解析のためのシーケンシャルおよび幾何学的アナライザをカスケードする,新しい事前学習フレームワークを提案する。
野生型タンパク質の自然選択をシミュレートすることにより、所望の形質に対する突然変異方向を誘導する。
提案手法は,多種多様な効果予測タスクに対して,パブリックデータベースと2つの新しいデータベースを用いて評価する。
論文 参考訳(メタデータ) (2023-06-08T03:00:50Z) - Protein Representation Learning by Geometric Structure Pretraining [27.723095456631906]
既存のアプローチは通常、多くの未ラベルアミノ酸配列で事前訓練されたタンパク質言語モデルである。
まず,タンパク質の幾何学的特徴を学習するための単純かつ効果的なエンコーダを提案する。
関数予測と折り畳み分類の両タスクの実験結果から,提案した事前学習法は,より少ないデータを用いた最先端のシーケンスベース手法と同等あるいは同等であることがわかった。
論文 参考訳(メタデータ) (2022-03-11T17:52:13Z) - Leveraging Sequence Embedding and Convolutional Neural Network for
Protein Function Prediction [27.212743275697825]
タンパク質機能予測の主な課題は、大きなラベル空間とラベル付きトレーニングデータの欠如である。
これらの課題を克服するために、教師なしシーケンス埋め込みと深部畳み込みニューラルネットワークの成功を活用する。
論文 参考訳(メタデータ) (2021-12-01T08:31:01Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。