論文の概要: Approximating SPR Distance Between Phylogenetic Trees with Graph Neural Networks
- arxiv url: http://arxiv.org/abs/2607.18311v1
- Date: Fri, 17 Jul 2026 16:25:07 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-22 19:05:05.177704
- Title: Approximating SPR Distance Between Phylogenetic Trees with Graph Neural Networks
- Title(参考訳): グラフニューラルネットワークを用いた系統樹間のSPR距離の近似
- Authors: Renata Martins Castanheira, Miguel Bugalho, Cátia Vaz,
- Abstract要約: グラフニューラルネットワーク (GNN) が, トレーニング後の比較において, ほぼ一定時間で, サブツリープーンとリグラフトの距離を近似できるかどうかを検討する。
まず,4種以上の細菌種を対象に,UPGMAとNeighbor-Joiningで推定される864種の系統樹のデータセットを構築し,公開する。
第2に、中間点再ルートを含む再現可能な前処理パイプラインを構築し、樹木の深さを減少させる。
第三に、監督対象を検証する: 正確なSPRが抽出可能な小さな木では、未根のファンゴルン::SPR.dist はほぼ完全に相関している。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Comparing phylogenetic tree topologies is essential for understanding epidemic dynamics, yet biologically meaningful distances such as the Subtree Prune and Regraft (SPR) distance are NP-hard to compute and intractable on large datasets. We investigate whether a Graph Neural Network (GNN) can approximate SPR distances in near-constant time per comparison after training. Our contributions are fourfold. First, we build and publicly release a dataset of 864 phylogenetic trees inferred with UPGMA and Neighbor-Joining over four bacterial species, spanning up to 9{,}500 isolates, together with 388 labelled tree pairs. Second, we establish a reproducible pre-processing pipeline including midpoint re-rooting, which reduces tree depth and supplies the rooting required for exact distance computation and for the model's root-based features. Third, we validate the supervision target: on small trees, where exact SPR is tractable, the unrooted phangorn::SPR.dist heuristic correlates almost perfectly with the exact rooted distance computed by rspr (Pearson $0.98$--$0.99$), making it an excellent monotonic surrogate. Lastly, we train a Siamese Graph Isomorphism Network (GIN) regressor. In-distribution, i.e., held-out trees from the same species and size range as training, it explains roughly 87--90% of the variance ($R^2 \approx 0.87$ on a held-out split; $0.90 \pm 0.19$ under stratified cross-validation), with about four times lower error than a mean-predictor baseline, and shows partial transfer to unseen species ($R^2 \approx 0.37$). Its main limitation is extrapolation to trees larger than those seen in training, where accuracy collapses. The released dataset and the validated heuristic versus exact relationship provide a reproducible basis for scaling learned SPR approximation.
- Abstract(参考訳): 系統樹のトポロジーを比較することは、流行のダイナミクスを理解するのに不可欠であるが、Subtree Prune や Regraft (SPR) のような生物学的に意味のある距離は、大きなデータセット上で計算し、抽出できるNPハードである。
グラフニューラルネットワーク (GNN) が, 訓練後の比較において, ほぼ一定時間でSPR距離を近似できるかどうかを検討する。
私たちの貢献は4倍です。
まず,UPGMAおよびNeighbor-Joiningによって推定される864種の系統樹のデータセットを,最大9{,}500種の分離株と388種のラベル付き樹木ペアで構築し,公開する。
第2に、中間点再ルートを含む再現可能な前処理パイプラインを構築し、樹木の深さを低減し、正確な距離計算とモデルの根元的特徴に必要とされるルートを提供する。
第3に, 厳密なSPRが抽出可能である小木では, rspr (Pearson $0.98$-$0.99$) によって計算された厳密なルート距離とほぼ完全に相関し, 優れた単調サロゲートとなる。
最後に、Siamese Graph Isomorphism Network (GIN) 回帰器を訓練する。
In-distriion, in-distriion, which, held-out trees from the same species and size range as training, described about 87---90% of the variance (R^2 \approx 0.87$ on a held-out split; $0.90 \pm 0.19$ under stratified cross-validation), with almost 4 times lower error than a mean-predictor baseline, and showed partial transfer to unseen species (R^2 \approx 0.37$)。
その主な制限は、トレーニングで見られるものよりも大きい木への外挿であり、精度は崩壊する。
リリースデータセットと検証されたヒューリスティックと正確な関係は、学習したSPR近似をスケーリングするための再現可能な基盤を提供する。
関連論文リスト
- Multi-Modal Spatio-Temporal Graph Neural Network with Mixture of Experts for Soil Organic Carbon Prediction [0.0]
既存のアプローチでは、手作りのコモーダルを古典的なMLとペアリングするか、豊富なスペクトルと時間情報を見逃す単一モードのディープモデルである。
本稿では,SpTGNNとグリッドベースアーキテクチャの両方に対処するマルチテンポラルグラフニューラルネットワークであるSpTG-NNを紹介する。
微調整されたTerraMindエンコーダは、Sentinel-2、Sentinel-1、DEM信号からノード特徴を抽出する。
論文 参考訳(メタデータ) (2026-06-15T11:25:38Z) - Fast unsupervised ground metric learning with tree-Wasserstein distance [14.235762519615175]
教師なしの地上距離学習アプローチが導入されました
一つの有望な選択肢はワッサーシュタイン特異ベクトル(WSV)であり、特徴量とサンプルの間の最適な輸送距離を同時に計算する際に現れる。
木にサンプルや特徴を埋め込むことでWSV法を強化し,木-ワッサーシュタイン距離(TWD)を計算することを提案する。
論文 参考訳(メタデータ) (2024-11-11T23:21:01Z) - Biology-inspired joint distribution neurons based on Hierarchical Correlation Reconstruction allowing for multidirectional neural networks [0.49728186750345144]
低レベル差を除去できるHCR(Arnold correlation Reconstruction)に基づく新しい人工ニューロンが提案されている。
このような HCR ネットワークは $rho(y,z|x)$ のような確率分布(ジョイント)を伝播することもできる。
また、テンソル分解によるdirect $(a_mathbfj)$ Estimationのような追加のトレーニングアプローチも可能である。
論文 参考訳(メタデータ) (2024-05-08T14:49:27Z) - PhyloGFN: Phylogenetic inference with generative flow networks [57.104166650526416]
本稿では,系統学における2つの中核的問題に対処するための生成フローネットワーク(GFlowNets)の枠組みを紹介する。
GFlowNetsは複雑な構造をサンプリングするのに適しているため、木トポロジー上の多重モード後部分布を探索し、サンプリングするのに自然な選択である。
我々は, 実際のベンチマークデータセット上で, 様々な, 高品質な進化仮説を生成できることを実証した。
論文 参考訳(メタデータ) (2023-10-12T23:46:08Z) - Active-LATHE: An Active Learning Algorithm for Boosting the Error
Exponent for Learning Homogeneous Ising Trees [75.93186954061943]
我々は、$rho$が少なくとも0.8$である場合に、エラー指数を少なくとも40%向上させるアルゴリズムを設計し、分析する。
我々の分析は、グラフの一部により多くのデータを割り当てるために、微小だが検出可能なサンプルの統計的変動を巧みに活用することに基づいている。
論文 参考訳(メタデータ) (2021-10-27T10:45:21Z) - Towards an Understanding of Benign Overfitting in Neural Networks [104.2956323934544]
現代の機械学習モデルは、しばしば膨大な数のパラメータを使用し、通常、トレーニング損失がゼロになるように最適化されている。
ニューラルネットワークの2層構成において、これらの良質な過適合現象がどのように起こるかを検討する。
本稿では,2層型ReLUネットワーク補間器を極小最適学習率で実現可能であることを示す。
論文 参考訳(メタデータ) (2021-06-06T19:08:53Z) - Visualizing hierarchies in scRNA-seq data using a density tree-biased
autoencoder [50.591267188664666]
本研究では,高次元scRNA-seqデータから意味のある木構造を同定する手法を提案する。
次に、低次元空間におけるデータのツリー構造を強調する木バイアスオートエンコーダDTAEを紹介する。
論文 参考訳(メタデータ) (2021-02-11T08:48:48Z) - SGA: A Robust Algorithm for Partial Recovery of Tree-Structured
Graphical Models with Noisy Samples [75.32013242448151]
ノードからの観測が独立しているが非識別的に分散ノイズによって破損した場合、Ising Treeモデルの学習を検討する。
Katiyarら。
(2020) は, 正確な木構造は復元できないが, 部分木構造を復元できることを示した。
統計的に堅牢な部分木回復アルゴリズムであるSymmetrized Geometric Averaging(SGA)を提案する。
論文 参考訳(メタデータ) (2021-01-22T01:57:35Z) - Neural Enhanced Belief Propagation on Factor Graphs [85.61562052281688]
グラフィカルモデルは局所依存確率変数の構造的表現である。
最初にグラフニューラルネットワークを拡張してグラフを分解する(FG-GNN)。
そこで我々は,FG-GNNを連立して動作させるハイブリッドモデルを提案する。
論文 参考訳(メタデータ) (2020-03-04T11:03:07Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。