論文の概要: MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
- arxiv url: http://arxiv.org/abs/2607.01982v1
- Date: Thu, 02 Jul 2026 10:13:19 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-03 19:45:08.785145
- Title: MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
- Title(参考訳): MolSight: 統一された化学画像理解のためのグラフ認識型視覚言語モデル
- Authors: Wenda Wang, Yihan Tong, Yuwei Hu, Zhewei Wei,
- Abstract要約: MolSightは、分子画像の理解を強化するために設計された、グラフ対応の視覚言語モデルフレームワークである。
実験の結果,MollSightは既存のVLM,分子LLM,および複数の化学視覚理解タスクにおける特殊ツールよりも優れていた。
- 参考スコア(独自算出の注目度): 25.768840369692608
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in tasks such as molecular design and drug discovery. However, these models struggle to fully capture the visual representation of molecular structures, limiting their potential. While existing molecular vision-language models (VLMs) show promise, they still face challenges in structural alignment and lack the necessary topological modeling for accurate molecular understanding. To address this, we propose MolSight, a graph-aware vision-language model framework designed to enhance the understanding of molecular images by VLMs. MolSight integrates a Molecular Topology Module to inject chemical-bond adjacency information into vision tokens, and a Molecular Grounding Module to align visual features with chemical symbolic semantics. Our experiments demonstrate that MolSight significantly outperforms existing VLMs, molecular LLMs, and specialized tools across multiple chemical visual understanding tasks, achieving a new level of molecular image reasoning.
- Abstract(参考訳): 分子構造と機能を理解するための統一的な枠組みとして分子大言語モデル(LLM)を用いることは、分子設計や薬物発見といったタスクにおける新しいトレンドとして現れつつある。
しかし、これらのモデルは分子構造の視覚的表現を完全に捉え、そのポテンシャルを制限するのに苦労する。
既存の分子ビジョン言語モデル(VLM)は将来性を示すが、構造的アライメントの課題に直面し、正確な分子理解に必要なトポロジ的モデリングを欠いている。
これを解決するために,VLMによる分子画像の理解を高めるために,グラフ対応の視覚言語モデルフレームワークであるMollSightを提案する。
MolSightは、化学結合した隣接情報を視覚トークンに注入する分子トポロジーモジュールと、視覚的特徴を化学記号的意味論と整合させる分子接地モジュールを統合している。
実験の結果,MollSightは既存のVLM,分子LLM,および複数の化学視覚理解タスクにおける特殊ツールを著しく上回り,新しいレベルの分子画像推論を実現していることがわかった。
関連論文リスト
- $\text{M}^{2}$LLM: Multi-view Molecular Representation Learning with Large Language Models [59.125833618091846]
分子構造ビュー,分子タスクビュー,分子規則ビューの3つの視点を統合した多視点フレームワークを提案する。
実験によると、$textM2$LLMは、分類タスクと回帰タスクをまたいだ複数のベンチマークで最先端のパフォーマンスを達成する。
論文 参考訳(メタデータ) (2025-08-12T05:46:47Z) - Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model [52.84455878597969]
Mol-LLaMAは、分子を中心とした一般的な知識を把握した大きな分子言語モデルである。
分子理解を改善するために,分子エンコーダの相補的な情報を統合するモジュールを提案する。
論文 参考訳(メタデータ) (2025-02-19T05:49:10Z) - MolMetaLM: a Physicochemical Knowledge-Guided Molecular Meta Language Model [19.458584012046646]
本稿では,分子メタ言語フレームワーク MolMetaLM を提案する。
我々は、同じS(分子)を共有する複数のS,P,O>知識トリプルとしてフォーマットされた分子特化メタ言語パラダイムを設計する。
異なる分子知識とノイズを導入することで、メタ言語パラダイムは数万の事前学習タスクを生成する。
論文 参考訳(メタデータ) (2024-11-23T09:27:38Z) - Learning Multi-view Molecular Representations with Structured and Unstructured Knowledge [14.08112359246334]
本稿では, 化学構造から多視点分子知識を抽出する表現学習モデルMV-Mol, バイオメディカルテキストからの非構造化知識, 知識グラフからの構造化知識について述べる。
MV-Molは分子特性予測に有効であることを示す。
論文 参考訳(メタデータ) (2024-06-14T08:48:10Z) - MultiModal-Learning for Predicting Molecular Properties: A Framework Based on Image and Graph Structures [2.5563339057415218]
MolIGは、画像とグラフ構造に基づいて分子特性を予測するための、新しいMultiModaL分子事前学習フレームワークである。
両者の分子表現の強さを融合させる。
ベンチマークグループ内の分子特性予測に関連する下流タスクでは、パフォーマンスが向上する。
論文 参考訳(メタデータ) (2023-11-28T10:28:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。