論文の概要: ProtStructQA: A Denotation Threshold in Protein Structural Reasoning
- arxiv url: http://arxiv.org/abs/2606.00451v1
- Date: Sat, 30 May 2026 00:42:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-02 21:34:28.467453
- Title: ProtStructQA: A Denotation Threshold in Protein Structural Reasoning
- Title(参考訳): ProtStructQA:タンパク質構造推論における記述閾値
- Authors: Aravind Mandiga, Guoming Li, Jin Lu, Ismailcem Budak Arpinar, Khaled Rasheed, Samuel E. Aggrey,
- Abstract要約: ProtStructQAはタンパク質構造質問応答の実行可能なベンチマークである。
信頼性、距離、予測誤差(PAE)、溶媒暴露、二次構造、トポロジ、接触に関する382.2Kの質問を公表する。
我々はQwen3-1.7BとQwen3-4Bの間に能力依存的な記述しきい値を求める。
- 参考スコア(独自算出の注目度): 2.997268954610233
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Protein-language systems are often evaluated by whether they generate plausible biological text, but a structural question has a sharper semantics: it denotes a measurement in a 3D coordinate system. We introduce ProtStructQA, an executable benchmark for protein structural question answering in which each natural-language question is generated from a hidden typed domain-specific language (DSL) program and the answer is obtained by executing that program on an AlphaFold-predicted structure. ProtStructQA releases 382.2K questions covering confidence, distances, predicted aligned error (PAE), solvent exposure, secondary structure, topology and contacts, and held-out compositions: a 330K active benchmark over 10K proteins from four species, plus a 52.2K hard-negative robustness pool. Without fine-tuning, we evaluate Qwen3 models from 0.6B to 8B under direct prompting, chain-of-thought, grammar-constrained executable voting, executable voting with chain-of-thought, and multi-turn ReAct-style tool use, and replicate the headline finding on Gemma-3-1B and Gemma-3-12B. We find a capability-dependent denotation threshold between Qwen3-1.7B and Qwen3-4B: below it, tool-mediated ReAct dominates because models often fail to produce executable denotations; above it, chain-of-thought flips from mostly harmful to strongly beneficial and becomes the strongest strategy on most splits. Parse-failure and family-level analyses show that the threshold is a transition from unparseable language to executable structural denotation, while grammar and execution remain selectively valuable for PAE and secondary-structure queries. ProtStructQA reframes scientific QA as compilation from language to measurement and provides a diagnostic testbed for when language models can map words to executable 3D structural measurements.
- Abstract(参考訳): タンパク質言語システムは、可塑性な生物学的テキストを生成するかどうかによってしばしば評価されるが、構造的疑問はよりシャープな意味を持ち、それは3D座標系における測定を意味する。
ProtStructQAはタンパク質構造質問応答の実行可能なベンチマークで、各自然言語質問が隠れ型付きドメイン固有言語(DSL)プログラムから生成され、そのプログラムをAlphaFoldで予測された構造上で実行することで解が得られる。
ProtStructQAは、信頼度、距離、予測整列誤差(PAE)、溶媒暴露、二次構造、トポロジー、接触、保留成分を含む382.2Kの質問を公表している。
微調整を行なわず,直接的プロンプト,チェーンオブシント,文法制約付き実行可能投票,チェーンオブシントによる実行可能投票,マルチターンReActスタイルツールの使用,Gemma-3-1BとGemma-3-12Bの見出し検索の再現により,0.6Bから8BまでのQwen3モデルを評価した。
Qwen3-1.7B と Qwen3-4B の間には機能依存的な記述しきい値が存在し、その下にはツールを介する ReAct が支配的である。
Parse-failure および family-level analysis は、しきい値がパーセブル言語から実行可能な構造記述への遷移であり、一方文法と実行は、PAE および二次構造クエリーに対して選択的に価値のあるままであることを示している。
ProtStructQAは、言語から測定へのコンパイルとして科学的なQAを再構成し、言語モデルが単語を実行可能な3D構造計測にマッピングできる際の診断テストベッドを提供する。
関連論文リスト
- Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies [4.738949927143789]
言語モデル自身の出力に対する連続的な自己学習は、フラット化のプロセスとして広く特徴づけられる。
この特徴が不完全であることを示す。
5つのモデルでの11世代にわたるセルフトレーニングでは、言語は均一にフラット化されていない。
論文 参考訳(メタデータ) (2026-05-20T01:44:47Z) - Structured Intent as a Protocol-Like Communication Layer: Cross-Model Robustness, Framework Comparison, and the Weak-Model Compensation Effect [0.0]
本稿では、AIモデル、言語、プロンプトフレームワーク間で、確実に構造化された意図表現がいかにユーザ目標を保っているかを検討する。
構造的プロンプトは、非構造的ベースラインに対する言語間スコアのばらつきを著しく低減する。
ユーザ調査では、AIが拡張した5W3Hは、インタラクションラウンドを60%削減し、ユーザの満足度を3.16から4.04に向上させる。
論文 参考訳(メタデータ) (2026-03-31T16:20:28Z) - StructLens: A Structural Lens for Language Models via Maximum Spanning Trees [52.040177523973334]
StructLensは、内部構造が全体構造とどのように関係しているかを明らかにするために設計された分析フレームワークである。
以上の結果から,StructLensは従来のコサイン類似性とは大きく異なる層間類似性パターンを呈することが明らかとなった。
論文 参考訳(メタデータ) (2026-02-10T11:30:32Z) - Multi-Agent Procedural Graph Extraction with Structural and Logical Refinement [66.51979814832332]
モデル式は、専用の構造的および論理的洗練を伴う多ラウンド推論プロセスとして手続きグラフ抽出を定式化する。
実験により、モデルが強いベースラインに対して構造的正当性と論理的整合性の両方において大幅に改善されることが示されている。
論文 参考訳(メタデータ) (2026-01-27T04:00:48Z) - How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction [26.53848099802812]
大言語モデル(LLM)はセマンティック理解に優れるが、スクランブルされた入力から内部構造を再構築する能力は未解明のままである。
中国語,日本語,韓国語で固定された4文字表現を用いた構造復元のための決定論的ベンチマークであるOrderProbeを紹介する。
回復精度を超えるモデルを評価するための診断枠組みを提案し,その内容は意味的忠実度,論理的妥当性,堅牢性,感度,情報密度などである。
論文 参考訳(メタデータ) (2026-01-13T15:03:38Z) - BRIDGE: Building Representations In Domain Guided Program Verification [67.36686119518441]
BRIDGEは、検証をコード、仕様、証明の3つの相互接続ドメインに分解する。
提案手法は, 標準誤差フィードバック法よりも精度と効率を著しく向上することを示す。
論文 参考訳(メタデータ) (2025-11-26T06:39:19Z) - Test Case Generation from Bug Reports via Large Language Models: A Cognitive Layered Evaluation Framework [10.919459368597295]
テストケース生成におけるLarge Language Models(LLM)推論の体系的評価について述べる。
言語的・意味的課題を導入した欠陥4J, GHRB, 変異変種についてStarCoderとGPT-4oを評価した。
論文 参考訳(メタデータ) (2025-10-06T20:47:12Z) - DISPROTBENCH: A Disorder-Aware, Task-Rich Benchmark for Evaluating Protein Structure Prediction in Realistic Biological Contexts [76.59606029593085]
DisProtBenchは、構造障害および複雑な生物学的条件下でタンパク質構造予測モデル(PSPM)を評価するためのベンチマークである。
DisProtBenchはデータの複雑さ、タスクの多様性、解釈可能性という3つの重要な軸にまたがっている。
その結果,機能的予測障害と相関する低信頼領域を有する障害下でのモデルロバスト性に有意な変動が認められた。
論文 参考訳(メタデータ) (2025-06-18T23:58:22Z) - StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs [78.84060166851805]
StructTestは、大規模な言語モデル(LLM)を合成命令に従って構造化出力を生成する能力に基づいて評価する、新しいベンチマークである。
評価はルールベースの評価器を用いて決定的に行われ、新しいタスクやデータセットに容易に拡張できる。
StructTestは、Deepseek-V3/R1やGPT-4oといったトップパフォーマンスモデルでも、依然として難しいままです。
論文 参考訳(メタデータ) (2024-12-23T22:08:40Z) - FoldToken: Learning Protein Language via Vector Quantization and Beyond [56.19308144551836]
タンパク質配列構造を離散シンボルとして表現するために textbfFoldTokenizer を導入する。
学習したシンボルを textbfFoldToken と呼び、FoldToken の配列が新しいタンパク質言語として機能する。
論文 参考訳(メタデータ) (2024-02-04T12:18:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。