論文の概要: TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
- arxiv url: http://arxiv.org/abs/2607.08803v1
- Date: Thu, 09 Jul 2026 06:56:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-13 14:47:12.719185
- Title: TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
- Title(参考訳): TheBioCollection:Unified Pre-Training Scale LLM Corpus for Biology
- Authors: Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung,
- Abstract要約: 生物学のための大規模言語モデル(BioLM)への推進は、生物学を真に理解したモデルを支援するコーパスのトレーニングの必要性を生み出した。
52.6Bの事前学習型コーパスであるTheBioCollectionは、異なる資源を小さな分子、タンパク質、ゲノム配列、細胞、経路にまたがる統一的で訓練可能な形式に変換する。
TheBioCollection-Evalは、分子、タンパク質、ゲノム、細胞、およびドメイン間における認識、生成、予測をマッチングしたスイートである。
- 参考スコア(独自算出の注目度): 27.715802943211504
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.
- Abstract(参考訳): 生物学のための大規模言語モデル(BioLM)への推進は、生物学を真に理解したモデルを支援するコーパスのトレーニングの必要性を生み出した。
しかし、分子データベース、タンパク質リポジトリ、ゲノムアノテーション、単細胞アトラス、経路データベースといった既存の生物資源は、異質な形式に分散しており、言語モデルトレーニングのための凝集性コーパスに組織化されていない。
52.6BのプレトレーニングスケールコーパスであるTheBioCollectionは、これらの異なるリソースを、小さな分子、タンパク質、ゲノム配列、細胞、経路にまたがる統一的で訓練可能な形式に変換する。
既存のデータの統合以外にも、TheBioCollectionは各レコードにツール計算された生物学的プロパティを付加し、現在のコーパスがほとんどカバーしていない機能のための新しい命令タスクを導入している。
TheBioCollection-Evalは、分子、タンパク質、ゲノム、細胞、およびドメイン間における認識、生成、予測をマッチングしたスイートである。
ベースとなるGravity-16B-A3Bアーキテクチャの修正を保ちながら、TheBioCollectionのトレーニングは、TheBioCollection-Evalの総スコアを2倍以上にし、すべてのドメインでゲインを獲得し、一般的な言語能力はほとんど残っていない。
関連論文リスト
- BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language [47.151057913560074]
BioMatrixは、分子とタンパク質の両方の配列、構造、自然言語を統合するマルチモーダル基盤モデルである。
Qwen3言語モデルに基づいて構築されたBioMatrixは、一般的なテキストとドメイン固有のテキストにまたがる304億のトークンを継続的に事前訓練している。
BioMatrixは80タスク中77タスクで最先端または競合的なパフォーマンスを達成する。
論文 参考訳(メタデータ) (2026-06-20T16:38:59Z) - Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models [55.74944165932666]
本稿では,生物配列の大規模学習データセットであるBiology-Instructionsを紹介する。
このデータセットは、大きな言語モデル(LLM)と複雑な生物学的シーケンス関連タスクをブリッジし、その汎用性と推論を強化する。
また,マルチオミクスタスクにおける現状のLLMの,専門訓練なしでの大幅な制限を強調した。
論文 参考訳(メタデータ) (2024-12-26T12:12:23Z) - Large-Scale Multi-omic Biosequence Transformers for Modeling Protein-Nucleic Acid Interactions [4.36852565205713]
OmniBioTEは,250億以上のタンパク質と核酸を混合したトークンをトレーニングした,オープンソースのマルチオミックモデルである。
我々は,OmbiBioTEが与えられた核酸とタンパク質の結合相互作用のギブス自由エネルギーの変化を予測できることを示す。
論文 参考訳(メタデータ) (2024-08-29T03:56:40Z) - BioT5: Enriching Cross-modal Integration in Biology with Chemical
Knowledge and Natural Language Associations [54.97423244799579]
$mathbfBioT5$は、化学知識と自然言語の関連性によって生物学のクロスモーダルな統合を強化する事前学習フレームワークである。
$mathbfBioT5$は構造化知識と非構造化知識を区別し、より効果的な情報利用につながる。
論文 参考訳(メタデータ) (2023-10-11T07:57:08Z) - BioGPT: Generative Pre-trained Transformer for Biomedical Text
Generation and Mining [140.61707108174247]
本稿では,大規模生物医学文献に基づいて事前学習したドメイン固有生成型トランスフォーマー言語モデルであるBioGPTを提案する。
BC5CDRでは44.98%、38.42%、40.76%のF1スコア、KD-DTIとDDIの関係抽出タスクでは78.2%、PubMedQAでは78.2%の精度が得られた。
論文 参考訳(メタデータ) (2022-10-19T07:17:39Z) - BioALBERT: A Simple and Effective Pre-trained Language Model for
Biomedical Named Entity Recognition [9.05154470433578]
既存のBioNERアプローチはこれらの問題を無視し、最先端(SOTA)モデルを直接採用することが多い。
本稿では,大規模バイオメディカルコーパスを用いた効果的なドメイン固有言語モデルであるALBERTを提案する。
論文 参考訳(メタデータ) (2020-09-19T12:58:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。