論文の概要: Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts
- arxiv url: http://arxiv.org/abs/2609.22455v2
- Date: Tue, 22 Sep 2026 02:00:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-23 18:04:03.872206
- Title: Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts
- Title(参考訳): Apollo Restore: 古代ギリシアのテキストの中間修復のために最適化された歴史的ギリシア語のための基盤 LLM
- Abstract要約: 我々は、古代ギリシアのテキストでラグネーを復元するための大きな言語モデルであるアポロ・レストアを提示する。
古代ギリシア語では最初の大規模なデコーダモデルであり、古代地中海語では最初のものである。
- 参考スコア(独自算出の注目度): 4.657316971327409
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore reconstructs missing spans without requiring oracle knowledge of their length. To our knowledge, it is the first large-scale decoder model for historical Greek, and the first for any ancient Mediterranean language. Evaluated as in prior work, on short gaps of up to ten characters, Apollo Restore places the correct restoration among its top twenty candidates for 80.6%/54.6%/61.0% of documentary-papyrus, literary-papyrus, and stone-inscription lacunae, exceeding the strongest published models by $1.6\times$/$2.6\times$/$1.4\times$. Prior evaluation protocols, however, inflate scores through a bias toward trivially short gaps; under a length-balanced metric Apollo Restore's advantage over the strongest published models grows to $2.3\times$/$3.5\times$/$1.6\times$ and degrades gracefully, even given incorrect length hints. In a blind study, 20 expert papyrologists, epigraphists, and philologists strongly preferred Apollo Restore to the strongest baseline and judged its performance at least as good as human restorations in 77% of cases. Apollo Restore also improves the published reading of PHerc. 1667---a papyrus roll carbonised in the eruption of Vesuvius in 79 CE and digitally unrolled and edited after Apollo Restore's training data was compiled. Apollo Restore is an output of the Decoding Antiquity initiative to build specialized LLMs for historical languages and manuscripts, led by the Austrian Academy of Sciences.
- Abstract(参考訳): 本稿は,古代ギリシアの断片的なテキストにおいて,ラグネー語を復元するための24ビリオンの大規模言語モデルであるApollo Restoreを提示する。
ミストラル・スモール(Mistral Small)のアポロ・レストア(Apollo Restore)は、ミストラル・スモール(Mistral Small)のミストラル・スモール(Mistral Small)から、ミストラル・スモール(Mistral Small)のミストラル・スモール(Apollo Restore)を改造した。
我々の知る限り、これは古代ギリシア語にとって初めての大規模なデコーダモデルであり、古代地中海語では初めてのものである。
アポロ・レストア(Apollo Restore)は、前作と同様に、最大10文字の短い間隔で、80.6%/54.6%/61.0%のドキュメンタリー・パピルス、文学・パピルス、および石碑文のラグナの上位20人の候補者に正しい復元を施し、1.6\times$/$2.6\times$/$1.4\times$を上回った。
アポロ・レストアの最も強力なモデルに対する優位性は、2.3\times$/$3.5\times$/$1.6\times$に成長し、誤った長さのヒントを与えられたとしても優雅に劣化する。
盲目的調査では、20人のパピルス学者、エピグラフィスト、文献学者がアポロ・レストアを最強の基準として強く支持し、少なくとも77%の症例でヒトの修復に匹敵する性能を判断した。
Apollo Restore では PHerc の読み込みも改善されている。
1667年 - 紀元前79年にヴェスウィウスの噴火で石炭化したパピルスが、アポロ・レストアのトレーニングデータが編纂された後、デジタルでアンロールされ、編集された。
アポロ・リストア(Apollo Restore)は、オーストリア科学アカデミー(英語版)が主導する歴史言語と写本のための特殊なLSMを構築するためのデコード古代イニシアチブの成果である。
関連論文リスト
- Stoicheia: Character-Level Masked Diffusion for Ancient Greek Textual Restoration, Parsing, and Metrical Scansion [0.0]
ストイチェア (Stoicheia) は古代ギリシアの文字レベルのマスク付き拡散エンコーダである。
オープンでリビジョンされた380万ワードのコーパスで事前トレーニングを行い、11個のチェックポイントをリリースする。
イサカ自身のテスト分割では、同じ凍結サンプルと厳密なスコアで、Stoicheiaは両方の従来の最先端システムと比較して文字エラーを減らす。
論文 参考訳(メタデータ) (2026-08-07T14:07:43Z) - Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains [0.0]
現代のギリシャ語はNVIDIAのネモトロン検索モデルに欠落している。
現代ギリシア語に対するネモトロン検索スタックのエンドツーエンド適応について述べる。
HERAは,検索拡張生成のためのギリシャ初の大規模ベンチマークである。
論文 参考訳(メタデータ) (2026-08-05T17:56:40Z) - EpiAgent: An Agent-Centric System for Ancient Inscription Restoration [60.77886293107454]
EpiAgentは、階層的な計画問題として碑文復元を定式化するエージェント中心のシステムである。
EpiAgentは、既存の方法よりも優れた復元品質とより強力な一般化を実現している。
我々の研究は、専門家レベルのエージェント主導による文化遺産の復元に向けた重要な一歩である。
論文 参考訳(メタデータ) (2026-04-10T14:37:54Z) - The Patrologia Graeca Corpus: OCR, Annotation, and Open Release of Noisy Nineteenth-Century Polytonic Greek Editions [0.0]
パトログア・グラエカ・コーパス(Patrologia Graeca Corpus)は、古代ギリシアの19世紀の版において、最初の大規模なオープンなOCRと言語資源である。
このコレクションは、複雑なバイリンガル(ギリシャ・ラテン語)のレイアウトで印刷されたPatrologia Graeca(PG)の残されている未デジタル化の巻をカバーしており、高度に劣化したポリトニック・ギリシャのタイポグラフィーが特徴である。
We achieve a character error rate (CER) of 1.05% and a word error rate (WER) of 4.69%。
その結果得られたコーパスには、約600万の補修と音声タグ付きトークンが含まれており、フルに整列している。
論文 参考訳(メタデータ) (2026-03-10T10:21:54Z) - Breccia and basalt classification of thin sections of Apollo rocks with deep learning [0.6282171844772422]
月の岩石分類器は、月の岩石サンプルを分析するために宇宙飛行士に必要な情報を提供するツールである。
我々は、アポロ計画からの大量の薄切片画像を活用し、平面偏光(PPL)、横偏光(XPL)、反射光を様々な倍率で捉えた。
微調整されたInception-Resnet-v2ネットワークは、アポロの岩石の薄い断面画像から重要な特徴を効果的に抽出することができる。
論文 参考訳(メタデータ) (2024-10-28T13:45:22Z) - Instruct-Tuning Pretrained Causal Language Models for Ancient Greek Papyrology and Epigraphy [0.0]
本稿は、古代ギリシアの碑文やドキュメンタリーパピルスの欠落した文字を復元するために、事前訓練された因果関係言語モデルを微調整する実験について述べる。
最新技術モデル (Ithaca) と比較すると、テキスト復元に優れた命令調整モデルである。
以上の結果から,修正および予想のための命令テンプレートを用いた事前学習型因果言語モデルの微調整が有望であることが示唆された。
論文 参考訳(メタデータ) (2024-09-20T19:49:45Z) - Puzzle Pieces Picker: Deciphering Ancient Chinese Characters with Radical Reconstruction [73.26364649572237]
Oracle Bone Inscriptionsは、世界で最も古い書式である。
多くのOracle Bone Inscriptions (OBI) は未解読のままであり、今日の古生物学におけるグローバルな課題の1つとなっている。
本稿では, 急進的再構成によってこれらの謎的文字を解読する新しい手法, Puzzle Pieces Picker (P$3$) を提案する。
論文 参考訳(メタデータ) (2024-06-05T07:34:39Z) - An open dataset for oracle bone script recognition and decipherment [66.35957530824872]
古代中国最古の書体の一つ、Oracleの骨書は、3000年前にさかのぼる上海王朝の人文・地理を研究する学者にとって、貴重な研究資料を提示している。
時間の経過はそれらの意味の多くを曖昧にしており、これらの古代のテキストを解読する上で重要な課題が提示されている。
人工知能(AI)の出現により、Oracle Bone Characters(OBC)の解読を支援するAIが実現可能な選択肢となっている。
このデータセットは1,588個の解読文字の77,064個の画像と9,411個の未解読文字の62,989個の画像を含む。
論文 参考訳(メタデータ) (2024-01-27T09:54:16Z) - An open dataset for the evolution of oracle bone characters: EVOBC [72.91231825135665]
現存する最古の漢字は、他の東アジアの言語と密接に関連しているオラクルの骨碑文に由来する。
本研究では,6つの歴史的段階にまたがる権威あるテキストやウェブサイトから,古代の文字を体系的に収集した。
我々は13,714の異なる文字カテゴリを表す229,170の画像からなる広範囲なデータセットを構築した。
論文 参考訳(メタデータ) (2024-01-23T03:30:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。