Fugu-MT 論文翻訳(概要): Benchmarking Large Language Models for Geolocating Colonial Virginia Land Grants

論文の概要: Benchmarking Large Language Models for Geolocating Colonial Virginia Land Grants

arxiv url: http://arxiv.org/abs/2508.08266v1
Date: Sun, 27 Jul 2025 21:49:58 GMT
ステータス: 翻訳完了
システム内更新日: 2025-08-17 22:58:06.150092
Title: Benchmarking Large Language Models for Geolocating Colonial Virginia Land Grants
Title（参考訳）: コロニアルバージニア土地の地理化のための大規模言語モデルのベンチマーク
Authors: Ryan Mioduski,
Abstract要約: バージニアの17世紀から18世紀の土地特許は、主に物語のメッツ・アンド・バウンドの記述として残っている。本研究では、これらの散文を地理的に正確な緯度・経度座標に変換する際に、現在世代の大言語モデル(LLM)を体系的に評価する。
参考スコア（独自算出の注目度）: 0.0
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Virginia's seventeenth- and eighteenth-century land patents survive primarily as narrative metes-and-bounds descriptions, limiting spatial analysis. This study systematically evaluates current-generation large language models (LLMs) in converting these prose abstracts into geographically accurate latitude/longitude coordinates within a focused evaluation context. A digitized corpus of 5,471 Virginia patent abstracts (1695-1732) is released, with 43 rigorously verified test cases serving as an initial, geographically focused benchmark. Six OpenAI models across three architectures (o-series, GPT-4-class, and GPT-3.5) were tested under two paradigms: direct-to-coordinate and tool-augmented chain-of-thought invoking external geocoding APIs. Results were compared with a GIS-analyst baseline, the Stanford NER geoparser, Mordecai-3, and a county-centroid heuristic. The top single-call model, o3-2025-04-16, achieved a mean error of 23 km (median 14 km), outperforming the median LLM (37.4 km) by 37.5%, the weakest LLM (50.3 km) by 53.5%, and external baselines by 67% (GIS analyst) and 70% (Stanford NER). A five-call ensemble further reduced errors to 19 km (median 12 km) at minimal additional cost (approx. USD 0.20 per grant), outperforming the median LLM by 48.6%. A patentee-name-redaction ablation increased error by about 9%, indicating reliance on textual landmark and adjacency descriptions rather than memorization. The cost-efficient gpt-4o-2024-08-06 model maintained a 28 km mean error at USD 1.09 per 1,000 grants, establishing a strong cost-accuracy benchmark; external geocoding tools offered no measurable benefit in this evaluation. These findings demonstrate the potential of LLMs for scalable, accurate, and cost-effective historical georeferencing.
Abstract（参考訳）: バージニア州の17世紀から18世紀にかけての土地特許は、主に物語のミート・アンド・バウンドの説明として存続し、空間分析を制限している。本研究では,これらの散文を地理的に正確な緯度・経度座標に変換する際に,現在の大言語モデル (LLM) を集中評価文脈内で体系的に評価する。 5,471のバージニア特許抽象化のデジタルコーパス(1695-1732)がリリースされ、43の厳密に検証されたテストケースが初期的、地理的に焦点を絞ったベンチマークとして機能している。 3つのアーキテクチャ(oシリーズ、GPT-4クラス、GPT-3.5)にわたる6つのOpenAIモデルを、2つのパラダイムでテストした。その結果、GIS分析系ベースライン、スタンフォードNERジオパーサー、モルデカイ-3、および郡中心のヒューリスティックと比較された。最上位のシングルコールモデルであるo3-2025-04-16は平均誤差23 km (median 14 km)、中央値LLM (37.4 km) の37.5%、最も弱いLLM (50.3 km) の53.5%、外部ベースラインの67% (GISアナリスト) と70% (Stanford NER) を上回った。 5発のアンサンブルにより、最小追加コストで19 km (median 12 km) の誤差が減少し、中央値のLLMを48.6%上回った。特許出願人のリアクション・アブレーションはエラーを約9%増加させ、暗記よりもテキストのランドマークと隣接性の記述に依存することを示した。コスト効率のよいgpt-4o-2024-08-06モデルでは、1,000の補助金に対して平均28kmの誤差がUSD 1.09で維持され、強力なコスト精度のベンチマークが確立された。これらの結果は,拡張性,正確性,費用対効果を有する歴史的ジオレファレンスにおけるLCMsの可能性を示している。

論文の概要: Benchmarking Large Language Models for Geolocating Colonial Virginia Land Grants

関連論文リスト