論文の概要: Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls
- arxiv url: http://arxiv.org/abs/2607.13636v1
- Date: Wed, 15 Jul 2026 09:29:24 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-16 16:39:12.719791
- Title: Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls
- Title(参考訳): クロールが見るもの: 縦断ウェブクロールにおける発見曲線、コアパーシステンス、シェルダイナミクス
- Authors: Michael Paris, Hande Celikkanat, Luca Foppiano,
- Abstract要約: 長手ウェブクロールは、進化するURL人口の部分的なサンプルのシーケンスである。
クロールの単純な例文モデルの下で、各ラウンドはURLのごく一部をサンプリングし、断片を置換する。
この分析を、emphdiscovery curve $U(s, T)$, 累積URLフットプリントで拡張します。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: A longitudinal web crawl is a sequence of partial samples of an evolving URL population. Pairwise containment between two crawls is the standard probe; under a simple \emph{urn} model of the crawl -- each round samples a fraction of the URLs and replaces a fraction -- it recovers two interpretable rates, per-round survival $α$ and coverage $c$, but treats the population as uniform and consumes one pair at a time. In this work, we define a formal language for talking about a crawl. We extend this analysis with the \emph{discovery curve} $U(s, T)$, the cumulative URL footprint over a sliding window of $T$ crawls starting at $s$, which under the same urn model is also a closed-form function of $(α, c)$. Containment and the discovery curve are then two projections of one process: independent fits agree on $(α, c)$ when the urn is homogeneous, so any disagreement is itself a measurement. Applied to Common Crawl (2020--2025, domain granularity) and to the German Academic Web (GAW, URL granularity), the two projections disagree on both archives, and a two-component urn with a persistent core fraction $κ$ alongside shell parameters $(α_\partial, c_\partial)$ reconciles the disagreement. A residual on $c_\partial$ remains, signaling that the shell itself is not homogeneous; $κ$ is recorded as the scalar entry point to a rank-resolved generalization, which is left to follow-up work. \keywords{web archive \and crawl coverage \and discovery curve \and urn model \and two-component model \and URL lifetime}
- Abstract(参考訳): 長手ウェブクロールは、進化するURLの集団の部分的なサンプルのシーケンスである。
クロールの単純な 'emph{urn} モデル -- それぞれのラウンドでURLのごく一部をサンプリングし、その分を置き換える -- では、2つの解釈可能なレート、ラウンドサバイバルあたりの$α$とカバレッジ$c$を回復するが、人口を均一に扱い、一度に1つのペアを消費する。
本研究では,クローリングについて話すための形式言語を定義する。
この解析は、$U(s, T)$, 累積URLフットプリントを$T$のスライディングウィンドウ上で$s$から開始し、同じurnモデルの下でも$(α, c)$のクローズドフォーム関数である。
内包と発見曲線は1つの過程の2つの射影である: 独立適合は urn が均質であるときに$(α, c)$ に一致するので、いかなる不一致もそれ自体は測度である。
Common Crawl (2020-2025, domain Granity), and the German Academic Web (GAW, URL Granity), and the two-component urn with a persistent core fraction $κ$ along shell parameters $(α_\partial, c_\partial)$ reconciles the disagreement.
c_\partial$ の残余は、シェル自体が均質でないことを示唆するものであり、$κ$ は階数分解された一般化へのスカラーエントリーポイントとして記録され、後続の作業に残されている。
\keywords{web Archive \and crawl coverage \and discovery curve \and urn model \and two-component model \and URL lifetime
関連論文リスト
- Bond-dimension scaling of a local-refinement advantage over hyperoptimized tensor-network contraction on Sycamore like topologies [0.0]
我々は,コテングラテンソル-ネットワーク収縮パイプラインにおける局所再分極の欠如を同定した。
我々は、その影響がシカモア型トポロジーの直交性グラフ上の結合次元とともに単調に増加することを示す。
論文 参考訳(メタデータ) (2026-04-28T11:59:31Z) - Contextuality from the Projector Overlap Matrix [0.6372261626436676]
我々は、Kochen-Speckerコンテキストのいくつかの既知の指標を単一の射影幾何学的フレームワークに配置する。
Tcal_ij = d-1tr[(hat P_i hat Q_j)2]$, $hat P_i$ と $hat Q_j$ は、2つの互換性のある可観測対の合同固有空間射影である。
KCBSペンタゴンのスピン-$1$実現において、各文脈における共有$m_s=0$固有状態が$c_MU = であることを示す。
論文 参考訳(メタデータ) (2026-04-26T21:56:11Z) - Canonicalizing Multimodal Contrastive Representation Learning [76.15228959754727]
ここでは,CLIP,SigLIP,FLAVAなどのモデルファミリにおいて,埋め込み空間間の幾何学的関係が存在することを示す。
この発見は、後方互換性のあるモデルアップグレードを可能にし、コストのかかる再埋め込みを回避し、学習された表現のプライバシに影響を及ぼす。
論文 参考訳(メタデータ) (2026-02-19T18:09:36Z) - Separating Oblivious and Adaptive Models of Variable Selection [13.61388474201292]
最適$ell_infty$誤差は、ほぼ直線時間で$gtrsim k2$サンプルで達成可能であることを示す。
本研究は,一括適応型 $ モデルの予備試験で結論付ける。
論文 参考訳(メタデータ) (2026-02-18T16:10:35Z) - On approximating the $f$-divergence between two Ising models [3.577310844634503]
2つのIsingモデル間の$f$-divergenceを近似する問題について検討する。
我々のアルゴリズムはKullback-Leiblerの発散など他の$f$-divergencesにも拡張できる。
論文 参考訳(メタデータ) (2025-09-05T11:25:22Z) - Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent Analysis [54.57279006229212]
情報指数は、オンライン勾配降下のサンプルの複雑さを予測する上で重要な役割を担っている。
本研究では,2次項と高次項の両方を考慮することで,まず2次項を用いて関連する空間を学習できることを示す。
オンラインSGDの全体サンプルと複雑さは$tildeO(d PL-1 )$である。
論文 参考訳(メタデータ) (2024-10-13T00:14:08Z) - Online Learning with Adversaries: A Differential-Inclusion Analysis [52.43460995467893]
我々は,完全に非同期なオンラインフェデレート学習のための観察行列ベースのフレームワークを提案する。
我々の主な結果は、提案アルゴリズムがほぼ確実に所望の平均$mu.$に収束することである。
新たな差分包摂型2時間スケール解析を用いて,この収束を導出する。
論文 参考訳(メタデータ) (2023-04-04T04:32:29Z) - Detection-Recovery Gap for Planted Dense Cycles [72.4451045270967]
期待帯域幅$n tau$とエッジ密度$p$をエルドホス=R'enyiグラフ$G(n,q)$に植え込むモデルを考える。
低次アルゴリズムのクラスにおいて、関連する検出および回復問題に対する計算しきい値を特徴付ける。
論文 参考訳(メタデータ) (2023-02-13T22:51:07Z) - Reward-Mixing MDPs with a Few Latent Contexts are Learnable [75.17357040707347]
報酬混合マルコフ決定過程(RMMDP)におけるエピソード強化学習の検討
我々のゴールは、そのようなモデルにおける時間段階の累積報酬をほぼ最大化する、ほぼ最適に近いポリシーを学ぶことである。
論文 参考訳(メタデータ) (2022-10-05T22:52:00Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。