Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks
- URL: http://arxiv.org/abs/2510.25623v2
- Date: Thu, 30 Oct 2025 13:49:22 GMT
- Title: Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks
- Authors: Davide Romano, Jonathan Schwarz, Daniele Giofré,
- Abstract summary: Test-time scaling (TTS) techniques can improve the performance of large language models (LLMs) at the expense of additional computation and latency.<n>We present an empirical study of verifier-based TTS methods for legal multiple-choice QA (MCQA) across five benchmarks.
- Score: 2.1049704239329152
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: Test-time scaling (TTS) techniques can improve the performance of large language models (LLMs) at the expense of additional computation and latency. While TTS has proven effective in formal domains such as mathematics and programming, its value in argumentative domains such as law remains underexplored. We present an empirical study of verifier-based TTS methods for legal multiple-choice QA (MCQA) across five benchmarks. Using a family of 7 reward models, we evaluate both outcome-level (Best-of-$N$) and process-level (tree search) verification under realistic low-$N$ budgets. Our analysis systematically investigates how verifier utility is affected by key properties such as domain specialization, model size, and supervision type (process-supervised PRMs vs. outcome-only ORMs), even when applied across different roles.
Related papers
- Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents [20.29427807019999]
Large Language Model (LLM) code agents increasingly resolve repository-level issues by iteratively editing code, invoking tools, and validating candidate patches.<n>In these, agents often write tests on the fly, a paradigm adopted by many high-ranking agents on the SWE-bench leaderboard.<n>This raises the critical question: whether such tests meaningfully improve issue resolution or merely mimic human testing practices while consuming a substantial interaction budget.
arXiv Detail & Related papers (2026-02-08T10:26:31Z) - MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems [59.20800753428596]
We present MAS-ProVe, a systematic empirical study of process verification for multi-agent systems (MAS)<n>Our study spans three verification paradigms (LLM-as-a-Judge, reward models, and process reward models)<n>We find that process-level verification does not consistently improve performance and frequently exhibits high variance.
arXiv Detail & Related papers (2026-02-03T03:30:36Z) - Limits and Gains of Test-Time Scaling in Vision-Language Reasoning [8.76012279865596]
Test-time scaling (TTS) has emerged as a powerful paradigm for improving the reasoning ability of Large Language Models (LLMs) by allocating additional computation at inference.<n>We present a systematic empirical study of inference time reasoning methods applied across both open-source and closed-source Vision-Language Models (VLMs) on different benchmarks.
arXiv Detail & Related papers (2025-12-11T20:48:54Z) - Test-Time Scaling of Reasoning Models for Machine Translation [16.317481079574065]
Test-time scaling (TTS) has enhanced the performance of Reasoning Models (RMs) on various tasks such as math and coding.<n>This paper investigates whether increased inference-time computation improves translation quality.
arXiv Detail & Related papers (2025-10-07T21:15:18Z) - Trust but Verify! A Survey on Verification Design for Test-time Scaling [8.428618801719198]
Test-time scaling (TTS) has emerged as a new frontier for scaling the performance of Large Language Models.<n>Verifiers serve as reward models that help score the candidate outputs from the decoding process.<n>Verifiers could be prompt-based, fine-tuned as a discriminative or generative model.
arXiv Detail & Related papers (2025-08-20T22:27:21Z) - Statistical Inference for Autoencoder-based Anomaly Detection after Representation Learning-based Domain Adaptation [7.10052009802944]
Anomaly detection plays a vital role across a wide range of domains, but its performance might deteriorate when applied to target domains with limited data.<n>We propose STAND-DA -- a novel framework for statistically rigorous Autoencoder-based AD after Representation Learning-based DA.
arXiv Detail & Related papers (2025-08-09T17:24:02Z) - T1: Tool-integrated Self-verification for Test-time Compute Scaling in Small Language Models [9.674458633565111]
We investigate whether small language models (sLMs) can reliably self-verify their outputs under test-time scaling.<n>We propose Tool-integrated self-verification (T1), which delegates-heavy verification steps to external tools, such as a code interpreter.<n>Our theoretical analysis shows that tool integration reduces memorization demands and improves test-time scaling performance.
arXiv Detail & Related papers (2025-04-07T04:01:17Z) - FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models [79.41859481668618]
Large Language Models (LLMs) have significantly advanced the fact-checking studies.<n>Existing automated fact-checking evaluation methods rely on static datasets and classification metrics.<n>We introduce FACT-AUDIT, an agent-driven framework that adaptively and dynamically assesses LLMs' fact-checking capabilities.
arXiv Detail & Related papers (2025-02-25T07:44:22Z) - Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning [59.25951947621526]
We propose an approach which can transform existing coding benchmarks into scoring and ranking datasets to evaluate the effectiveness of synthetic verifiers.<n>We release four new benchmarks (HE-R, HE-R+, MBPP-R, and MBPP-R+), and analyzed synthetic verification methods with standard, reasoning-based, and reward-based LLMs.<n>Our experiments show that reasoning can significantly improve test case generation and that scaling the number of test cases enhances the verification accuracy.
arXiv Detail & Related papers (2025-02-19T15:32:11Z) - The Surprising Effectiveness of Test-Time Training for Few-Shot Learning [59.309477460893916]
Language models (LMs) have shown impressive performance on tasks within their training distribution, but often struggle with structurally novel tasks.<n>We investigate the effectiveness of test-time training (TTT) as a mechanism for improving LMs' reasoning and few-shot learning capabilities.<n>Our findings highlight the limitations of in-context learning for novel tasks and demonstrate the potential of test-time training to enhance language model adaptability.
arXiv Detail & Related papers (2024-11-11T18:59:45Z) - Active Test-Time Adaptation: Theoretical Analyses and An Algorithm [51.84691955495693]
Test-time adaptation (TTA) addresses distribution shifts for streaming test data in unsupervised settings.
We propose the novel problem setting of active test-time adaptation (ATTA) that integrates active learning within the fully TTA setting.
arXiv Detail & Related papers (2024-04-07T22:31:34Z) - BLESS: Benchmarking Large Language Models on Sentence Simplification [55.461555829492866]
We present BLESS, a performance benchmark of the most recent state-of-the-art large language models (LLMs) on the task of text simplification (TS)
We assess a total of 44 models, differing in size, architecture, pre-training methods, and accessibility, on three test sets from different domains (Wikipedia, news, and medical) under a few-shot setting.
Our evaluation indicates that the best LLMs, despite not being trained on TS, perform comparably with state-of-the-art TS baselines.
arXiv Detail & Related papers (2023-10-24T12:18:17Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.