FuguReport

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

Authors Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt
Affiliations Adobe / Independent
Categories Evaluation / Model Evaluation / LLM behavior under input-output compression, Method / Evaluation Protocol / Two-channel evaluation methodology, Task / Compression Effects / Impact of linguistic input/output reduction
License CC BY 4.0

Abstract Overview

This paper introduces CAVEWOMAN, a two-channel evaluation protocol for studying large language models under linguistic compression of either the input prompt or the output response. The framework evaluates generations on three axes: task accuracy, realized per-item cost, and agreement with the model’s own unconstrained reference output, using bidirectional NLI and complementary semantic measures. The authors test eight models on five benchmarks across five reduction levels, applying the same reduction family to both channels for direct comparison. The evidence shows that compressing prompts and compressing responses have qualitatively different effects on cost, accuracy, and semantic agreement.

Novelty

The paper’s main novelty is treating linguistic compression as a two-channel problem and evaluating input compression and output compression on the same items with the same reduction levels. It also adds reference-text agreement against the model’s unconstrained output, alongside realized cost and accuracy, to reveal divergences that accuracy-only evaluations miss.

Results

Output compression reduced realized cost on most API models and all four open-weight models under the pricing assumptions studied, while input compression generally increased net cost because models compensated with longer responses and accuracy deteriorated at stronger reductions. Under L1 output compression, many non-reasoning-model outputs remained correct while no longer entailing the model’s own unconstrained reference, with a pooled correct-but-divergent rate of 51.9% across six non-reasoning models. Robustness to compression varied substantially by model and did not track unconstrained accuracy or parameter count.

Key Points

  1. The study finds a strong asymmetry between channels: output compression can lower realized inference cost, whereas input compression often raises cost through compensatory output expansion.
  2. Accuracy alone is insufficient for evaluating compression, because outputs can stay correct while diverging semantically from the model’s own unconstrained reference text.
  3. Model rankings change under compression constraints, so systems should be evaluated at the constraint level expected in deployment rather than only in unconstrained settings.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.