FuguReport

Revisiting N2DCG: An Empirically Grounded Reformulation of Carousel Recommendation Evaluation

Authors Jingwei Kang, Santiago de Leon-Martinez, Maarten de Rijke, Harrie Oosterhuis
Affiliations University of Amsterdam / Brno University of Technology / Kempelen Institute of Intelligent Technologies
Categories Method / Ranking Metrics / Normalized discounted cumulative gain reformulation, Evaluation / User Behavior Analysis / Empirical validation with gaze data, Application / Carousel Recommendation / Evaluation of carousel recommendation models
License CC BY 4.0

Abstract Overview

This paper re-examines N2DCG, a two-dimensional adaptation of NDCG for carousel recommendation interfaces, and argues that the original formulation makes assumptions that do not hold in real carousel layouts. The authors identify two core problems: the unconstrained ideal ranking used for normalization can violate category constraints of carousel rows, and the discount function fails to reflect empirically observed browsing behavior after horizontal swipes. They propose a reformulated metric that constrains the ideal layout to valid category-specific rows and introduces mirrored F-pattern discounting with several action-aware variants. The approach is validated using the RecGaze eye-tracking dataset and simulated pairwise layout comparisons.

Novelty

The paper's main novelty is a category-aware and empirically grounded reformulation of N2DCG for carousel interfaces. It restricts the ideal ranking to valid layouts where rows respect category boundaries and introduces mirrored F-pattern discount functions reflecting attention shifts observed in eye-tracking studies.

Results

On held-out eye-tracking data from RecGaze, the proposed row-page discount formulation achieved the highest correlation and lowest mean squared error among all tested discount variants. In simulations across 20,000 trials, the reformulated 2DCG consistently outperformed the original metric in identifying correct pairwise layout preferences, reaching 100.0% accuracy at the largest ground-truth difference threshold (Δ ≥ 0.10) under both binary and graded relevance.

Key Points

  1. The authors show that the original N2DCG normalization can be invalid for carousel interfaces because it ignores the requirement that each row contains items from a single category.
  2. They propose mirrored F-pattern discounting to model the empirically observed right-to-left attention reset on later carousel pages after swiping.
  3. Experiments with real eye-tracking data and simulated layout comparisons indicate that the reformulated metric better reflects user examination behavior and more reliably ranks carousel layouts.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.