FuguReport

Investigating Knowledge Transfer Across Interactive Dialogue Games

Authors Filippo Momentè, Mir Nafis Sharear Shopnil, Andrea de Varda, Pavel Merinov, Raffaella Bernardi, Oswald Lanz, Alessandro Suglia, Alessandro Torcinovich
Affiliations The University of Edinburgh / Free University of Bozen Bolzano / Massachusetts Institute of Technology / University of Trento / Technovative Solutions Ltd
Categories Application / Dialogue Systems / Interactive dialogue game tasks, Method / Knowledge Transfer / Cross-game knowledge transfer study, Evaluation / Model Evaluation / Transferability evaluation on game suite
License CC BY 4.0

Abstract Overview

This paper studies how knowledge transfers across interactive dialogue games by fine-tuning a common LLM on tasks from the clembench suite and evaluating cross-task performance. The authors adapt the Taskonomy framework to build directed transfer taxonomies from a transfer matrix derived from specialist models, complementing this with a weight-space analysis based on task vectors. Their dataset is constructed from 17 textual clembench games, filtered and balanced down to 15 role-specific tasks across 9 games. Through these analyses, the paper investigates which games serve as effective transfer sources and whether parameter-space similarity can explain cross-game transfer behavior.

Novelty

The study applies a structured Taskonomy-style transfer optimization framework systematically across a broad suite of interactive dialogue games rather than static benchmarks or isolated task pairs. It also combines behavioral transfer taxonomies with task-vector subspace analyses to assess whether parameter-space adaptations align with empirical transferability.

Results

Empirical results show that visuospatial dialogue games transfer more effectively than verbal games across training budgets, with tasks like adventuregame acting as strong sources that sometimes improve target performance more than self-training. The discovered transfer taxonomies significantly outperform randomly sampled transfer baselines. Conversely, weight-space task vector similarity reflects game and role identity rather than transferability, showing that symmetric parameter metrics fail to predict directional transfer patterns.

Key Points

  1. The study constructs transferability taxonomies across 15 role-specific tasks from 9 clembench games by optimizing source task subsets using quality-score gains over a baseline LLM.
  2. Visuospatial tasks exhibit the strongest cross-game transferability across budget regimes and provide versatile capabilities that generalize well to verbal games.
  3. Task-vector representations align strongly between roles within the same game, but symmetric parameter-space similarity does not reliably predict asymmetric task transferability.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.