FuguReport

DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue

Authors Qi Zhang, Heajun An, Prakriti Dumaru, Sang Won Lee, Lifu Huang, Pamela J. Wisniewski, Jin-Hee Cho
Affiliations Virginia Tech / University of California, Davis / International Computer Science Institute
Categories Method / Reinforcement Learning / Stage-aware dialogue modeling, Application / Mental Health / Cognitive behavioral therapy dialogue, Evaluation / Dialogue Efficiency / Engagement and goal completion balance
License CC BY 4.0

Abstract Overview

DeepSAGE is a hybrid LLM–DRL framework for conducting the first CBT counseling session as an explicit sequence of eleven stages, each with a therapeutic objective and automated stage-completion criterion. An external controller decides when a stage is complete using semantic similarity and NLI-based goal satisfaction, while a reinforcement learning policy selects therapeutic intentions that guide the LLM's responses. The system is evaluated against six retrieval-, prompting-, stage-, and policy-based baselines using simulated clients representing anxiety and major depressive disorder conditions. Across these simulations, the paper reports stronger engagement and a better balance between stage-goal completion and dialogue efficiency for DeepSAGE, while emphasizing that the evidence reflects comparative dialogue-control performance rather than clinical effectiveness.

Novelty

The paper's main novelty is framing a full first-session CBT dialogue as an eleven-stage control problem with explicit therapeutic goals, automated transition criteria, and DRL-based therapeutic-intention selection. This distinguishes it from prior LLM counseling work focused on single turns, prompt-only CBT guidance, or unstructured supportive conversation.

Results

In simulated anxiety and depression sessions, DeepSAGE achieved the highest user utterance length and self-disclosure among the compared systems, and it produced the highest anxiety-condition emotional intensity drop while remaining competitive on depression-condition distress reduction. Among stage-structured systems, it achieved the highest session success rates with fewer counselor utterances than Stage-Prompt and DeepSAGE_R, demonstrating a stronger completion-efficiency tradeoff. Expert review further indicated that the generated dialogues exhibited broadly plausible emotional trajectories and recognizable CBT processes.

Key Points

  1. DeepSAGE combines an external stage controller with a PPO-trained policy that selects therapeutic intentions for LLM response generation within an eleven-stage CBT session structure.
  2. Stage transitions are determined through explicit goal-success scoring combining semantic similarity and NLI entailment, rather than relying solely on prompts or fixed dialogue flows.
  3. The demonstrated advantages appear in simulated engagement, self-disclosure, and protocol-completion efficiency, while the authors explicitly disclaim clinical effectiveness and highlight safety limitations for high-risk cases.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.