Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
2609.39183

Authors

Arjun Reddy Akula,Angela Yao,Pengzhan Sun,Shiu-hong Kao,Shijie Li

Abstract

This paper studies thinking--answer consistency in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box.

We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from "thinking drift", where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose Rita (ReInforcing Thinking--Answer consistency) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a thinking reward and a consistency reward. It also adopts a difficulty-aware data filtering strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.

Resources

Ray graphicRay graphicRay graphicRay graphic

Stay in the loop

Every AI paper that matters, free in your inbox daily.

Details

  • takara.ai
  • Custom AI and machine learning from the Frontier Research Team.
  • © 2026 takara.ai Ltd
  • Content is sourced from third-party publications.
Ray graphicRay graphicRay graphic