2 papers
cs.CV2026
STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning
Xiaowen Zhang, Zhi Gao, Licheng Jiao +2
In vision-language models (VLMs), misalignment between textual descriptions and visual coordinates often induces hallucinations. This issue becomes particularly severe in dense pre…
cs.LG2025
Fantastic Multi-Task Gradient Updates and How to Find Them In a Cone
Negar Hassanpour, Muhammad Kamran Janjua, Kunlin Zhang +4
Balancing competing objectives remains a fundamental challenge in multi-task learning (MTL), primarily due to conflicting gradients across individual tasks. A common solution relie…