2 papers
cs.SD2026
Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping
Yining Wang
Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's out…
cs.CV2026
LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents
Zijian Wang, Junnan Zhu, Rongzhen Li +7
Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video tool-use agents address this ch…