2 papers
cs.CV2026
Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling
Yuqi Hou, Zhuo Chen, Han Hu +3
Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals inde…
cs.AI2026
SceneActBench: Can Agents Act on the 3D Scenes They See?
Yifei Zhao, Xiangxin Zhou, Wenhao Yang +11
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operat…