6 papers
Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Dazhao Du, Shiyan Du, Jian Liu +8
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs…
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
Dazhao Du, Jian Liu, Jialong Qin +7
Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather…
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
Dazhao Du, Liao Duan, Jian Liu +5
Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs)…
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
Jiawei Ren, Yan Zhuang, Xiaokang Ye +20
While LLM/VLM-powered AI agents have advanced rapidly in math, coding, and computer use, their applications in complex physical and social environments remain challenging. Building…
MotionScript: Natural Language Descriptions for Expressive 3D Human Motions
Payam Jome Yazdian, Rachel Lagasse, Hamid Mohammadi +3
We introduce MotionScript, a novel framework for generating highly detailed, natural language descriptions of 3D human motions. Unlike existing motion datasets that rely on broad a…
Evolutionary and Coevolutionary Multi-Agent Design Choices and Dynamics
Erik Hemberg, Eric Liu, Lucille Fuller +2
We investigate two representation alternatives for the controllers of teams of cyber agents. We combine these controller representations with different evolutionary algorithms, one…