Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
Junbin Xiao, Shenglang Zhang, Pengxiang Zhu +1
We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the came…
cs.CV2025
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
Wendong Bu, Kaihang Pan, Yuze Lin +6
Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods a…
cs.CV2025
Towards Physically Executable 3D Gaussian for Embodied Navigation
Bingchen Miao, Rong Wei, Zhiqi Ge +8
3D Gaussian Splatting (3DGS), a 3D representation method with photorealistic real-time rendering capabilities, is regarded as an effective tool for narrowing the sim-to-real gap. H…