133 citations · 210 across the 40 of their papers we have counts for
59 papers
TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta +5
Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We i…
GeomVLA: Unifying Scene, Motion, and Action in 3D
Ziyin Xiong, Nikolaos Gkanatsios, Nikos Gkanatsios +2
We present GeomVLA, a Vision-Language-Action (VLA) model that unifies perception, latent scene motion prediction, and action generation within a shared robot-centric 3D coordinate…
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
Lucy Lin, Ayush Jain, Yifan Liu +1
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization an…
Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
Mihir Prabhudesai, Aryan Satpathy, Yangmin Li +6
We have witnessed remarkable advances in LLM reasoning capabilities with the advent of DeepSeek-R1. However, much of this progress has been fueled by the abundance of internet ques…
Iterative Refinement Improves Compositional Image Generation
Shantanu Jaiswal, Mihir Prabhudesai, Nikash Bhardwaj +5
Text-to-image (T2I) models have achieved remarkable progress, yet they continue to struggle with complex prompts that require simultaneously handling multiple objects, relations, a…
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
Yifan Liu, Fangneng Zhan, Wanhua Li +3
Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on top of 2D visual backbones and depend…