activity
20242026
collaborators

9 papers

cs.CV2026

SAW-Bench: Learning Situated Awareness in the Real World

Chuhan Li, Rilyn Han, Joy Hsu +5

A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. Howeve…

cs.CV2026

Pursuing Minimal Sufficiency in Spatial Reasoning

Yejie Guo, Yunzhong Hou, Wufei Ma +2

Spatial reasoning, the ability to ground language in 3D understanding, remains a persistent challenge for Vision-Language Models (VLMs). We identify two fundamental bottlenecks: in…

cs.CV2025

Edit3r: Instant 3D Scene Editing from Sparse Unposed Images

Jiageng Liu, Weijie Lyu, Xueting Li +2

We present Edit3r, a feed-forward framework that reconstructs and edits 3D scenes in a single pass from unposed, view-inconsistent, instruction-edited images. Unlike prior methods…

cs.CL2025

TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Frank F. Xu, Yufan Song, Boxuan Li +18

We interact with computers on an everyday basis, be it in everyday life or work, and many aspects of work can be done entirely with access to a computer and the Internet. At the sa…

cs.CV2025

CoCo4D: Comprehensive and Complex 4D Scene Generation

Junwei Zhou, Xueting Li, Lu Qi +1

Existing 4D synthesis methods primarily focus on object-level generation or dynamic scene synthesis with limited novel views, restricting their ability to generate multi-view consi…

cs.CV2025

HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis

Xiaoyuan Wang, Yizhou Zhao, Botao Ye +6

We propose HoliGS, a novel deformable Gaussian splatting framework that addresses embodied view synthesis from long monocular RGB videos. Unlike prior 4D Gaussian splatting and dyn…