46 citations · 136 across the 137 of their papers we have counts for
100 papers · 1 filter
Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions
Zhongyao Wang, Wanli Ouyang, Taoyong Cui +1
Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can rewar…
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Zhen Fang, Yu Zeng, Wenxuan Huang +17
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding couple…
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
Xiaopei Wu, Chenshu Hou, Liang Peng +9
3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene. Despite the impressive results achieved by previous m…
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
Haitao Wu, Qirui Zhang, Zhouheng Yao +8
Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approa…
Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation
Xiaopei Wu, Yuenan Hou, Junkai Xu +7
Multi-modal fusion and multi-model ensembling are prevalent in enhancing the performance of 3D semantic segmentation. Despite the impressive performance, these methods either rely…
LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory
Jianbao Cao, Zhangrui Zhao, Bohan Feng +15
Automated laboratories hold the promise of accelerating scientific discovery, yet their deployment is bottlenecked by the difficulty of designing safe and executable environments.…