4 papers
ViThinker: Active Vision-Language Reasoning via Dynamic Perceptual Querying
Weihang You, Qingchan Zhu, David Liu +3
Chain-of-Thought (CoT) reasoning excels in language models but struggles in vision-language models due to premature visual-to-text conversion that discards continuous information s…
Achieving Fine-grained Cross-modal Understanding through Brain-inspired Hierarchical Representation Learning
Weihang You, Hanqi Jiang, Yi Pan +3
Understanding neural responses to visual stimuli remains challenging due to the inherent complexity of brain representations and the modality gap between neural data and visual inp…
FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI
Yuhang Peng, Yizhou Pan, Xinning He +6
As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex,…
Benchmarking Table Comprehension In The Wild
Yikang Pan, Yi Zhu, Rand Xie +1
Large Language Models (LLMs), while being increasingly dominant on a myriad of knowledge-intensive activities, have only had limited success understanding lengthy table-text mixtur…