From the 1 of 30 linked papers with an AI index.
30 papers
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu +6
The paper proposes Ego Scene Augmentation (ESA), a framework that uses an Ego-element Graph to improve the spatial perception of multimodal large language models for egocentric vis…
Kwai Summary Attention Technical Report
Chenglong Chu, Guorui Zhou, Guowang Zhang +35
Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agen…
OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning
Haocong He, Chenfei Liao, Zichen Wen +13
Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of p…
OneReason Technical Report
OneRec Team, Biao Yang, Boyang Ding +81
Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. Howev…
Panoramic Affordance Prediction
Zixin Zhang, Chenfei Liao, Hongfei Zhang +10
Affordance prediction serves as a critical bridge between perception and action in embodied AI. However, existing research is confined to pinhole camera models, which suffer from n…
DVD: Deterministic Video Depth Estimation with Generative Priors
Hongfei Zhang, Harold Haodong Chen, Chenfei Liao +12
Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand…