4 papers
LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search
Yulun Zhang, Zixu Li, Zhiwei Chen +6
Traditional Text-based Person Search (TPS) is typically limited to matching static appearance attributes, severely neglecting dynamic action information. The Text-based Person Anom…
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
Yuqian Fu, Tianwen Qian, Yanjun Li +30
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…
OmniEgo-R: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026
Zixu Li, Zhiwei Chen, Zhiheng Fu +4
The 1st Cross-Domain EgoCross Challenge at EgoVis, CVPR 2026 evaluates whether multimodal large language models can reason over egocentric videos across surgery, industry, extreme…
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
Daxiang Dong, Mingming Zheng, Dong Xu +17
We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It…