4 papers
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Lai Wei, Chengqi Li, Jiapeng Li +3
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as co…
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
Zhengbo Jiao, Yiming Cheng, Yilei Jiang +15
Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search…
DiagramNet: An End-to-End Recognition Framework and Dataset for Non-Standard System-Level Diagrams
Jincheng Lou, Ruohan Xu, Jiapeng Li +6
System-level diagrams encode the architectural blueprint of chip design, specifying module functions, dataflows, and interface protocols. However, non-standardized symbols and the…
R^3-VQA: "Read the Room" by Video Social Reasoning
Lixing Niu, Jiapeng Li, Xingping Yu +6
"Read the room" is a significant social reasoning capability in human daily life. Humans can infer others' mental states from subtle social cues. Previous social reasoning tasks an…