11 papers
OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning
Maonan Wang, Zhengyan Huang, Kemou Jiang +13
Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. Howev…
InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain
Tiancheng Han, Yong Li, Wuzhou Yu +2
Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts. Chunk-wise memory agents address this issue by sequentially reading docume…
SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction
Zhixiong Zhang, Yizhuo Li, Shuangrui Ding +6
Referring segmentation grounds natural-language queries to pixel-level masks, but extending it to complex scenarios with multiple instances, cross-category groups, or open-ended ta…
Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation
Zhiyao Cui, Chenxu Wang, Shuyue Hu +4
Producing presentation slides automatically entails coordinating narrative structure with page-level graphic design under strict spatial constraints. For such structured multimodal…
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
Wenhan Wang, Zhixiang Zhou, Zhongtian Ma +8
The connoisseurship of antique Chinese porcelain demands extensive historical expertise, material understanding, and aesthetic sensitivity, making it difficult for non-specialists…
Intern-S1: A Scientific Multimodal Foundation Model
Lei Bai, Zhongrui Cai, Yuhang Cao +173
In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that…