9 papers
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
Dongping Chen, Xuanao Huang, Zhihan Hu +3
As multimodal LLMs increasingly target video and audio, it is often assumed that such tasks require native omnimodal models. We show that this is not always the case: coding agents…
On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
Yue Huang, Chujie Gao, Siyuan Wu +63
Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions.…
Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
Chenlong Wang, Yuhang Chen, Zhihan Hu +4
Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are gen…
Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models
Rong Zhou, Dongping Chen, Zihan Jia +24
Digital twins, as precise digital representations of physical systems, have evolved from passive simulation tools into intelligent and autonomous entities through the integration o…
Optimizing Length Compression in Large Reasoning Models
Zhengxiang Cheng, Dongping Chen, Mingyang Fu +1
Large Reasoning Models (LRMs) have achieved remarkable success, yet they often suffer from producing unnecessary and verbose reasoning chains. We identify a core aspect of this iss…
Reinforced Visual Perception with Tools
Zetong Zhou, Dongping Chen, Zixian Ma +6
Visual reasoning, a cornerstone of human intelligence, encompasses complex perceptual and logical processes essential for solving diverse visual problems. While advances in compute…