7 papers
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5
Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, an…
LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling
Wenkai Chen, Tianshu Li, Wenyong Huang +3
Mixture-of-Experts (MoE) and looped architectures scale models along two orthogonal axes, namely parameter capacity and effective depth. However, mainstream looped architectures re…
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning
Jianbo Lin, Xiaomin Yu, Yi Xin +7
Large language model-based agents make mistakes, yet critique can often guide the same model toward correct behavior. However, when critique is removed, the model may fail again on…
Anisotropic Modality Align
Xiaomin Yu, Yijiang Li, Yuhui Zhang +8
Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the shared representation space of…
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
Xiaomin Yu, Yi Xin, Yuhui Zhang +12
Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of d…
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
Chengwen Liu, Xiaomin Yu, Zhuoyue Chang +15
In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore ne…