collaborators

7 papers

cs.CV2026

Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models

Jitai Hao, Hao Liu, Xinyan Xiao +2

Unified Multimodal Models (UMMs) built on shared autoregressive (AR) transformers are attractive for their architectural simplicity. However, we identify a critical limitation: whe…

cs.CL2025

A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone

Jitai Hao, Qiang Huang, Hao Liu +3

Training high-performing Small Language Models (SLMs) remains costly, even with knowledge distillation and pruning from larger teacher models. Existing work often faces three key c…

cs.CV2025

SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design Prior

Haoran Wang, Bo Zhao, Jinghui Wang +5

In this paper, we study the content-aware layout generation problem, which aims to automatically generate layouts that are harmonious with a given background image. Existing method…

cs.CL2025

Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration

Ante Wang, Yujie Lin, Jingyao Liu +4

Critical thinking is essential for building robust AI systems, preventing them from blindly accepting flawed data or biased reasoning. However, prior work has primarily focused on…

cs.CL2025

Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study

Yujie Lin, Ante Wang, Moye Chen +4

Recently, inference-time scaling of chain-of-thought (CoT) has been demonstrated as a promising approach for addressing multi-modal reasoning tasks. While existing studies have pre…

cs.CL2025

UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning

Hongxuan Tang, Hao Liu, Xinyan Xiao

We introduce UGen, a unified autoregressive multimodal model that demonstrates strong performance across text processing, image understanding, and image generation tasks simultaneo…