13 papers
OmniPresent: Generating Coherent Presentation Suites from Scientific Papers
Qianli Ma, Jipeng Xiao, Siyu Wang +7
Transforming static research papers into dynamic media such as posters, slides, and videos is essential for effective dissemination but remains a labor-intensive challenge. Existin…
Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression
Shuochen Chang, Qingyang Liu, Shaobo Wang +8
Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and extended inference time. La…
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
Shuochen Chang, Tong Bai, Xiaofeng Zhang +5
Latent reasoning enables Large Language Models (LLMs) to perform multi-step inference within continuous hidden states, offering efficiency gains over explicit Chain-of-Thought (CoT…
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
Qingyang Liu, Bingjie Gao, Canmiao Fu +9
Recent unified models integrate multimodal understanding and generation within a single framework. However, an "understanding-generation gap" persists, where models can capture use…
The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates
Shaobo Wang, Guo Chen, Ziyue Wang +5
With the rapid progress of large language models (LLMs), reliably evaluating the capabilities of pre-trained LLMs has become increasingly important. The challenge is that base pre-…
RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
Bingjie Gao, Qianli Ma, Xiaoxue Wu +9
Prompt design plays a crucial role in text-to-video (T2V) generation, yet user-provided prompts are often short, unstructured, and misaligned with training data, limiting the gener…