collaborators

6 papers

cs.AI2026

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

Chunyang Jiang, Pingping Zhang, Yuzhi Zhao +9

Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimod…

cs.CV2026

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

Yiyang Cai, Nan Chen, Rongchang Xie +8

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most app…

cs.SD2026

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

Sitong Cheng, Weizhen Bian, Songjun Cao +9

Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue),…

cs.CL2026

Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks

Chunyang Jiang, Yonggang Zhang, Yiyang Cai +5

The rising cost of acquiring supervised data has driven significant interest in self-improvement for large language models (LLMs). Straightforward unsupervised signals like majorit…

cs.AI2025

PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents

Junjie Wang, Yuxiang Zhang, Minghao Liu +19

Recent advancements in large multimodal models (LMMs) have leveraged extensive multimodal datasets to enhance capabilities in complex knowledge-driven tasks. However, persistent ch…

cs.LG2025

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge

Chi-Min Chan, Chunpu Xu, Jiaming Ji +7

The current focus of AI research is shifting from emphasizing model training towards enhancing evaluation quality, a transition that is crucial for driving further advancements in…