27 papers
Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch
Shuyang Xie, Shuxiao Xie, Feng Zhu +2
Online-judge verdicts and the datasets and benchmarks built on them are treated as ground truth for evaluating and training large language models for code. Yet prior audits have so…
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
Shihao Yuan, Yuanze Li, Ruyi Zhang +2
Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue t…
PreferThinker: Reasoning-based Personalized Image Preference Assessment
Shengqi Xu, Xinpeng Zhou, Yabo Zhang +6
Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing m…
Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization
Zhe Yang, Ruyi Zhang, Hongtao Chen +4
Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training.…
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization
Feng Zhu, Shuyang Xie, Yihan Zeng +2
Real-world image restoration is challenging due to complex and interacting mixed degradations. Recent agent-based approaches address this problem by composing multiple task-specifi…
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
Haodong Yu, Yabo Zhang, Donglin Di +2
While diffusion models excel at generating images with conventional dimensions, pushing them to synthesize ultra-high-resolution imagery at extreme aspect ratios (EAR) often trigge…