2 papers
cs.AI2026
Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic
Chuou Xu, Liya Ji, Qifeng Chen
Reinforcement learning (RL) as post-training is crucial for enhancing the reasoning ability of large language models (LLMs) in coding and math. However, their capacity for visual s…
cs.CV2026
Instruction-based Image Editing with Planning, Reasoning, and Generation
Liya Ji, Chenyang Qi, Qifeng Chen
Editing images via instruction provides a natural way to generate interactive content, but it is a big challenge due to the higher requirement of scene understanding and generation…