5 papers
Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding
Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan +11
Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to acti…
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
Hongzhu Yi, Yujia Yang, Yuanxiang Wang +18
In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive manner through natural language inst…
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
Haopeng Jin, Hongzhu Yi, Wenlong Zhao +6
Long-video understanding with multimodal language models suffers from three compounding bottlenecks: heavy decode cost to obtain dense RGB frames, quadratic token growth with frame…
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
Yujia Yang, Yuanxiang Wang, Zhenyu Guan +11
While Instruction-based Image Editing (IIE) has achieved significant progress, existing benchmarks pursue task breadth via mixed evaluations. This paradigm obscures a critical fail…
RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization
Hongzhu Yi, Xinming Wang, Zhenghao zhang +12
Within the domain of large language models, reinforcement fine-tuning algorithms necessitate the generation of a complete reasoning trajectory beginning from the input query, which…