6 papers
LaGen: Towards Autoregressive LiDAR Scene Generation
Sizhuo Zhou, Xiaosong Jia, Fanrui Zhang +7
Generative world models for autonomous driving (AD) are of great value in applications such as data augmentation, closed-loop simulation, and safety-critical scenario evaluation. U…
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
Fanrui Zhang, Qiang Zhang, Sizhuo Zhou +10
Existing image forgery detection (IFD) methods either exploit low-level, semantics-agnostic artifacts or rely on multimodal large language models (MLLMs) with high-level semantic k…
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
Yukang Feng, Jianwen Sun, Chuanhao Li +8
Recent advancements in Large Multimodal Models (LMMs) have significantly improved multimodal understanding and generation. However, these models still struggle to generate tightly…
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
Yukang Feng, Jianwen Sun, Zelai Yang +16
Recent advances in AI-assisted programming have empowered agents to execute complex workflows via command-line interfaces, however, existing benchmarks are limited by short task ho…
Computational Caustic Design for Surface Light Source
Sizhuo Zhou, Yuou Sun, Bailin Deng +1
Designing freeform surfaces to control light based on real-world illumination patterns is challenging, as existing caustic lens designs often assume oversimplified point or paralle…
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
Jianwen Sun, Yukang Feng, Chuanhao Li +7
Unified multimodal understanding and generation have recently received much attention in the area of vision and language. Existing UniMs are designed to simultaneously learn both m…