collaborators

5 papers

cs.CV2026

:Parallel Kinematic Selective State Space Scanners for Efficient Video Understanding

Lingjie Zeng, Hailun Zhang, Xiwen Wang +1

Temporal modeling remains a fundamental challenge in video understanding, particularly as sequence lengths scale. Traditional video models relying on dense spatiotemporal attention…

cs.CV2026

Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios

Xuekang Zhu, Ji-Zhe Zhou, Kaiwen Feng +5

With the large models easing the labor-intensive manipulation process, image manipulations in today's real scenarios often entail a complex manipulation process, comprising a serie…

cs.CV2026

Dehallu3D: Hallucination-Mitigated 3D Generation from Single Image via Cyclic View Consistency Refinement

Xiwen Wang, Shichao Zhang, Hailun Zhang +5

Large 3D reconstruction models have revolutionized the 3D content generation field, enabling broad applications in virtual reality and gaming. Just like other large models, large 3…

cs.CV2024

Mesoscopic Insights: Orchestrating Multi-scale & Hybrid Architecture for Image Manipulation Localization

Xuekang Zhu, Xiaochen Ma, Lei Su +7

The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial techniq…

cs.CV2024

Saliency Guided Optimization of Diffusion Latents

Xiwen Wang, Jizhe Zhou, Xuekang Zhu +2

With the rapid advances in diffusion models, generating decent images from text prompts is no longer challenging. The key to text-to-image generation is how to optimize the results…