From the 1 of 24 linked papers with an AI index.
24 papers
One Query, Many Scales: Sparse Mixture-of-Experts for Efficient Hierarchical Cross-View Geo-Localization
Ruijie Fan, Junyan Ye, Qi Zhu +1
Cross-view geo-localization (CVGL) retrieves geo-tagged satellite imagery for a ground-view query. Most systems exhaustively search a flat, fixed-resolution gallery, incurring high…
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Haodong Li, Tianfei Ren, Xiaoxiao Ma +25
The paper presents VideoCoCo, a system that generates physically consistent videos by having a coding agent produce executable Blender code that defines the scene and its dynamics,…
GenClaw: Code-Driven Agentic Image Generation
Junyan Ye, Jun He, Zilong Huang +4
Image generation models have evolved from text-conditioned pixel synthesis toward multimodal agents endowed with visual comprehension and tool invocation capabilities. Yet, existin…
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
Leqi Zhu, Junyan Ye, Kaiqing Lin +3
The development of generative artificial intelligence technologies has propelled the visual realism of synthetic images to an unprecedented level. Although current interpretable de…
OmniAID: Decoupling Semantics and Artifacts for Universal AI-Generated Image Detection in the Wild
Yuncheng Guo, Junyan Ye, Chenjue Zhang +4
A truly universal AI-Generated Image (AIGI) detector must simultaneously generalize across diverse generative models and varied semantic content. Current methods learn a single, en…
SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors
Ruijie Fan, Junyan Ye, Huan Chen +3
Existing satellite video tracking methods often struggle with generalization, requiring scenario-specific training to achieve satisfactory performance, and are prone to track loss…