collaborators

20 papers

cs.AI2026

Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit Systems

Jusheng Zhang, Yijia Fan, Kaitong Cai +6

Vision-Language Models (VLMs) enable powerful multi-agent systems, but scaling them is economically unsustainable: coordinating heterogeneous agents under information asymmetry oft…

cs.CV2025

Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains

Jesen Zhang, Ningyuan Liu, Kaitong Cai +5

Multimodal LLMs often produce fluent yet unreliable reasoning, exhibiting weak step-to-step coherence and insufficient visual grounding, largely because existing alignment approach…

cs.CV2025

CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation

Qinglin Zeng, Kaitong Cai, Ruiqi Chen +2

Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independentl…

cs.LG2025

RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks

Ningyuan Liu, Jing Yang, Kaitong Cai +1

Full parameter fine tuning is a key technique for adapting large language models (LLMs) to downstream tasks, but it incurs substantial memory overhead due to the need to cache exte…

cs.CV2025

FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models

Kaitong Cai, Jusheng Zhang, Jing Yang +4

Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy…

cs.CV2025

SirenPose: Dynamic Scene Reconstruction via Geometric Supervision

Kaitong Cai, Jensen Zhang, Jing Yang +1

We introduce SirenPose, a geometry-aware loss formulation that integrates the periodic activation properties of sinusoidal representation networks with keypoint-based geometric sup…