collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

Arash Akbari, Arman Akbari, Masih Eskandar +11

Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressiv…

cs.CV2026

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

Yiming Liang, Yixiao Chen, Yiyang Zhou +8

Many video reasoning tasks require tracking motion, temporal order, and evolving visual states across frames. Existing methods built on large vision-language models (LVLMs) often a…

cs.CV2025

Dual-Flow: Transferable Multi-Target, Instance-Agnostic Attacks via In-the-wild Cascading Flow Optimization

Yixiao Chen, Shikun Sun, Jianshu Li +3

Adversarial attacks are widely used to evaluate model robustness, and in black-box scenarios, the transferability of these attacks becomes crucial. Existing generator-based attacks…

cs.CV2025

RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers

Min Zhao, Guande He, Yixiao Chen +3

Recent advancements in video generation have enabled models to synthesize high-quality, minute-long videos. However, generating even longer videos with temporal coherence remains a…

cs.CV2025

Context-Aware Autoregressive Models for Multi-Conditional Image Generation

Yixiao Chen, Zhiyuan Ma, Guoli Jia +3

Autoregressive transformers have recently shown impressive image generation quality and efficiency on par with state-of-the-art diffusion models. Unlike diffusion architectures, au…