collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

Qu Tang, Benhui Zhuang, Bo Yuan +3

Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes…

cs.CV2026

REC-RL: Referring expression counting via Gaussian and range-based reward optimization

Hui Liu, Yunlai Teng, Kunlong Bai +4

Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorporate language for visual un…

cs.CV2026

Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance

Song Wu, Xinyu Chen, Qian Wang +3

Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the…

cs.CV2026

GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning

Jiayin Sun, Caixia Sun, Boyu Yang +7

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structu…

cs.CV2025

ReDDiT: Rehashing Noise for Discrete Visual Generation

Tianren Ma, Xiaosong Zhang, Boyu Yang +2

In the visual generative area, discrete diffusion models are gaining traction for their efficiency and compatibility. However, pioneered attempts still fall behind their continuous…