5 papers · 1 filter
World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
Qu Tang, Benhui Zhuang, Bo Yuan +3
Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes…
REC-RL: Referring expression counting via Gaussian and range-based reward optimization
Hui Liu, Yunlai Teng, Kunlong Bai +4
Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorporate language for visual un…
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
Song Wu, Xinyu Chen, Qian Wang +3
Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the…
GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
Jiayin Sun, Caixia Sun, Boyu Yang +7
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structu…
ReDDiT: Rehashing Noise for Discrete Visual Generation
Tianren Ma, Xiaosong Zhang, Boyu Yang +2
In the visual generative area, discrete diffusion models are gaining traction for their efficiency and compatibility. However, pioneered attempts still fall behind their continuous…