10 papers
CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs
Zihan Nan, Yang Gu, Wei Liu +4
Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primaril…
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows
Wei Liu, Yang Gu, Xi Yan +5
Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines. While recent LLM-based approac…
How Mobile World Model Guides GUI Agents?
Weikai Xu, Kun Huang, Yunren Feng +10
Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable prediction of action consequences…
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
Ruicheng Zhang, Kaixi Cong, Jun Zhou +5
Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rely on noise-based exploration…
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
Zishan Liu, Ruoxi Zang, Yanglin Zhang +5
Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle with spatial reasoning, ofte…
V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
Zhiwei Ning, Xuanang Gao, Jiaxi Cao +6
Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a persistent challenge. Although re…