From the 1 of 20 linked papers with an AI index.
20 papers
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
Zedong Yu, Qianxing Li, Zhi Gao +8
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as cl…
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
Yigui Feng, Qinglin Wang, Yang Liu +1
The paper introduces Fre-Res, a dual‑track video token compression method for video multimodal large language models that keeps a few high‑fidelity spatial anchor tokens while enco…
Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework
Jiayi Gao, Qingchao Chen, Yuxin Peng +1
Current image editing methods excel at static attributes but fail at complex Human-Object Interactions (HOI), a critical challenge unaddressed by existing benchmarks that conflate…
Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning
Chengzhen Yu, Canran Xiao, Siyuan Ma +1
Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly con…
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics
Xiaomeng Fu, Junfan Lin, Yang Liu +4
Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and…
Learning with a Single Rollout via Monte Carlo Pass@k Critic
Fengdi Che, Yang Liu, Lei Yu +4
Estimating token-level advantages in reinforcement learning (RL) for language models remains challenging because scaling up episodic experience collection is expensive. The difficu…