cross-platform interaction 1foundation models 1gui agents 1mobile automation 1reinforcement learning 1
From the 1 of 11 linked papers with an AI index.
Showing cs.MMShow all
3 papers · 1 filter
cs.MM2026
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production
Huanran Hu, Zihui Ren, Dingyi Yang +4
Real-world video creation often involves a complex reasoning workflow of selecting relevant shots from noisy materials, planning missing shots for narrative completeness, and organ…
cs.MM2025
ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
Liangyu Chen, Yichen Xu, Jianzhe Ma +5
Chart editing reduces manual effort in visualization design. Typical benchmarks limited in data diversity and assume access to complete chart code, which is seldom in real-world sc…
cs.MM2024
Unveiling Visual Biases in Audio-Visual Localization Benchmarks
Liangyu Chen, Zihao Yue, Boshen Xu +1
Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding obj…