2 papers
cs.SE2026
Omni2Web: Benchmarking Audiovisual Website Development
Minghao Han, Zhenghao Xing, Xize Cheng +7
Screen-recorded web editing requests contain weak deictic expressions such as ``this'' and ``there,'' whose referents depend on speech, cursor trajectories, page state, and edit hi…
cs.CL2026
Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
Qi Chen, Yunfei Chu, Haolin He +15
Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed tex…