activity
20242026
collaborators

9 papers

cs.LG2026

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

Rui Yang, Qianhui Wu, Yuxi Chen +7

Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despite rapid progress, the stronges…

cs.CV2026

Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA

Yuyang Sun, Yongliang Wu, Xingyu Zhu +8

TimeLogicQA evaluates whether video question answering systems can reason over temporal relations such as event existence, ordering, persistence, boundary conditions, and overlap.…

cs.CV2026

Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge

Yuyang Sun, Yongliang Wu, Xingyu Zhu +8

VRR-QA evaluates whether video-language systems can infer spatial, temporal, viewpoint, depth, and visibility relations that are not always resolved by a single frame. We present a…

cs.CV2026

Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R

Yuyang Sun, Yongliang Wu, Xingyu Zhu +8

We describe \emph{Dual-Route Top-K Retrieval with 1v1 VLM Reranking} for the CoVR-R challenge. The method treats composed video retrieval as two coupled problems: finding a suffici…

cs.CY2026

Measuring and Mitigating Bias in Code Generated by Large Language Models

Yuxi Chen, Yutian Tang, Timothy Storer

Large language models (LLMs) are widely recognised for their applications in natural language generation and are increasingly used for code generation tasks. However, concerns abou…

cs.CV2026

STORM: End-to-End Referring Multi-Object Tracking in Videos

Zijia Lu, Jingru Yi, Jue Wang +4

Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMO…