works on

From the 1 of 34 linked papers with an AI index.

collaborators

34 papers

cs.CV2026

EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent

Junlong Li, Junxi Li, Yuxiang Yang +3

The paper introduces EgoProceVQA, a new egocentric video question answering benchmark focused on procedural reasoning, and proposes EgoProceAgent, a self-exploring agent that learn…

cs.CV2026

TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts

Boyuan Chen, Zichen Dang, Chuang Yang +2

In real-world deployments, scene text detectors inevitably face distribution shifts beyond the training distribution. Prior work often depends on large-scale scene-text pretraining…

cs.CV2026

GUI-C: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning

Junlong Li, Chao Hao, Lap-Pui Chau +1

Existing agentic reinforcement learning methods for GUI grounding have limitations at two levels. At the data level, current approaches typically treat all training samples equally…

cs.CV2026

SP-MoMamba: Superpixel-driven Mixture of State Space Experts for Efficient Image Super-Resolution

Wenbin Zou, Yawen Cui, Yi Wang +5

State space models (SSMs) have emerged as a powerful paradigm for efficient single-image super-resolution (SR) due to their linear complexity and long-range modeling capabilities.…

cs.RO2026

RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation

Sixu Lin, Junliang Chen, Huaiyuan Xu +8

Planning and acting in 3D environments is a fundamental capability for robotic manipulation in the real world. Although prior work has explored predictive flow planners to guide 3D…

cs.CV2026

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding

Yuejiao Su, Xinshen Zhang, Zhen Ye +3

Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing multimodal large language mod…