activity
20242026
collaborators

7 papers

cs.CV2026

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

Dingyu Yao, Junhao Zhou, Chenxu Yang +12

Many moments in the real world do not wait for a user to ask. A fire starts on a security monitor, an expression flickers across a video call, or a product a viewer wants flashes b…

cs.CV2026

AdaCodec: A Predictive Visual Code for Video MLLMs

Haowen Hou, Zhen Huang, Zheming Liang +8

Video is temporally redundant: adjacent frames usually share most objects, background, and layout. Yet existing video multimodal large language models (video MLLMs) usually encode…

cs.SE2026

SEMAG: Self-Evolutionary Multi-Agent Code Generation

Yulin Peng, Haowen Hou, Xinxin Zhu +2

Large Language Models (LLMs) have made significant progress in handling complex programming tasks. However, current methods rely on manual model selection and fixed workflows, whic…

cs.CV2025

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

Nianbo Zeng, Haowen Hou, Fei Richard Yu +2

Despite recent advances in retrieval-augmented generation (RAG) for video understanding, effectively understanding long-form video content remains underexplored due to the vast sca…

cs.CV2025

RWKV-UI: UI Understanding with Enhanced Perception and Reasoning

Jiaxi Yang, Haowen Hou

Existing Visual Language Modelsoften struggle with information loss and limited reasoning abilities when handling high-resolution web interfaces that combine complex visual, textua…

cs.CV2024

VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models

Haowen Hou, Peigen Zeng, Fei Ma +1

Visual Language Models (VLMs) have rapidly progressed with the recent success of large language models. However, there have been few attempts to incorporate efficient linear Recurr…