works on

From the 1 of 14 linked papers with an AI index.

activity
20242026
collaborators

14 papers

cs.CV2026

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

Dayong Liu, Chao Xu, Weihong Chen +5

The paper introduces CFG-Bench, a benchmark of videos and QA pairs to evaluate how well multimodal language models can generate fine-grained action instructions and higher-order re…

cs.CV2026

Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models

Yijie Qian, Juncheng Wang, Chao Xu +6

As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality…

cs.CV2026

Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies

Yuxiang Feng, Juncheng Wang, Chao Xu +7

Forecasting the future anatomy of slow-evolving neurodegenerative diseases could enable earlier, more targeted intervention and improve clinical trial design, but it remains challe…

cs.LG2026

GIPO: Gaussian Importance Sampling Policy Optimization

Chengxuan Lu, Zhenquan Zhang, Shukuan Wang +3

Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor da…

cs.CV2026

DenseControl: Instance-Level Controllable Synthesis of Dense Crowd Image

Juncheng Wang, Lei Shang, Wang Lu +2

In this paper, we introduce DenseControl, a novel pipeline for generating dense crowd images. Specifically, DenseControl meticulously positions and sizes each generated instance to…

cs.LG2026

AcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models

Chengxuan Lu, Shukuan Wang, Yanjie Li +10

Reinforcement learning (RL) for large-scale Vision-Language-Action (VLA) models is severely bottlenecked by synchronization barriers and the high cost of environment data acquisiti…