activity
20242026
collaborators

10 papers

cs.AI2026

Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning

Wenxuan Zhang, Yuhui Wang, Donggang Jia +5

Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end t…

cs.AI2026

Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation

Xiaohan Ye, Xu Chen, Zihan Gong +15

The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to complex multimodal interactions that sea…

cs.CV2026

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

Xuyang Wang, Zhenyu Li, Jian Ding +4

Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typically separate geometry recons…

cs.AI2026

Chinese Labor Law Large Language Model Benchmark

Zixun Lan, Maochun Xu, Yifan Ren +7

Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose mod…

cs.CV2025

iMotion-LLM: Instruction-Conditioned Trajectory Generation

Abdulwahab Felemban, Nussair Hroub, Jian Ding +4

We introduce iMotion-LLM, a large language model (LLM) integrated with trajectory prediction modules for interactive motion generation. Unlike conventional approaches, it generates…

cs.CV2025

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

Kirolos Ataallah, Eslam Abdelrahman, Mahmoud Ahmed +4

Understanding long-form videos, such as movies and TV episodes ranging from tens of minutes to two hours, remains a significant challenge for multi-modal models. Existing benchmark…