works on

From the 1 of 40 linked papers with an AI index.

activity
20242026
collaborators

40 papers

cs.CV2026

UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets

Xian Li, Rong Wei, Lujie Yang +6

The paper presents UniPhys, a framework that automatically converts raw 3D models into simulation-ready assets with unified physical semantics, and UniPhysGen, a model that jointly…

cs.AI2026

SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments

Yundaichuan Zhan, Minghe Gao, Zhongqi Yue +7

Recent works have explored integrating Vision-Language Models (VLMs) with classical planners that rely on symbolic representations of planning problems to generate long-horizon pla…

cs.CL2026

CoMoL: Efficient Mixture of LoRA Experts via Dynamic Core Space Merging

Jie Cao, Zhenxuan Fan, Zhuonan Wang +8

Large language models (LLMs) achieve remarkable performance on diverse downstream and domain-specific tasks via parameter-efficient fine-tuning (PEFT). However, existing PEFT metho…

cs.CV2026

InstructSAM: Segment Any Instance with Any Instructions

Yuqian Yuan, Wentong Li, Zhaocheng Li +6

In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions. We formulates instruction-driven…

cs.AI2026

Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration

Weile Chen, Bingchen Miao, Qifan Yu +6

Recent advances in Multimodal Large Language Models (MLLMs) have led to promising progress in web agents. However, existing web agents often rely on handcrafted execution pipelines…

cs.CV2026

VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

Mingjian Gao, Wenqiao Zhang, Yuqian Yuan +9

Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irr…