19 papers
Terminal Agents: A Survey of AI Agents in Command-Line Environments
Yi Bin, Xiaoyang Yuan, Haoxi Zeng +9
Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software engineering, tool use, and computer-use res…
WSA: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control
Jiahao Jiang, Jianing Zhang, Zhenhan Yin +8
Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for generalist robotic systems to date. By leveraging imitation learning on…
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
Cheng Chen, Yuyu Guo, Pengpeng Zeng +4
Vision-Language Models (VLMs) create a severe visual feature bottleneck by using a crude, asymmetric connection that links only the output of the vision encoder to the input of the…
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
Xiao Cai, Pengpeng Zeng, Ji Zhang +3
Precise spatial fidelity in Image-to-3D multi-instance generation is critical for downstream real-world applications. Recent work attempts to address this by fine-tuning pre-traine…
Janus-LoRA: A Balanced Low-Rank Adaptation for Continual Learning
Cheng Chen, Pengpeng Zeng, Yuyu Guo +3
Low-Rank Adaptation (LoRA) has emerged as a promising paradigm for Continual Learning. It independently updates its low-rank factors ( and ), creating a composite update to t…
Reversible Inversion for Training-Free Exemplar-guided Image Editing
Yuke Li, Lianli Gao, Ji Zhang +5
Exemplar-guided Image Editing (EIE) aims to modify a source image according to a visual reference. Existing approaches often require large-scale pre-training to learn relationships…