9 papers
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
Tianci Liu, Zihan Dong, Tianchun Li +8
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a…
PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation
Pengyuan Guo, Zhonghao Mai, Zhengtong Xu +8
Recent advances in vision-language models (VLMs) have enabled increasing progress in real-world robot manipulation. However, long-horizon manipulation in unstructured environments…
In-Context Compositional Learning via Sparse Coding Transformer
Wei Chen, Jingxi Yu, Zichen Miao +1
Transformer architectures have achieved remarkable success across language, vision, and multimodal tasks, and there is growing demand for them to address in-context compositional l…
DiffOG: Differentiable Policy Trajectory Optimization with Generalizability
Zhengtong Xu, Zichen Miao, Qiang Qiu +2
Imitation learning-based visuomotor policies excel at manipulation tasks but often produce suboptimal action trajectories compared to model-based methods. Directly mapping camera d…
TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription
Guo Yutong, Wanying Wang, Yue Wu +2
Table Visual Question Answering (Table VQA) is typically addressed by large vision-language models (VLMs). While such models can answer directly from images, they often miss fine-g…
Sparse Fine-Tuning of Transformers for Generative Tasks
Wei Chen, Jingxi Yu, Zichen Miao +1
Large pre-trained transformers have revolutionized artificial intelligence across various domains, and fine-tuning remains the dominant approach for adapting these models to downst…