2 papers
cs.CV2026
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
Linquan Wu, Tianxiang Jiang, Yifei Dong +6
Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critica…
cs.SE2025
Towards Engineering Multi-Agent LLMs: A Protocol-Driven Approach
Zhenyu Mao, Jacky Keung, Fengji Zhang +3
The increasing demand for software development has driven interest in automating software engineering (SE) tasks using Large Language Models (LLMs). Recent efforts extend LLMs into…