12 papers
Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents
Kaituo Zhang, Zhen Xiong, Mingyu Zhong +4
Tool-augmented reasoning has become a popular direction for LLM-based agents, and it is widely assumed to improve reasoning and reliability. However, we demonstrate that this conse…
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
Zhecheng Li, Guoxian Song, Yiwei Wang +3
Grounding natural language queries in graphical user interfaces (GUIs) presents a challenging task that requires models to comprehend diverse UI elements across various application…
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
Zhen Xiong, Yujun Cai, Zhecheng Li +2
Recent Large Audio-Language Models (LALMs) have shown strong performance on various audio understanding tasks such as speech translation and Audio Q\&A. However, they exhibit signi…
Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
Zhen Xiong, Yujun Cai, Zhecheng Li +1
Controllable generation is a fundamental task in NLP with many applications, providing a basis for function calling to agentic communication. However, even state-of-the-art autoreg…
: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
Zhecheng Li, Guoxian Song, Yiwei Wang +3
Img2LaTeX is a practically important task that involves translating mathematical expressions and structured visual content from images into LaTeX code. In recent years, vision-lang…
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
Zhecheng Li, Guoxian Song, Yujun Cai +3
Modern Vision-Language Models (VLMs) exhibit remarkable visual and linguistic capabilities, achieving impressive performance in various tasks such as image recognition and object l…