1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.AI2025
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
Hang Wu, Hongkai Chen, Yujun Cai +4
Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of langua…
cs.CL2025★ 1 cited
Structured Attention Matters to Multimodal LLMs in Document Understanding
Chang Liu, Hongkai Chen, Yujun Cai +4
Document understanding remains a significant challenge for multimodal large language models (MLLMs). While previous research has primarily focused on locating evidence pages throug…