From the 1 of 15 linked papers with an AI index.
5 papers · 1 filter
Infinity-Parser2 Technical Report
Zuming Huang, Jun Huang, Kexuan Ren +12
Infinity-Parser2 is a large multimodal model that uses a controllable synthetic data pipeline and multi‑task reinforcement learning to parse documents, offering two variants (Flash…
Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding
Jingyuan Huang, Zuming Huang, Yucheng Shi +4
Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in high-resolution screenshots and predict precise screen coordina…
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
Haozhe Wang, Long Li, Chao Qu +4
Recent advances in mathematical problem-solving with language models (LMs) integrate chain-of-thought (CoT) reasoning and code execution to harness their complementary strengths. H…
Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents
Rui Xu, MingYu Wang, XinTao Wang +4
Recent advances in LLM-based role-playing language agents (RPLAs) have attracted broad attention in various applications. While chain-of-thought reasoning has shown importance in m…
MINDECHO: Role-Playing Language Agents for Key Opinion Leaders
Rui Xu, Dakuan Lu, Xiaoyu Tan +5
Large language models~(LLMs) have demonstrated impressive performance in various applications, among which role-playing language agents (RPLAs) have engaged a broad user base. Now,…