2 papers
cs.CL2026
What to Format and How: A Benchmark and Workflow Approach for Document Formatting
Shihao Rao, Liang Li, Jiapeng Liu +6
Recent advances in large language models (LLMs) have opened up new possibilities for automated document formatting. However, real-world formatting often requires identifying target…
cs.LG2026
EXaMCaP: Subset Selection with Entropy Gain Maximization for Probing Capability Gains of Large Chart Understanding Training Sets
Jiapeng Liu, Liang Li, Bing Li +5
Recent works focus on synthesizing Chart Understanding (ChartU) training sets to inject advanced chart knowledge into Multimodal Large Language Models (MLLMs), where the sufficienc…