3 papers
cs.CL2026
What to Format and How: A Benchmark and Workflow Approach for Document Formatting
Shihao Rao, Liang Li, Jiapeng Liu +6
Recent advances in large language models (LLMs) have opened up new possibilities for automated document formatting. However, real-world formatting often requires identifying target…
cs.LG2026
EXaMCaP: Subset Selection with Entropy Gain Maximization for Probing Capability Gains of Large Chart Understanding Training Sets
Jiapeng Liu, Liang Li, Bing Li +5
Recent works focus on synthesizing Chart Understanding (ChartU) training sets to inject advanced chart knowledge into Multimodal Large Language Models (MLLMs), where the sufficienc…
cs.CR2025
Reconstruction of Differentially Private Text Sanitization via Large Language Models
Shuchao Pang, Zhigang Lu, Haichen Wang +3
Differential privacy (DP) is the de facto privacy standard against privacy leakage attacks, including many recently discovered ones against large language models (LLMs). However, w…