2 citations · 3 across the 5 of their papers we have counts for
6 papers
Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning
Shi-Yu Tian, Zhuo-Xia Wang, Xuan-Yi Zhu +6
Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that demand both precise spatial per…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
PRISM: A Unified Framework for Post-Training LLMs Without Verifiable Rewards
Mukesh Ghimire, Aosong Feng, Liwen You +3
Current techniques for post-training Large Language Models (LLMs) rely either on costly human supervision or on external verifiers to boost performance on tasks such as mathematica…
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator
Zhuotong Chen, Fang Liu, Xuan Zhu +2
Existing studies on preference optimization (PO) have centered on constructing pairwise preference data following simple heuristics, such as maximizing the margin between preferred…
Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
Xiao-Wen Yang, Xuan-Yi Zhu, Wen-Da Wei +5
The integration of slow-thinking mechanisms into large language models (LLMs) offers a promising way toward achieving Level 2 AGI Reasoners, as exemplified by systems like OpenAI's…
TaeBench: Improving Quality of Toxic Adversarial Examples
Xuan Zhu, Dmitriy Bespalov, Liwen You +2
Toxicity text detectors can be vulnerable to adversarial examples - small perturbations to input text that fool the systems into wrong detection. Existing attack algorithms are tim…