activity
20032024
most citedPanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback

13 citations · 26 across the 14 of their papers we have counts for

collaborators

17 papers

cs.CL20242 cited

CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin +7

Due to the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Despite the datasets l…

cs.CL2024

SPUQ: Perturbation-Based Uncertainty Quantification for Large Language Models

Xiang Gao, Jiaxin Zhang, Lalla Mouatadid +1

In recent years, large language models (LLMs) have become increasingly prevalent, offering remarkable text generation capabilities. However, a pressing challenge is their tendency…

cs.LG2024

Discriminant Distance-Aware Representation on Deterministic Uncertainty Quantification Methods

Jiaxin Zhang, Kamalika Das, Sricharan Kumar

Uncertainty estimation is a crucial aspect of deploying dependable deep learning models in safety-critical systems. In this study, we introduce a novel and efficient method for det…

cs.CL2024

DCR-Consistency: Divide-Conquer-Reasoning for Consistency Evaluation and Improvement of Large Language Models

Wendi Cui, Jiaxin Zhang, Zhuohang Li +4

Evaluating the quality and variability of text generated by Large Language Models (LLMs) poses a significant, yet unresolved research challenge. Traditional evaluation methods, suc…

cs.CV2023

On the Quantification of Image Reconstruction Uncertainty without Training Data

Sirui Bi, Victor Fung, Jiaxin Zhang

Computational imaging plays a pivotal role in determining hidden information from sparse measurements. A robust inverse solver is crucial to fully characterize the uncertainty indu…

cs.CL20234 cited

Interactive Multi-fidelity Learning for Cost-effective Adaptation of Language Model with Sparse Human Supervision

Jiaxin Zhang, Zhuohang Li, Kamalika Das +1

Large language models (LLMs) have demonstrated remarkable capabilities in various tasks. However, their suitability for domain-specific tasks, is limited due to their immense scale…