7 papers
How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective
Chengpiao Huang, Yuhang Wu, Kaizheng Wang
Large language models (LLMs) are increasingly used to simulate survey responses, but synthetic data can be misaligned with the human population, leading to unreliable inference. We…
Spend Wisely: Maximizing Post-Training Gains in Iterative Synthetic Data Bootstrapping
Pu Yang, Yunzhen Feng, Ziyuan Chen +2
Modern foundation models often undergo iterative ``bootstrapping'' in their post-training phase: a model generates synthetic data, an external verifier filters out low-quality samp…
\textsc{Gen2Real}: Towards Demo-Free Dexterous Manipulation by Harnessing Generated Video
Kai Ye, Yuhang Wu, Shuyuan Hu +4
Dexterous manipulation remains a challenging robotics problem, largely due to the difficulty of collecting extensive human demonstrations for learning. In this paper, we introduce…
Performance of LLMs on Stochastic Modeling Operations Research Problems: From Theory to Practice
Akshit Kumar, Tianyi Peng, Yuhang Wu +1
Large language models (LLMs) have exhibited expert-level capabilities across various domains. However, their abilities to solve problems in Operations Research (OR) -- the analysis…
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
Yuhang Wu, Wenmeng Yu, Yean Cheng +5
Evaluating the alignment capabilities of large Vision-Language Models (VLMs) is essential for determining their effectiveness as helpful assistants. However, existing benchmarks pr…
Grasp What You Want: Embodied Dexterous Grasping System Driven by Your Voice
Junliang Li, Kai Ye, Haolan Kang +6
In recent years, as robotics has advanced, human-robot collaboration has gained increasing importance. However, current robots struggle to fully and accurately interpret human inte…