14 papers
ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services
Fengxian Ji, Jingpu Yang, Zirui Song +5
Recent image generation and editing models demonstrate robust adherence to instructions and high visual quality on academic benchmarks. However, their performance on paid, real-wor…
One Model, Multiple Goals: Adaptive Multi-Objective Learning for E-commerce Dialogue Systems
Mingzhe Li, Jing Xiang, Enguo Zhou +5
Dialogue systems in e-commerce scenarios often need to satisfy multiple objectives: accurately reasoning over user profiles (e.g., eligibility, credit limit) to ensure correct deci…
The Cylindrical Representation Hypothesis for Language Model Steering
Lang Gao, Jinghui Zhang, Wei Liu +7
Steering is a widely used technique for controlling large language models, yet its effects are often unstable and hard to predict. Existing theoretical accounts are largely based o…
A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
Kaiyang Wan, Lang Gao, Honglin Mu +3
Multi-Hop Question Answering (MHQA) requires integrating dispersed, interdependent evidence through sequential reasoning under noise. This task is challenging for LLMs as they have…
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
Zirui Song, Yuan Huang, Junchang Liu +7
Social deduction games like Werewolf combine language, reasoning, and strategy, providing a testbed for studying natural language and social intelligence. However, most studies red…
Do LLMs "Feel"? Emotion Circuits Discovery and Control
Chenxi Wang, Yixuan Zhang, Ruiji Yu +8
As the demand for emotional intelligence in large language models (LLMs) grows, a key challenge lies in understanding the internal mechanisms that give rise to emotional expression…