3 papers
cs.CV2026
State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading
Yuanze Hu, Gen Li, Yuqin Lan +5
Multimodal large language models (MLLMs) have achieved impressive progress on general multimodal tasks, yet they remain brittle on dial-based measurement reading. In this paper, we…
cs.CL2026
DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing
Hongzhi Zhang, Yuanze Hu, Tinghai Zhang +9
The evolution of Large Language Models (LLMs) towards autonomous agents has catalyzed progress in Deep Research. While retrieval capabilities are well-benchmarked, the post-retriev…
cs.CV2025
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
Guanwen Feng, Zhiyuan Ma, Yunan Li +3
Recent advances in audio-driven talking head generation have achieved impressive results in lip synchronization and emotional expression. However, they largely overlook the crucial…