21 papers
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
Zhixiang Liang, Yifei Liu, Yidan Huang +5
Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long…
PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding
Zheng Wang, Zhifan Ye, Qi Cheng +8
Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, generating tokens in parallel. This makes them effective draft models f…
From Context to Skills: Can Language Models Learn from Context Skillfully?
Shuzheng Si, Haozhe Zhao, Yu Lei +10
Many real-world tasks require language models (LMs) to reason over complex contexts that exceed their parametric knowledge. This calls for context learning, where LMs directly lear…
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
Haozhe Zhao, Shuzheng Si, Zhenhailong Wang +6
Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains one of the most labor-intensi…
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
Haozhe Zhao, Zefan Cai, Shuzheng Si +5
Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex m…
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
Haozhe Zhao, Shuzheng Si, Liang Chen +4
Large vision-language models (LVLMs) have achieved impressive results in various vision-language tasks. However, despite showing promising performance, LVLMs suffer from hallucinat…