activity
20242026
collaborators

5 papers

cs.CV2026

Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation

Yuhan Liu, Lianhui Qin, Shengjie Wang

Large Vision-Language Models (VLMs) have achieved remarkable progress in multimodal understanding, yet they struggle when reasoning over information-intensive images that densely i…

cs.AI2025

WebRenderBench: Enhancing Web Interface Generation through Layout-Style Consistency and Reinforcement Learning

Peichao Lai, Jinhui Zhuang, Kexuan Zhang +6

Automating the conversion of UI images into web code is a critical task for front-end development and rapid prototyping. Advances in multimodal large language models (MLLMs) have m…

cs.CL2025

Reveal and Release: Iterative LLM Unlearning with Self-generated Data

Linxi Xie, Xin Teng, Shichang Ke +2

Large language model (LLM) unlearning has demonstrated effectiveness in removing the influence of undesirable data (also known as forget data). Existing approaches typically assume…

cs.CL2025

Inter-Passage Verification for Multi-evidence Multi-answer QA

Bingsen Chen, Shengjie Wang, Xi Ye +1

Multi-answer question answering (QA), where questions can have many valid answers, presents a significant challenge for existing retrieval-augmented generation-based QA systems, as…

cs.CV2024

Diffusion Cocktail: Mixing Domain-Specific Diffusion Models for Diversified Image Generations

Haoming Liu, Yuanhe Guo, Shengjie Wang +1

Diffusion models, capable of high-quality image generation, receive unparalleled popularity for their ease of extension. Active users have created a massive collection of domain-sp…