20 papers
ChatImage: Navigating Long-Form LLM Answers through Interactive Images
Wencan Jiang, Jiangning Zhang, Yong Liu
Large Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented as dense linear text, which makes fine-grained inspection, n…
R-Searcher: Calibrating Retrieval and Reasoning Boundaries for Agentic Search
Sheng Zhang, Junyi Li, Wenlin Zhang +6
Recent search agents for multi-hop reasoning often fail by either retrieving incomplete evidence or reasoning over irrelevant portions of the retrieved content, leading to a retrie…
RAGR: Review-Augmented Generative Recommendation
Yingyi Zhang, Junyi Li, Yejing Wang +8
Sequential recommendation (SR) is traditionally formulated as next-item prediction over chronological item interactions. Although recent generative recommendation (GR) methods intr…
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
Junyao Yang, Chen Qian, Kun Wang +4
The advancement of Large Reasoning Models (LRMs) has catalyzed a paradigm shift from reactive ``fast thinking'' text generation to systematic, step-by-step ``slow thinking'' reason…
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
Yu Yang, Yue Liao, Jianbiao Mei +11
Long-horizon action-conditioned video generation aims to synthesize temporally coherent videos that follow complex action instructions over extended horizons, requiring procedural…
PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset
Haojun Chen, Haoyang He, Chengming Xu +11
Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imagin…