6 papers
VizRAG: Enhancing Retrieval-Augmented Generation with Hypergraph Visualization
Yanbin Wei, Yang Chen, Renling Gan +7
Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic facts among entities, rather than relying solely on binary relationships.…
SoftSkill: Behavioral Compression for Contextual Adaptation
Xijia Tao, Yihua Teng, Xinyu Fu +6
Agent skills are commonly deployed as natural-language Markdown files that encode answer policies, evidence-use habits, and task procedures. These files are readable and portable,…
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
Xijia Tao, Yihua Teng, Xinxing Su +7
Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verific…
CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
Jinpeng Chen, Cheng Gong, Hanbo Li +9
Developing multi-turn interactive tool-use agents is challenging because real-world user needs are often complex and ambiguous, yet agents must execute deterministic actions to sat…
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
Kecheng Chen, Ziru Liu, Xijia Tao +7
Diffusion Language Models (DLMs) have recently achieved significant success due to their any-order generation capabilities. However, existing inference methods typically rely on lo…
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
Weikang Shi, Aldrich Yu, Rongyao Fang +11
While Large Language Models (LLMs) have excelled in textual reasoning, they struggle with mathematical domains like geometry that intrinsically rely on visual aids. Existing approa…