6 papers
SoftSkill: Behavioral Compression for Contextual Adaptation
Xijia Tao, Yihua Teng, Xinyu Fu +6
Agent skills are commonly deployed as natural-language Markdown files that encode answer policies, evidence-use habits, and task procedures. These files are readable and portable,…
Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models
Kecheng Chen, Ziru Liu, Xijia Tao +9
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive language models, offering stronger global awareness and highly parallel generati…
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
Xijia Tao, Yihua Teng, Xinxing Su +7
Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verific…
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
Kecheng Chen, Ziru Liu, Xijia Tao +7
Diffusion Language Models (DLMs) have recently achieved significant success due to their any-order generation capabilities. However, existing inference methods typically rely on lo…
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
Fuqing Bie, Shiyu Huang, Xijia Tao +6
While generalist foundation models like Gemini and GPT-4o demonstrate impressive multi-modal competence, existing evaluations fail to test their intelligence in dynamic, interactiv…
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
Xijia Tao, Shuai Zhong, Lei Li +2
There has been an increasing interest in the alignment of large language models (LLMs) with human values. However, the safety issues of their integration with a vision module, or v…