activity
20242026
collaborators

9 papers

cs.CL2026

Improving Diffusion Language Model Decoding through Joint Search in Generation Order and Token Space

Yangyi Shen, Tianjian Feng, Jiaqi Han +5

Diffusion Language Models (DLMs) offer order-agnostic generation that can explore many possible decoding trajectories. However, current decoding methods commit to a single trajecto…

cs.AI2025

LLM-based Content Classification Approach for GitHub Repositories by the README Files

Malik Uzair Mehmood, Shahid Hussain, Wen Li Wang +1

GitHub is the world's most popular platform for storing, sharing, and managing code. Every GitHub repository has a README file associated with it. The README files should contain p…

cs.LG2025

GUI-G: Gaussian Reward Modeling for GUI Grounding

Fei Tang, Zhangxuan Gu, Zhengxi Lu +9

Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use…

eess.AS2025

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

Huadai Liu, Kaicheng Luo, Jialei Wang +4

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like p…

cs.CL2025

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models

He Zhu, Junyou Su, Minxin Chen +4

In the field of urban planning, existing Vision-Language Models (VLMs) frequently fail to effectively analyze and evaluate planning maps, despite the critical importance of these v…

eess.AS2025

OmniAudio: Generating Spatial Audio from 360-Degree Video

Huadai Liu, Tianyi Luo, Kaicheng Luo +11

Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing so…