14 papers
Generative World Renderer at the Speed of Play
Guixu Lin, Zheng-Hui Huang, Siqi Yang +3
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/cont…
H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions
Shiping Zhu, Yibo Yang, Zhengyang Wang +3
Large language model agents are increasingly deployed in human-human interaction settings, such as meeting assistants and clinical documentation systems, where they must observe co…
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
Bowen Sun, Yujun Cai, Ming-Hsuan Yang +2
Video LLMs suffer from temporal inconsistency: small shifts in frame timing can flip attention and suppress relevant frames. We trace this instability to the common extension of Ro…
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
Pengrui Lu, Shiqi Zhang, Yunzhong Hou +8
Recent coding agents can generate complete codebases from simple prompts, yet existing evaluations focus on issue-level bug fixing and lag behind end-to-end development. We introdu…
Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
Bowen Sun, Yujun Cai, Ming-Hsuan Yang +1
Discrete diffusion language models have shown strong potential for text generation, yet standard supervised fine-tuning (SFT) misaligns with their semi-autoregressive inference: tr…
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
Haonan Ge, Yiwei Wang, Ming-Hsuan Yang +1
Large Vision-Language Models (LVLMs) have shown strong performance across multimodal tasks. However, they often produce hallucinations -- text that is inconsistent with visual inpu…