7 papers
SubdivAR: Autoregressive Next-Scale Prediction for Neural Mesh Subdivision
Huipeng Guo, Zikai Song, Hang Long +7
Mesh subdivision is a fundamental operation for converting coarse, editable meshes into high-resolution surfaces, with broad applications in digital asset creation. Classical rule-…
SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models
Zhengxuan Wei, Yi Dong, Zonghui Li +6
Low-Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, existing LoRA merging techniq…
Large Language Model as Token Compressor and Decompressor
Wenbing Li, Yiran Wang, Zikai Song +4
In this paper, we study whether an off-the-shelf LLM can be adapted into a discrete, variable-length token compressor and decompressor for long-context processing. To this end, we…
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
Wenbing Li, Zikai Song, Hang Zhou +3
Recent attempts to combine low-rank adaptation (LoRA) with mixture-of-experts (MoE) for multi-task adaptation of Large Language Models (LLMs) often replace whole attention/FFN laye…
Semantic-Aware Logical Reasoning via a Semiotic Framework
Yunyao Zhang, Xinglang Zhang, Junxi Sheng +5
Logical reasoning is a fundamental capability of large language models. However, existing studies often overlook the interaction between logical complexity and semantic complexity,…
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
Zhe Gao, Shiyu Shen, Taifeng Chai +7
Existing Multimodal Large Language Models (MLLMs) often suffer from hallucinations in long video understanding (LVU), primarily due to the imbalance between textual and visual toke…