9 papers
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm
Yaofang Liu, Kangning Cui, Meng Chu +7
Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generators still ask users to seria…
DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions
Zishan Shao, Lixun Zhang, Kangning Cui +10
Large language models (LLMs) handle many tasks with one set of parameters, but under KV-cached inference it is unclear what task-general structure, if any, is used at decode time r…
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
Wenhao Wu, Zishan Shao, Kangning Cui +5
SVD-based Low-rank compression reduces transformer parameters and nominal FLOPs, but these savings often translate poorly into real LLM serving speedups. We show that this gap is l…
SphUnc: Hyperspherical Uncertainty Decomposition and Causal Identification via Information Geometry
Rong Fu, Chunlei Meng, Jinshuo Liu +8
Reliable decision-making in complex multi-agent systems requires calibrated predictions and interpretable uncertainty. We introduce SphUnc, a unified framework combining hyperspher…
Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection
Yangchen Zeng, Zhenyu Yu, Dongming Jiang +5
Transformer-based detectors have advanced small-object detection, but they often remain inefficient and vulnerable to background-induced query noise, which motivates deep decoders…
StoryState: Agent-Based State Control for Consistent and Editable Storybooks
Ayushman Sarkar, Zhenyu Yu, Wei Tang +3
Large multimodal models have enabled one-click storybook generation, where users provide a short description and receive a multi-page illustrated story. However, the underlying sto…