7 papers
HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment
Chuangxin Zhao, Boyan Shi, Yanling Wang +7
Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate reasoning step appears in noi…
An LMM for Precisely Grounding Elements in Documents
Yijian Lu, Chuangxin Zhao, Kai Sun +3
Visual grounding in documents is a crucial ability for Large Multimodal Models (LMMs) in areas such as document understanding, deep research and document error detection. However,…
Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs
Chuangxin Zhao, Canran Xiao, Siyuan Ma +5
Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-t…
Reasoning emerges from constrained inference manifolds in large language models
Yanbiao Ma, Fei Luo, Linfeng Zhang +10
Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inference. Here we study reasonin…
ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and Expression
Mingxuan Wang, Cheng Chen, Gaoyang Jiang +4
Single-cell RNA-seq profiles are high-dimensional, sparse, and unordered, causing autoregressive generation to impose an artificial ordering bias and suffer from error accumulation…
Curiosity Meets Cooperation: A Game-Theoretic Approach to Long-Tail Multi-Label Learning
Canran Xiao, Chuangxin Zhao, Zong Ke +1
Long-tail imbalance is endemic to multi-label learning: a few head labels dominate the gradient signal, while the many rare labels that matter in practice are silently ignored. We…