2 papers
cs.AI2025
SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation
Xin Liang, Xiang Zhang, Yiwei Xu +2
Generating academic slides from scientific papers is a challenging multimodal reasoning task that requires both long context understanding and deliberate visual planning. Existing…
cs.CV2025
Images are Worth Variable Length of Representations
Lingjun Mao, Rodolfo Corona, Xin Liang +2
Most existing vision encoders map images into a fixed-length sequence of tokens, overlooking the fact that different images contain varying amounts of information. For example, a v…