activity
20232025
most citedTrack the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues

1 citations · 2 across the 7 of their papers we have counts for

collaborators

7 papers

cs.LG2025

Multi-Modal Molecular Representation Learning via Structure Awareness

Rong Yin, Ruyue Liu, Xiaoshuai Hao +4

Accurate extraction of molecular representations is a critical step in the drug discovery process. In recent years, significant progress has been made in molecular representation l…

cs.CV2025

Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition

Yifei Zhang, Chang Liu, Jin Wei +4

Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while t…

cs.LG2024

Communication-Efficient Personalized Federal Graph Learning via Low-Rank Decomposition

Ruyue Liu, Rong Yin, Xiangzhen Bo +5

Federated graph learning (FGL) has gained significant attention for enabling heterogeneous clients to process their private graph data locally while interacting with a centralized…

cs.CV20241 cited

Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues

Yan Zhang, Gangyan Zeng, Huawen Shen +3

Video text-based visual question answering (Video TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. I…

cs.CL2024

Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation

Xunyu Zhu, Jian Li, Can Ma +1

Large Language Models (LLMs) demonstrate exceptional reasoning capabilities, often achieving state-of-the-art performance in various tasks. However, their substantial computational…

cs.CL2024

Unifying Structured Data as Graph for Data-to-Text Pre-Training

Shujie Li, Liang Li, Ruiying Geng +8

Data-to-text (D2T) generation aims to transform structured data into natural language text. Data-to-text pre-training has proved to be powerful in enhancing D2T generation and yiel…