papers

Publications (28)

cs.CV2024

Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?

Xingyu Fu, Muyu He, Yujie Lu +2

We present a novel task and benchmark for evaluating the ability of text-to-image(T2I) generation models to produce images that align with commonsense in real life, which we call C…

cs.CL2020

ASTRAL: Adversarial Trained LSTM-CNN for Named Entity Recognition

Jiuniu Wang, Wenjia Xu, Xingyu Fu +2

Named Entity Recognition (NER) is a challenging task that extracts named entities from unstructured text data, including news, articles, social comments, etc. The NER system has be…

cs.CV2022

There is a Time and Place for Reasoning Beyond the Image

Xingyu Fu, Ben Zhou, Ishaan Preetam Chandratreya +2

Images are often more significant than only the pixels to human eyes, as we can infer, associate, and reason with contextual information from other sources to establish a more comp…

cs.CV2025

FormCoach: Lift Smarter, Not Harder

Xiaoye Zuo, Nikos Athanasiou, Ginger Delmas +3

Good form is the difference between strength and strain, yet for the fast-growing community of at-home fitness enthusiasts, expert feedback is often out of reach. FormCoach transfo…

cs.CL2024

Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering

Xingyu Fu, Ben Zhou, Sihao Chen +2

Recent advances in multimodal large language models (LLMs) have shown extreme effectiveness in visual question answering (VQA). However, the design nature of these end-to-end model…

cs.CL2026

Reinforced Attention Learning

Bangzheng Li, Jianmo Ni, Chen Qu +5

Post-training with Reinforcement Learning (RL) has substantially improved reasoning in Large Language Models (LLMs) via test-time scaling. However, extending this paradigm to Multi…