5 papers
RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
Shuichiro Haruta, Kazunori Matsumoto, Zhi Li +2
In this paper, we propose a rotation-constrained compensation method to address the errors introduced by structured pruning of large language models (LLMs). LLMs are trained on mas…
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
Yanan Wang, Julio Vizcarra, Zhi Li +2
Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fi…
Top-down Activity Representation Learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Capturing complex hierarchical human activities, from atomic actions (e.g., picking up one present, moving to the sofa, unwrapping the present) to contextual events (e.g., celebrat…
Multi-object event graph representation learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Video question answering (VideoQA) is a task to predict the correct answer to questions posed about a given video. The system must comprehend spatial and temporal relationships amo…
QWalkVec: Node Embedding by Quantum Walk
Rei Sato, Shuichiro Haruta, Kazuhiro Saito +1
In this paper, we propose QWalkVec, a quantum walk-based node embedding method. A quantum walk is a quantum version of a random walk that demonstrates a faster propagation than a r…