Publications (28)
CTS-PLL: A Robust and Anytime Framework for Collaborative Task Sequencing and Multi-Agent Path Finding
Junkai Jiang, Yitao Xu, Ruochen Li +2
The Collaborative Task Sequencing and Multi-Agent Path Finding (CTS-MAPF) problem requires agents to accomplish sequences of tasks while avoiding collisions, posing significant cha…
Motion-Adaptive Multi-Scale Temporal Modelling with Skeleton-Constrained Spatial Graphs for Efficient 3D Human Pose Estimation
Ruochen Li, Shuang Chen, Wenke E +2
Accurate 3D human pose estimation from monocular videos requires effective modelling of complex spatial and temporal dependencies. However, existing methods often face challenges i…
The Ramon Llull's Thinking Machine for Automated Ideation
Xinran Zhao, Boyuan Zheng, Chenglei Si +8
This paper revisits Ramon Llull's Ars combinatoria - a medieval framework for generating knowledge through symbolic recombination - as a conceptual foundation for building a modern…
AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control
Ruosen Li, Ziming Luo, Quan Zhang +4
Large reasoning models (LRMs) achieve impressive reasoning capabilities by generating lengthy chain-of-thoughts, but this "overthinking" incurs high latency and cost without commen…
ViTE: Virtual Graph Trajectory Expert Router for Pedestrian Trajectory Prediction
Ruochen Li, Zhanxing Zhu, Tanqiu Qiao +1
Pedestrian trajectory prediction is critical for ensuring safety in autonomous driving, surveillance systems, and urban planning applications. While early approaches primarily focu…
Geometric Visual Fusion Graph Neural Networks for Multi-Person Human-Object Interaction Recognition in Videos
Tanqiu Qiao, Ruochen Li, Frederick W. B. Li +3
Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features…
VRUD: A Drone Dataset for Complex Vehicle-VRU Interactions within Mixed Traffic
Ziyu Wang, Hongrui Kou, Cheng Wang +4
The Operational Design Domain (ODD) of urbanoriented Level 4 (L4) autonomous driving, especially for autonomous robotaxis, confronts formidable challenges in complex urban mixed tr…
Classification, Regression and Segmentation directly from k-Space in Cardiac MRI
Ruochen Li, Jiazhen Pan, Youxiang Zhu +2
Cardiac Magnetic Resonance Imaging (CMR) is the gold standard for diagnosing cardiovascular diseases. Clinical diagnoses predominantly rely on magnitude-only Digital Imaging and Co…
Beyond Scalar Scores: Exploring LLM-based Metrics for Clinical Significance Evaluation in Radiology Reports
Qingyu Lu, Ruochen Li, Liang Ding +3
Reliable evaluation of generated radiology reports requires strict clinical accuracy, as omitted critical findings or mischaracterized radiographic observations can directly affect…
LDC: Learning to Generate Research Idea with Dynamic Control
Ruochen Li, Liqiang Jing, Chi Han +2
Recent advancements in large language models (LLMs) have demonstrated their potential in automating the scientific research ideation. Existing approaches primarily focus on prompti…
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
Ruosen Li, Ruochen Li, Barry Wang +1
To evaluate Large Language Models (LLMs) for question answering (QA), traditional methods typically focus on assessing single-turn responses to given questions. However, this appro…
CTS-CBS: A New Approach for Multi-Agent Collaborative Task Sequencing and Path Finding
Junkai Jiang, Ruochen Li, Yibin Yang +4
This paper addresses a generalization problem of Multi-Agent Pathfinding (MAPF), called Collaborative Task Sequencing - Multi-Agent Pathfinding (CTS-MAPF), where agents must plan c…
Degradation-Aware Model Predictive Control for Battery Swapping Stations under Energy Arbitrage
Ruochen Li, Zhichao Chen, Zhaoting Zhang +4
Battery swapping stations (BSS) offer a fast and scalable alternative to conventional electric vehicle (EV) charging, gaining growing policy support worldwide. However, existing BS…
Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts
Youxiang Zhu, Ruochen Li, Danqing Wang +2
Long-context large language models (LLMs) are prone to be distracted by irrelevant contexts. The reason for distraction remains poorly understood. In this paper, we first identify…
Multiclass-SGCN: Sparse Graph-based Trajectory Prediction with Agent Class Embedding
Ruochen Li, Stamos Katsigiannis, Hubert P. H. Shum
Trajectory prediction of road users in real-world scenarios is challenging because their movement patterns are stochastic and complex. Previous pedestrian-oriented works have been…
UIESNN: A Scale-Aware Spiking Network for Underwater Image Enhancement
Shuang Chen, Ruochen Li, Zihan Zhu +3
Underwater image enhancement (UIE) is a practically important yet underexplored application of spiking neural networks (SNNs), where the dominant degradations are large-scale and l…
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
Shuo Yan, Ruochen Li, Ziming Luo +11
Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reprodu…
BP-SGCN: Behavioral Pseudo-Label Informed Sparse Graph Convolution Network for Pedestrian and Heterogeneous Trajectory Prediction
Ruochen Li, Stamos Katsigiannis, Tae-Kyun Kim +1
Trajectory prediction allows better decision-making in applications of autonomous vehicles or surveillance by predicting the short-term future movement of traffic agents. It is cla…
ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment
Ruochen Li, Jun Li, Bailiang Jian +2
Automatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians' trust. This gap reveals fundamental flaws in how c…
SurgGoal: Rethinking Surgical Planning Evaluation via Goal-Satisfiability
Ruochen Li, Kun Yuan, Yufei Xia +5
Surgical planning integrates visual perception, long-horizon reasoning, and procedural knowledge, yet it remains unclear whether current evaluation protocols reliably assess vision…
Boundary Guided Semantic Learning for Real-time COVID-19 Lung Infection Segmentation System
Runmin Cong, Yumo Zhang, Ning Yang +6
The coronavirus disease 2019 (COVID-19) continues to have a negative impact on healthcare systems around the world, though the vaccines have been developed and national vaccination…
ART: Adaptive Relational Transformer for Pedestrian Trajectory Prediction with Temporal-Aware Relations
Ruochen Li, Ziyi Chang, Junyan Hu +3
Accurate prediction of real-world pedestrian trajectories is crucial for a wide range of robot-related applications. Recent approaches typically adopt graph-based or transformer-ba…
MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents
Ruochen Li, Teerth Patel, Qingyun Wang +1
Autonomous machine learning research has gained significant attention recently. We present MLR-COPILOT, an autonomous Machine Learning Research framework powered by large language…
Unified Spatial-Temporal Edge-Enhanced Graph Networks for Pedestrian Trajectory Prediction
Ruochen Li, Tanqiu Qiao, Stamos Katsigiannis +2
Pedestrian trajectory prediction aims to forecast future movements based on historical paths. Spatial-temporal (ST) methods often separately model spatial interactions among pedest…
From Category to Scenery: An End-to-End Framework for Multi-Person Human-Object Interaction Recognition in Videos
Tanqiu Qiao, Ruochen Li, Frederick W. B. Li +1
Video-based Human-Object Interaction (HOI) recognition explores the intricate dynamics between humans and objects, which are essential for a comprehensive understanding of human be…
ESCoT: An Enhanced Step-based Coordinate Trajectory Planning Method for Multiple Car-like Robots
Junkai Jiang, Yihe Chen, Yibin Yang +3
Multi-vehicle trajectory planning (MVTP) is one of the key challenges in multi-robot systems (MRSs) and has broad applications across various fields. This paper presents ESCoT, an…
TimeFlow: Temporal Conditioning for Longitudinal Brain MRI Registration and Aging Analysis
Bailiang Jian, Jiazhen Pan, Yitong Li +5
Longitudinal brain analysis is essential for understanding healthy aging and identifying pathological deviations. Longitudinal registration of sequential brain MRI underpins such a…
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
Omer Faruk Deniz, Ruiyu Mao, Ruochen Li +2
Multimodal Large Language Models (MLLMs) incur significant computational cost from processing numerous vision tokens through all LLM layers. Prior pruning methods operate either be…