papers

Publications (17)

cs.CL2024

Enhancing Temporal Sensitivity and Reasoning for Time-Sensitive Question Answering

Wanqi Yang, Yanda Li, Meng Fang +1

Time-Sensitive Question Answering (TSQA) demands the effective utilization of specific temporal contexts, encompassing multiple time-evolving facts, to address time-sensitive quest…

cs.SD2026

Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models

Yanda Li, Yuhan Liu, Zirui Song +3

Large audio-language models (LALMs) generalize across speech, sound, and music, but unified decoders can exhibit a \emph{temporal smoothing bias}: transient acoustic cues may be un…

cs.AI2025

MMAC-Copilot: Multi-modal Agent Collaboration Operating Copilot

Zirui Song, Yaohang Li, Meng Fang +6

Large language model agents that interact with PC applications often face limitations due to their singular mode of interaction with real-world environments, leading to restricted…

cs.IR2026

PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation

Kirill Dubovikov, Omar El Mansouri, Hachem Madmoun +12

Petroleum-engineering search exposes a supervision gap for strong general retrievers: relevant evidence exists in public web text, but domain relevance labels are scarce. To addres…

cs.CV2023

StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

Yanda Li, Chi Zhang, Gang Yu +6

The remarkable multimodal capabilities demonstrated by OpenAI's GPT-4 have sparked significant interest in the development of multimodal Large Language Models (LLMs). A primary res…

cs.CR2026

Glitch in the Sky: Exploiting Voltage Fault Injection in UAV Flight Controllers

Yun-Ping Hsiao, Yanda Li, Youssef Gamal +2

As Cyber-Physical Systems (CPS) become increasingly pervasive and autonomous, ensuring the resilience of their embedded logic is critical to maintaining safety and integrity. Among…

cs.SD2025

Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models

Wanqi Yang, Yanda Li, Meng Fang +2

Adversarial audio attacks pose a significant threat to the growing use of large audio-language models (LALMs) in voice-based human-machine interactions. While existing research foc…

cs.CL2024

Adaptive Reinforcement Learning Planning: Harnessing Large Language Models for Complex Information Extraction

Zepeng Ding, Ruiyang Ke, Wenhao Huang +4

Existing research on large language models (LLMs) shows that they can solve information extraction tasks through multi-step planning. However, their extraction behavior on complex…

cs.CV2026

AppAgent: Multimodal Agents as Smartphone Users

Chi Zhang, Zhao Yang, Jiaxuan Liu +6

Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper introduces a novel LLM-based mult…

cs.CL2025

Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization

Dixuan Wang, Yanda Li, Junyuan Jiang +5

Large Language Models (LLMs) have shown remarkable capabilities in language understanding and generation. Nonetheless, it was also witnessed that LLMs tend to produce inaccurate re…

cs.CV2023

Disentangled Pre-training for Image Matting

Yanda Li, Zilong Huang, Gang Yu +3

Image matting requires high-quality pixel-level human annotations to support the training of a deep model in recent literature. Whereas such annotation is costly and hard to scale,…

cs.HC2025

AppAgent v2: Advanced Agent for Flexible Mobile Interactions

Yanda Li, Chi Zhang, Wenjia Jiang +6

With the advancement of Multimodal Large Language Models (MLLM), LLM-driven visual agents are increasingly impacting software interfaces, particularly those with graphical user int…

cs.CL2024

Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding

Yanda Li, Dixuan Wang, Jiaqing Liang +4

Large Language Models (LLMs) have demonstrated good performance in many reasoning tasks, but they still struggle with some complicated reasoning tasks including logical reasoning.…

cs.CL2025

MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents

Wanqi Yang, Yanda Li, Meng Fang +1

Understanding temporal dynamics is critical for conversational agents, enabling effective content analysis and informed decision-making. However, time-aware datasets, particularly…

cs.AI2025

Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Biao Wu, Yanda Li, Zhiwei Zhang +3

Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and pro…

cs.CL2025

SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models

Wanqi Yang, Yanda Li, Yunchao Wei +2

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surf…

cs.CL2024

Continual Learning for Temporal-Sensitive Question Answering

Wanqi Yang, Yunqiu Xu, Yanda Li +3

In this study, we explore an emerging research area of Continual Learning for Temporal Sensitive Question Answering (CLTSQA). Previous research has primarily focused on Temporal Se…