From the 2 of 15 linked papers with an AI index.
15 papers
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
Yunxin Li, Jinchao Li, Shibo Su +7
The paper introduces KnowAct-GUIClaw, an agent framework that combines accumulated user interaction knowledge with a self‑evolving memory and skill library to improve cross‑platfor…
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
Zhenyu Liu, Xuanyu Zhang, Yunxin Li +10
The paper identifies gradient conflicts between acoustic and semantic modeling as the cause of modality interference in full-duplex spoken language models and proposes Lychee-FD, a…
Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations
Xinxin Li, Huiyao Chen, Meishan Zhang +6
Automatic speech recognition (ASR) correction has traditionally focused on isolated utterances or short local contexts. However, as text and speech become increasingly interleaved…
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
Jinchao Li, Yunxin Li, Chenrui Zhao +3
While GUI agents have shown impressive capabilities in common computer-use tasks such as OSWorld, current benchmarks mainly focus on isolated and single-application tasks. This ove…
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
Haoyuan Shi, Yunxin Li, Nanhao Deng +5
The evolution of video generation toward complex, multi-shot narratives has exposed a critical deficit in current evaluation methods. Existing benchmarks remain anchored to single-…
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
Yunxin Li, Zhenyu Liu, Baotian Hu +4
Recent advancements in multimodal large language models (MLLMs) have achieved significant multimodal generation capabilities, akin to GPT-4. These models predominantly map visual i…