9 papers · 1 filter
PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning
Bo Su, Ankit Shah, Thai Le
Machine unlearning for large language models (LLMs) aims to remove specified knowledge while preserving the rest of the model's capabilities. However, the boundary between knowledg…
Inference Time Optimization with Confidence Dynamics
Yu Wang, Minghao Liu, Jiayun Wang +3
Inference time optimization techniques, such as repeated sampling, have significantly advanced the reasoning capabilities of Large Language Models (LLMs). However, the critical rol…
Training-Free Agentic AI: Probabilistic Control and Coordination in Multi-Agent LLM Systems
Mohammad Parsa Hosseini, Ankit Shah, Saiyra Qureshi +3
Multi-agent large language model (LLM) systems enable complex, long-horizon reasoning by composing specialized agents, but practical deployment remains hindered by inefficient rout…
ProRefine: Inference-Time Prompt Refinement with Textual Feedback
Deepak Pandita, Tharindu Cyril Weerasooriya, Ankit Parag Shah +3
Agentic workflows, where multiple AI agents collaborate to accomplish complex tasks like reasoning or planning, play a substantial role in many cutting-edge commercial applications…
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Zhenting Wang, Qi Chang, Hemani Patel +8
We introduce MCP-Bench, a benchmark for evaluating large language models (LLMs) on realistic, multi-step tasks that demand tool use, cross-tool coordination, precise parameter cont…
Enhancing Retrieval for ESGLLM via ESG-CID -- A Disclosure Content Index Finetuning Dataset for Mapping GRI and ESRS
Shafiuddin Rehan Ahmed, Ankit Parag Shah, Quan Hung Tran +5
Climate change has intensified the need for transparency and accountability in organizational practices, making Environmental, Social, and Governance (ESG) reporting increasingly c…