HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios

Evaluating the capabilities of large language models (LLMs) in human-LLM interactions remains challenging due to the inherent complexity and openness of dialogue processes. This paper introduces HammerBench, a novel benchmarking framework designed to assess the function-calling ability of LLMs more...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	arXiv.org 2024-12
Hauptverfasser:	Wang, Jun, Zhou, Jiamu, Wen, Muning, Mo, Xiaoyun, Zhang, Haoyu, Lin, Qiqiang, Cheng, Jin, Wang, Xihuai, Zhang, Weinan, Peng, Qiuying
Format:	Artikel
Sprache:	eng
Schlagworte:	Electronic devices Large language models Performance evaluation
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Schreiben Sie den ersten Kommentar!