MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants

LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lacks an objective and automatic evaluation on their memory capability, largely due to the challenges in constructing reliab...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Hauptverfasser:	Zhang, Zeyu, Dai, Quanyu, Chen, Luyu, Jiang, Zeren, Li, Rui, Zhu, Jieming, Chen, Xu, Xie, Yi, Dong, Zhenhua, Wen, Ji-Rong
Format:	Artikel
Sprache:	eng
Schlagworte:	Computer Science - Artificial Intelligence Computer Science - Computation and Language
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

container_end_page
container_issue
container_start_page
container_title
container_volume
creator	Zhang, Zeyu Dai, Quanyu Chen, Luyu Jiang, Zeren Li, Rui Zhu, Jieming Chen, Xu Xie, Yi Dong, Zhenhua Wen, Ji-Rong
description	LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lacks an objective and automatic evaluation on their memory capability, largely due to the challenges in constructing reliable questions and answers (QAs) according to user messages. In this paper, we propose MemSim, a Bayesian simulator designed to automatically construct reliable QAs from generated user messages, simultaneously keeping their diversity and scalability. Specifically, we introduce the Bayesian Relation Network (BRNet) and a causal generation mechanism to mitigate the impact of LLM hallucinations on factual information, facilitating the automatic creation of an evaluation dataset. Based on MemSim, we generate a dataset in the daily-life scenario, named MemDaily, and conduct extensive experiments to assess the effectiveness of our approach. We also provide a benchmark for evaluating different memory mechanisms in LLM-based agents with the MemDaily dataset. To benefit the research community, we have released our project at https://github.com/nuster1128/MemSim.
doi_str_mv	10.48550/arxiv.2409.20163
format	Article
fullrecord	<record><control><sourceid>arxiv_GOX</sourceid><recordid>TN_cdi_arxiv_primary_2409_20163</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2409_20163</sourcerecordid><originalsourceid>FETCH-arxiv_primary_2409_201633</originalsourceid><addsrcrecordid>eNpjYJA0NNAzsTA1NdBPLKrILNMzMjGw1DMyMDQz5mQI9k3NDc7MtVJwVHBKrEwtzkzMUwDyS3MSS_KLFNKA2LUsMac0sSQzL10BqDa_qFIhP03Bx8dXNymxODVFISC1qDg_LzFHwbG4OLO4JDGvpJiHgTUtMac4lRdKczPIu7mGOHvogq2PLyjKzE0sqowHOSMe7AxjwioAUEI87g</addsrcrecordid><sourcetype>Open Access Repository</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype></control><display><type>article</type><title>MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants</title><source>arXiv.org</source><creator>Zhang, Zeyu ; Dai, Quanyu ; Chen, Luyu ; Jiang, Zeren ; Li, Rui ; Zhu, Jieming ; Chen, Xu ; Xie, Yi ; Dong, Zhenhua ; Wen, Ji-Rong</creator><creatorcontrib>Zhang, Zeyu ; Dai, Quanyu ; Chen, Luyu ; Jiang, Zeren ; Li, Rui ; Zhu, Jieming ; Chen, Xu ; Xie, Yi ; Dong, Zhenhua ; Wen, Ji-Rong</creatorcontrib><description>LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lacks an objective and automatic evaluation on their memory capability, largely due to the challenges in constructing reliable questions and answers (QAs) according to user messages. In this paper, we propose MemSim, a Bayesian simulator designed to automatically construct reliable QAs from generated user messages, simultaneously keeping their diversity and scalability. Specifically, we introduce the Bayesian Relation Network (BRNet) and a causal generation mechanism to mitigate the impact of LLM hallucinations on factual information, facilitating the automatic creation of an evaluation dataset. Based on MemSim, we generate a dataset in the daily-life scenario, named MemDaily, and conduct extensive experiments to assess the effectiveness of our approach. We also provide a benchmark for evaluating different memory mechanisms in LLM-based agents with the MemDaily dataset. To benefit the research community, we have released our project at https://github.com/nuster1128/MemSim.</description><identifier>DOI: 10.48550/arxiv.2409.20163</identifier><language>eng</language><subject>Computer Science - Artificial Intelligence ; Computer Science - Computation and Language</subject><creationdate>2024-09</creationdate><rights>http://creativecommons.org/licenses/by/4.0</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>228,230,776,881</link.rule.ids><linktorsrc>$$Uhttps://arxiv.org/abs/2409.20163$$EView_record_in_Cornell_University$$FView_record_in_$$GCornell_University$$Hfree_for_read</linktorsrc><backlink>$$Uhttps://doi.org/10.48550/arXiv.2409.20163$$DView paper in arXiv$$Hfree_for_read</backlink></links><search><creatorcontrib>Zhang, Zeyu</creatorcontrib><creatorcontrib>Dai, Quanyu</creatorcontrib><creatorcontrib>Chen, Luyu</creatorcontrib><creatorcontrib>Jiang, Zeren</creatorcontrib><creatorcontrib>Li, Rui</creatorcontrib><creatorcontrib>Zhu, Jieming</creatorcontrib><creatorcontrib>Chen, Xu</creatorcontrib><creatorcontrib>Xie, Yi</creatorcontrib><creatorcontrib>Dong, Zhenhua</creatorcontrib><creatorcontrib>Wen, Ji-Rong</creatorcontrib><title>MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants</title><description>LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lacks an objective and automatic evaluation on their memory capability, largely due to the challenges in constructing reliable questions and answers (QAs) according to user messages. In this paper, we propose MemSim, a Bayesian simulator designed to automatically construct reliable QAs from generated user messages, simultaneously keeping their diversity and scalability. Specifically, we introduce the Bayesian Relation Network (BRNet) and a causal generation mechanism to mitigate the impact of LLM hallucinations on factual information, facilitating the automatic creation of an evaluation dataset. Based on MemSim, we generate a dataset in the daily-life scenario, named MemDaily, and conduct extensive experiments to assess the effectiveness of our approach. We also provide a benchmark for evaluating different memory mechanisms in LLM-based agents with the MemDaily dataset. To benefit the research community, we have released our project at https://github.com/nuster1128/MemSim.</description><subject>Computer Science - Artificial Intelligence</subject><subject>Computer Science - Computation and Language</subject><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2024</creationdate><recordtype>article</recordtype><sourceid>GOX</sourceid><recordid>eNpjYJA0NNAzsTA1NdBPLKrILNMzMjGw1DMyMDQz5mQI9k3NDc7MtVJwVHBKrEwtzkzMUwDyS3MSS_KLFNKA2LUsMac0sSQzL10BqDa_qFIhP03Bx8dXNymxODVFISC1qDg_LzFHwbG4OLO4JDGvpJiHgTUtMac4lRdKczPIu7mGOHvogq2PLyjKzE0sqowHOSMe7AxjwioAUEI87g</recordid><startdate>20240930</startdate><enddate>20240930</enddate><creator>Zhang, Zeyu</creator><creator>Dai, Quanyu</creator><creator>Chen, Luyu</creator><creator>Jiang, Zeren</creator><creator>Li, Rui</creator><creator>Zhu, Jieming</creator><creator>Chen, Xu</creator><creator>Xie, Yi</creator><creator>Dong, Zhenhua</creator><creator>Wen, Ji-Rong</creator><scope>AKY</scope><scope>GOX</scope></search><sort><creationdate>20240930</creationdate><title>MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants</title><author>Zhang, Zeyu ; Dai, Quanyu ; Chen, Luyu ; Jiang, Zeren ; Li, Rui ; Zhu, Jieming ; Chen, Xu ; Xie, Yi ; Dong, Zhenhua ; Wen, Ji-Rong</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-arxiv_primary_2409_201633</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2024</creationdate><topic>Computer Science - Artificial Intelligence</topic><topic>Computer Science - Computation and Language</topic><toplevel>online_resources</toplevel><creatorcontrib>Zhang, Zeyu</creatorcontrib><creatorcontrib>Dai, Quanyu</creatorcontrib><creatorcontrib>Chen, Luyu</creatorcontrib><creatorcontrib>Jiang, Zeren</creatorcontrib><creatorcontrib>Li, Rui</creatorcontrib><creatorcontrib>Zhu, Jieming</creatorcontrib><creatorcontrib>Chen, Xu</creatorcontrib><creatorcontrib>Xie, Yi</creatorcontrib><creatorcontrib>Dong, Zhenhua</creatorcontrib><creatorcontrib>Wen, Ji-Rong</creatorcontrib><collection>arXiv Computer Science</collection><collection>arXiv.org</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext_linktorsrc</fulltext></delivery><addata><au>Zhang, Zeyu</au><au>Dai, Quanyu</au><au>Chen, Luyu</au><au>Jiang, Zeren</au><au>Li, Rui</au><au>Zhu, Jieming</au><au>Chen, Xu</au><au>Xie, Yi</au><au>Dong, Zhenhua</au><au>Wen, Ji-Rong</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants</atitle><date>2024-09-30</date><risdate>2024</risdate><abstract>LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lacks an objective and automatic evaluation on their memory capability, largely due to the challenges in constructing reliable questions and answers (QAs) according to user messages. In this paper, we propose MemSim, a Bayesian simulator designed to automatically construct reliable QAs from generated user messages, simultaneously keeping their diversity and scalability. Specifically, we introduce the Bayesian Relation Network (BRNet) and a causal generation mechanism to mitigate the impact of LLM hallucinations on factual information, facilitating the automatic creation of an evaluation dataset. Based on MemSim, we generate a dataset in the daily-life scenario, named MemDaily, and conduct extensive experiments to assess the effectiveness of our approach. We also provide a benchmark for evaluating different memory mechanisms in LLM-based agents with the MemDaily dataset. To benefit the research community, we have released our project at https://github.com/nuster1128/MemSim.</abstract><doi>10.48550/arxiv.2409.20163</doi><oa>free_for_read</oa></addata></record>
fulltext	fulltext_linktorsrc
identifier	DOI: 10.48550/arxiv.2409.20163
ispartof
issn
language	eng
recordid	cdi_arxiv_primary_2409_20163
source	arXiv.org
subjects	Computer Science - Artificial Intelligence Computer Science - Computation and Language
title	MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
url	https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-02-01T07%3A46%3A28IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-arxiv_GOX&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=MemSim:%20A%20Bayesian%20Simulator%20for%20Evaluating%20Memory%20of%20LLM-based%20Personal%20Assistants&rft.au=Zhang,%20Zeyu&rft.date=2024-09-30&rft_id=info:doi/10.48550/arxiv.2409.20163&rft_dat=%3Carxiv_GOX%3E2409_20163%3C/arxiv_GOX%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_id=info:pmid/&rfr_iscdi=true