QCS: A system for querying, clustering and summarizing documents

Information retrieval systems consist of many complicated components. Research and development of such systems is often hampered by the difficulty in evaluating how each particular component would behave across multiple systems. We present a novel integrated information retrieval system—the Query, C...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Information processing & management 2007-11, Vol.43 (6), p.1588-1605
Hauptverfasser:	Dunlavy, Daniel M., O’Leary, Dianne P., Conroy, John M., Schlesinger, Judith D.
Format:	Artikel
Sprache:	eng
Schlagworte:	Automatic abstracting Clustering Document management Indexing Information retrieval Latent semantic indexing Sentence trimming Studies Summarization Text processing
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

container_end_page	1605
container_issue	6
container_start_page	1588
container_title	Information processing & management
container_volume	43
creator	Dunlavy, Daniel M. O’Leary, Dianne P. Conroy, John M. Schlesinger, Judith D.
description	Information retrieval systems consist of many complicated components. Research and development of such systems is often hampered by the difficulty in evaluating how each particular component would behave across multiple systems. We present a novel integrated information retrieval system—the Query, Cluster, Summarize (QCS) system—which is portable, modular, and permits experimentation with different instantiations of each of the constituent text analysis components. Most importantly, the combination of the three types of methods in the QCS design improves retrievals by providing users more focused information organized by topic. We demonstrate the improved performance by a series of experiments using standard test sets from the Document Understanding Conferences (DUC) as measured by the best known automatic metric for summarization system evaluation, ROUGE. Although the DUC data and evaluations were originally designed to test multidocument summarization, we developed a framework to extend it to the task of evaluation for each of the three components: query, clustering, and summarization. Under this framework, we then demonstrate that the QCS system (end-to-end) achieves performance as good as or better than the best summarization engines. Given a query, QCS retrieves relevant documents, separates the retrieved documents into topic clusters, and creates a single summary for each cluster. In the current implementation, Latent Semantic Indexing is used for retrieval, generalized spherical k-means is used for the document clustering, and a method coupling sentence “trimming” and a hidden Markov model, followed by a pivoted QR decomposition, is used to create a single extract summary for each cluster. The user interface is designed to provide access to detailed information in a compact and useful format. Our system demonstrates the feasibility of assembling an effective IR system from existing software libraries, the usefulness of the modularity of the design, and the value of this particular combination of modules.
doi_str_mv	10.1016/j.ipm.2007.01.003
format	Article
fullrecord	<record><control><sourceid>proquest_cross</sourceid><recordid>TN_cdi_proquest_miscellaneous_57695899</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><els_id>S0306457307000246</els_id><sourcerecordid>57695899</sourcerecordid><originalsourceid>FETCH-LOGICAL-c398t-d58fe61d8c59b71b00f5c7bf3d4e51652faf43456a631114a4ef9b20d4d4a6ea3</originalsourceid><addsrcrecordid>eNp9kEtLxDAUhYMoOD5-gLviwpWt9zZJH7pxGHzBgIi6DmkekjJtx6QVxl9vhnHlwtV98J3LuYeQM4QMAYurNnPrLssBygwwA6B7ZIZVSVNOS9wnM6BQpIyX9JAchdACAOOYz8jty-L1OpknYRNG0yV28MnnZPzG9R-XiVpNcetjn8heJ2HqOund93bWg5o604_hhBxYuQrm9Lcek_f7u7fFY7p8fnhazJeponU1pppX1hSoK8XrpsQGwHJVNpZqZjgWPLfSMsp4IQuKiEwyY-smB800k4WR9Jhc7O6u_RAdhlF0LiizWsneDFMQvCxqXtV1BM__gO0w-T56E1izOs8rWkYId5DyQwjeWLH2Lj63EQhiG6hoRQxUbAMVgCIGGjU3O42Jb34540VQzvTKaOeNGoUe3D_qH5PVfUw</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>194922837</pqid></control><display><type>article</type><title>QCS: A system for querying, clustering and summarizing documents</title><source>ScienceDirect Journals (5 years ago - present)</source><creator>Dunlavy, Daniel M. ; O’Leary, Dianne P. ; Conroy, John M. ; Schlesinger, Judith D.</creator><creatorcontrib>Dunlavy, Daniel M. ; O’Leary, Dianne P. ; Conroy, John M. ; Schlesinger, Judith D.</creatorcontrib><description>Information retrieval systems consist of many complicated components. Research and development of such systems is often hampered by the difficulty in evaluating how each particular component would behave across multiple systems. We present a novel integrated information retrieval system—the Query, Cluster, Summarize (QCS) system—which is portable, modular, and permits experimentation with different instantiations of each of the constituent text analysis components. Most importantly, the combination of the three types of methods in the QCS design improves retrievals by providing users more focused information organized by topic. We demonstrate the improved performance by a series of experiments using standard test sets from the Document Understanding Conferences (DUC) as measured by the best known automatic metric for summarization system evaluation, ROUGE. Although the DUC data and evaluations were originally designed to test multidocument summarization, we developed a framework to extend it to the task of evaluation for each of the three components: query, clustering, and summarization. Under this framework, we then demonstrate that the QCS system (end-to-end) achieves performance as good as or better than the best summarization engines. Given a query, QCS retrieves relevant documents, separates the retrieved documents into topic clusters, and creates a single summary for each cluster. In the current implementation, Latent Semantic Indexing is used for retrieval, generalized spherical k-means is used for the document clustering, and a method coupling sentence “trimming” and a hidden Markov model, followed by a pivoted QR decomposition, is used to create a single extract summary for each cluster. The user interface is designed to provide access to detailed information in a compact and useful format. Our system demonstrates the feasibility of assembling an effective IR system from existing software libraries, the usefulness of the modularity of the design, and the value of this particular combination of modules.</description><identifier>ISSN: 0306-4573</identifier><identifier>EISSN: 1873-5371</identifier><identifier>DOI: 10.1016/j.ipm.2007.01.003</identifier><identifier>CODEN: IPMADK</identifier><language>eng</language><publisher>Oxford: Elsevier Ltd</publisher><subject>Automatic abstracting ; Clustering ; Document management ; Indexing ; Information retrieval ; Latent semantic indexing ; Sentence trimming ; Studies ; Summarization ; Text processing</subject><ispartof>Information processing & management, 2007-11, Vol.43 (6), p.1588-1605</ispartof><rights>2007</rights><rights>Copyright Pergamon Press Inc. Nov 2007</rights><lds50>peer_reviewed</lds50><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed><citedby>FETCH-LOGICAL-c398t-d58fe61d8c59b71b00f5c7bf3d4e51652faf43456a631114a4ef9b20d4d4a6ea3</citedby><cites>FETCH-LOGICAL-c398t-d58fe61d8c59b71b00f5c7bf3d4e51652faf43456a631114a4ef9b20d4d4a6ea3</cites></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><linktohtml>$$Uhttps://dx.doi.org/10.1016/j.ipm.2007.01.003$$EHTML$$P50$$Gelsevier$$H</linktohtml><link.rule.ids>314,780,784,3550,27924,27925,45995</link.rule.ids></links><search><creatorcontrib>Dunlavy, Daniel M.</creatorcontrib><creatorcontrib>O’Leary, Dianne P.</creatorcontrib><creatorcontrib>Conroy, John M.</creatorcontrib><creatorcontrib>Schlesinger, Judith D.</creatorcontrib><title>QCS: A system for querying, clustering and summarizing documents</title><title>Information processing & management</title><description>Information retrieval systems consist of many complicated components. Research and development of such systems is often hampered by the difficulty in evaluating how each particular component would behave across multiple systems. We present a novel integrated information retrieval system—the Query, Cluster, Summarize (QCS) system—which is portable, modular, and permits experimentation with different instantiations of each of the constituent text analysis components. Most importantly, the combination of the three types of methods in the QCS design improves retrievals by providing users more focused information organized by topic. We demonstrate the improved performance by a series of experiments using standard test sets from the Document Understanding Conferences (DUC) as measured by the best known automatic metric for summarization system evaluation, ROUGE. Although the DUC data and evaluations were originally designed to test multidocument summarization, we developed a framework to extend it to the task of evaluation for each of the three components: query, clustering, and summarization. Under this framework, we then demonstrate that the QCS system (end-to-end) achieves performance as good as or better than the best summarization engines. Given a query, QCS retrieves relevant documents, separates the retrieved documents into topic clusters, and creates a single summary for each cluster. In the current implementation, Latent Semantic Indexing is used for retrieval, generalized spherical k-means is used for the document clustering, and a method coupling sentence “trimming” and a hidden Markov model, followed by a pivoted QR decomposition, is used to create a single extract summary for each cluster. The user interface is designed to provide access to detailed information in a compact and useful format. Our system demonstrates the feasibility of assembling an effective IR system from existing software libraries, the usefulness of the modularity of the design, and the value of this particular combination of modules.</description><subject>Automatic abstracting</subject><subject>Clustering</subject><subject>Document management</subject><subject>Indexing</subject><subject>Information retrieval</subject><subject>Latent semantic indexing</subject><subject>Sentence trimming</subject><subject>Studies</subject><subject>Summarization</subject><subject>Text processing</subject><issn>0306-4573</issn><issn>1873-5371</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2007</creationdate><recordtype>article</recordtype><recordid>eNp9kEtLxDAUhYMoOD5-gLviwpWt9zZJH7pxGHzBgIi6DmkekjJtx6QVxl9vhnHlwtV98J3LuYeQM4QMAYurNnPrLssBygwwA6B7ZIZVSVNOS9wnM6BQpIyX9JAchdACAOOYz8jty-L1OpknYRNG0yV28MnnZPzG9R-XiVpNcetjn8heJ2HqOund93bWg5o604_hhBxYuQrm9Lcek_f7u7fFY7p8fnhazJeponU1pppX1hSoK8XrpsQGwHJVNpZqZjgWPLfSMsp4IQuKiEwyY-smB800k4WR9Jhc7O6u_RAdhlF0LiizWsneDFMQvCxqXtV1BM__gO0w-T56E1izOs8rWkYId5DyQwjeWLH2Lj63EQhiG6hoRQxUbAMVgCIGGjU3O42Jb34540VQzvTKaOeNGoUe3D_qH5PVfUw</recordid><startdate>20071101</startdate><enddate>20071101</enddate><creator>Dunlavy, Daniel M.</creator><creator>O’Leary, Dianne P.</creator><creator>Conroy, John M.</creator><creator>Schlesinger, Judith D.</creator><general>Elsevier Ltd</general><general>Elsevier Science Ltd</general><scope>AAYXX</scope><scope>CITATION</scope><scope>E3H</scope><scope>F2A</scope></search><sort><creationdate>20071101</creationdate><title>QCS: A system for querying, clustering and summarizing documents</title><author>Dunlavy, Daniel M. ; O’Leary, Dianne P. ; Conroy, John M. ; Schlesinger, Judith D.</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-LOGICAL-c398t-d58fe61d8c59b71b00f5c7bf3d4e51652faf43456a631114a4ef9b20d4d4a6ea3</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2007</creationdate><topic>Automatic abstracting</topic><topic>Clustering</topic><topic>Document management</topic><topic>Indexing</topic><topic>Information retrieval</topic><topic>Latent semantic indexing</topic><topic>Sentence trimming</topic><topic>Studies</topic><topic>Summarization</topic><topic>Text processing</topic><toplevel>peer_reviewed</toplevel><toplevel>online_resources</toplevel><creatorcontrib>Dunlavy, Daniel M.</creatorcontrib><creatorcontrib>O’Leary, Dianne P.</creatorcontrib><creatorcontrib>Conroy, John M.</creatorcontrib><creatorcontrib>Schlesinger, Judith D.</creatorcontrib><collection>CrossRef</collection><collection>Library & Information Sciences Abstracts (LISA)</collection><collection>Library & Information Science Abstracts (LISA)</collection><jtitle>Information processing & management</jtitle></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Dunlavy, Daniel M.</au><au>O’Leary, Dianne P.</au><au>Conroy, John M.</au><au>Schlesinger, Judith D.</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>QCS: A system for querying, clustering and summarizing documents</atitle><jtitle>Information processing & management</jtitle><date>2007-11-01</date><risdate>2007</risdate><volume>43</volume><issue>6</issue><spage>1588</spage><epage>1605</epage><pages>1588-1605</pages><issn>0306-4573</issn><eissn>1873-5371</eissn><coden>IPMADK</coden><abstract>Information retrieval systems consist of many complicated components. Research and development of such systems is often hampered by the difficulty in evaluating how each particular component would behave across multiple systems. We present a novel integrated information retrieval system—the Query, Cluster, Summarize (QCS) system—which is portable, modular, and permits experimentation with different instantiations of each of the constituent text analysis components. Most importantly, the combination of the three types of methods in the QCS design improves retrievals by providing users more focused information organized by topic. We demonstrate the improved performance by a series of experiments using standard test sets from the Document Understanding Conferences (DUC) as measured by the best known automatic metric for summarization system evaluation, ROUGE. Although the DUC data and evaluations were originally designed to test multidocument summarization, we developed a framework to extend it to the task of evaluation for each of the three components: query, clustering, and summarization. Under this framework, we then demonstrate that the QCS system (end-to-end) achieves performance as good as or better than the best summarization engines. Given a query, QCS retrieves relevant documents, separates the retrieved documents into topic clusters, and creates a single summary for each cluster. In the current implementation, Latent Semantic Indexing is used for retrieval, generalized spherical k-means is used for the document clustering, and a method coupling sentence “trimming” and a hidden Markov model, followed by a pivoted QR decomposition, is used to create a single extract summary for each cluster. The user interface is designed to provide access to detailed information in a compact and useful format. Our system demonstrates the feasibility of assembling an effective IR system from existing software libraries, the usefulness of the modularity of the design, and the value of this particular combination of modules.</abstract><cop>Oxford</cop><pub>Elsevier Ltd</pub><doi>10.1016/j.ipm.2007.01.003</doi><tpages>18</tpages><oa>free_for_read</oa></addata></record>
fulltext	fulltext
identifier	ISSN: 0306-4573
ispartof	Information processing & management, 2007-11, Vol.43 (6), p.1588-1605
issn	0306-4573 1873-5371
language	eng
recordid	cdi_proquest_miscellaneous_57695899
source	ScienceDirect Journals (5 years ago - present)
subjects	Automatic abstracting Clustering Document management Indexing Information retrieval Latent semantic indexing Sentence trimming Studies Summarization Text processing
title	QCS: A system for querying, clustering and summarizing documents
url	https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2024-12-25T17%3A10%3A27IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest_cross&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=QCS:%20A%20system%20for%20querying,%20clustering%20and%20summarizing%20documents&rft.jtitle=Information%20processing%20&%20management&rft.au=Dunlavy,%20Daniel%20M.&rft.date=2007-11-01&rft.volume=43&rft.issue=6&rft.spage=1588&rft.epage=1605&rft.pages=1588-1605&rft.issn=0306-4573&rft.eissn=1873-5371&rft.coden=IPMADK&rft_id=info:doi/10.1016/j.ipm.2007.01.003&rft_dat=%3Cproquest_cross%3E57695899%3C/proquest_cross%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_pqid=194922837&rft_id=info:pmid/&rft_els_id=S0306457307000246&rfr_iscdi=true