Pir\'a: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

CIKM '21: Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 2021 Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as C...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Paschoal, André F. A, Pirozelli, Paulo, Freire, Valdinei, Delgado, Karina V, Peres, Sarajane M, José, Marcos M, Nakasato, Flávio, Oliveira, André S, Brandão, Anarosa A. F, Costa, Anna H. R, Cozman, Fabio G
Format: Artikel
Sprache:eng
Schlagworte:
Online-Zugang:Volltext bestellen
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
container_end_page
container_issue
container_start_page
container_title
container_volume
creator Paschoal, André F. A
Pirozelli, Paulo
Freire, Valdinei
Delgado, Karina V
Peres, Sarajane M
José, Marcos M
Nakasato, Flávio
Oliveira, André S
Brandão, Anarosa A. F
Costa, Anna H. R
Cozman, Fabio G
description CIKM '21: Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 2021 Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pir\'a dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pir\'a is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pir\'a dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pir\'a, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.
doi_str_mv 10.48550/arxiv.2202.02398
format Article
fullrecord <record><control><sourceid>arxiv_GOX</sourceid><recordid>TN_cdi_arxiv_primary_2202_02398</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2202_02398</sourcerecordid><originalsourceid>FETCH-arxiv_primary_2202_023983</originalsourceid><addsrcrecordid>eNpjYJA0NNAzsTA1NdBPLKrILNMzMjIw0jMwMra04GQIDcgsilFPtFJwVHDKzMnMSy9NzFEIyC8qKU0vTS1O1XXNS8_JLM5QcEksSSxOLVFIyy9SCATKlGTm5-k65hWXpxYBNSkkJuWXliiUZKQq-CenJubxMLCmJeYUp_JCaW4GeTfXEGcPXbAD4guKMnMTiyrjQQ6JBzvEmLAKAOmtPco</addsrcrecordid><sourcetype>Open Access Repository</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype></control><display><type>article</type><title>Pir\'a: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean</title><source>arXiv.org</source><creator>Paschoal, André F. A ; Pirozelli, Paulo ; Freire, Valdinei ; Delgado, Karina V ; Peres, Sarajane M ; José, Marcos M ; Nakasato, Flávio ; Oliveira, André S ; Brandão, Anarosa A. F ; Costa, Anna H. R ; Cozman, Fabio G</creator><creatorcontrib>Paschoal, André F. A ; Pirozelli, Paulo ; Freire, Valdinei ; Delgado, Karina V ; Peres, Sarajane M ; José, Marcos M ; Nakasato, Flávio ; Oliveira, André S ; Brandão, Anarosa A. F ; Costa, Anna H. R ; Cozman, Fabio G</creatorcontrib><description>CIKM '21: Proceedings of the 30th ACM International Conference on Information &amp; Knowledge Management, 2021 Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pir\'a dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pir\'a is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pir\'a dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pir\'a, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.</description><identifier>DOI: 10.48550/arxiv.2202.02398</identifier><language>eng</language><subject>Computer Science - Computation and Language</subject><creationdate>2022-02</creationdate><rights>http://creativecommons.org/licenses/by-nc-sa/4.0</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>228,230,776,881</link.rule.ids><linktorsrc>$$Uhttps://arxiv.org/abs/2202.02398$$EView_record_in_Cornell_University$$FView_record_in_$$GCornell_University$$Hfree_for_read</linktorsrc><backlink>$$Uhttps://doi.org/10.48550/arXiv.2202.02398$$DView paper in arXiv$$Hfree_for_read</backlink><backlink>$$Uhttps://doi.org/10.1145/3459637.3482012$$DView published paper (Access to full text may be restricted)$$Hfree_for_read</backlink></links><search><creatorcontrib>Paschoal, André F. A</creatorcontrib><creatorcontrib>Pirozelli, Paulo</creatorcontrib><creatorcontrib>Freire, Valdinei</creatorcontrib><creatorcontrib>Delgado, Karina V</creatorcontrib><creatorcontrib>Peres, Sarajane M</creatorcontrib><creatorcontrib>José, Marcos M</creatorcontrib><creatorcontrib>Nakasato, Flávio</creatorcontrib><creatorcontrib>Oliveira, André S</creatorcontrib><creatorcontrib>Brandão, Anarosa A. F</creatorcontrib><creatorcontrib>Costa, Anna H. R</creatorcontrib><creatorcontrib>Cozman, Fabio G</creatorcontrib><title>Pir\'a: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean</title><description>CIKM '21: Proceedings of the 30th ACM International Conference on Information &amp; Knowledge Management, 2021 Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pir\'a dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pir\'a is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pir\'a dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pir\'a, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.</description><subject>Computer Science - Computation and Language</subject><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2022</creationdate><recordtype>article</recordtype><sourceid>GOX</sourceid><recordid>eNpjYJA0NNAzsTA1NdBPLKrILNMzMjIw0jMwMra04GQIDcgsilFPtFJwVHDKzMnMSy9NzFEIyC8qKU0vTS1O1XXNS8_JLM5QcEksSSxOLVFIyy9SCATKlGTm5-k65hWXpxYBNSkkJuWXliiUZKQq-CenJubxMLCmJeYUp_JCaW4GeTfXEGcPXbAD4guKMnMTiyrjQQ6JBzvEmLAKAOmtPco</recordid><startdate>20220204</startdate><enddate>20220204</enddate><creator>Paschoal, André F. A</creator><creator>Pirozelli, Paulo</creator><creator>Freire, Valdinei</creator><creator>Delgado, Karina V</creator><creator>Peres, Sarajane M</creator><creator>José, Marcos M</creator><creator>Nakasato, Flávio</creator><creator>Oliveira, André S</creator><creator>Brandão, Anarosa A. F</creator><creator>Costa, Anna H. R</creator><creator>Cozman, Fabio G</creator><scope>AKY</scope><scope>GOX</scope></search><sort><creationdate>20220204</creationdate><title>Pir\'a: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean</title><author>Paschoal, André F. A ; Pirozelli, Paulo ; Freire, Valdinei ; Delgado, Karina V ; Peres, Sarajane M ; José, Marcos M ; Nakasato, Flávio ; Oliveira, André S ; Brandão, Anarosa A. F ; Costa, Anna H. R ; Cozman, Fabio G</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-arxiv_primary_2202_023983</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2022</creationdate><topic>Computer Science - Computation and Language</topic><toplevel>online_resources</toplevel><creatorcontrib>Paschoal, André F. A</creatorcontrib><creatorcontrib>Pirozelli, Paulo</creatorcontrib><creatorcontrib>Freire, Valdinei</creatorcontrib><creatorcontrib>Delgado, Karina V</creatorcontrib><creatorcontrib>Peres, Sarajane M</creatorcontrib><creatorcontrib>José, Marcos M</creatorcontrib><creatorcontrib>Nakasato, Flávio</creatorcontrib><creatorcontrib>Oliveira, André S</creatorcontrib><creatorcontrib>Brandão, Anarosa A. F</creatorcontrib><creatorcontrib>Costa, Anna H. R</creatorcontrib><creatorcontrib>Cozman, Fabio G</creatorcontrib><collection>arXiv Computer Science</collection><collection>arXiv.org</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext_linktorsrc</fulltext></delivery><addata><au>Paschoal, André F. A</au><au>Pirozelli, Paulo</au><au>Freire, Valdinei</au><au>Delgado, Karina V</au><au>Peres, Sarajane M</au><au>José, Marcos M</au><au>Nakasato, Flávio</au><au>Oliveira, André S</au><au>Brandão, Anarosa A. F</au><au>Costa, Anna H. R</au><au>Cozman, Fabio G</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>Pir\'a: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean</atitle><date>2022-02-04</date><risdate>2022</risdate><abstract>CIKM '21: Proceedings of the 30th ACM International Conference on Information &amp; Knowledge Management, 2021 Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pir\'a dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pir\'a is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pir\'a dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pir\'a, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.</abstract><doi>10.48550/arxiv.2202.02398</doi><oa>free_for_read</oa></addata></record>
fulltext fulltext_linktorsrc
identifier DOI: 10.48550/arxiv.2202.02398
ispartof
issn
language eng
recordid cdi_arxiv_primary_2202_02398
source arXiv.org
subjects Computer Science - Computation and Language
title Pir\'a: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean
url https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-28T21%3A33%3A15IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-arxiv_GOX&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=Pir%5C'a:%20A%20Bilingual%20Portuguese-English%20Dataset%20for%20Question-Answering%20about%20the%20Ocean&rft.au=Paschoal,%20Andr%C3%A9%20F.%20A&rft.date=2022-02-04&rft_id=info:doi/10.48550/arxiv.2202.02398&rft_dat=%3Carxiv_GOX%3E2202_02398%3C/arxiv_GOX%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_id=info:pmid/&rfr_iscdi=true