Media Cloud: Massive Open Source Collection of Global News on the Open Web

We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data M...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:arXiv.org 2021-05
Hauptverfasser: Roberts, Hal, Bhargava, Rahul, Valiukas, Linas, Jen, Dennis, Malik, Momin M, Bishop, Cindy, Ndulue, Emily, Aashka Dave, Clark, Justin, Etling, Bruce, Faris, Rob, Shah, Anushka, Rubinovitz, Jasmin, Hope, Alexis, D'Ignazio, Catherine, Bermejo, Fernando, Benkler, Yochai, Zuckerman, Ethan
Format: Artikel
Sprache:eng
Online-Zugang:Volltext
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
container_end_page
container_issue
container_start_page
container_title arXiv.org
container_volume
creator Roberts, Hal
Bhargava, Rahul
Valiukas, Linas
Jen, Dennis
Malik, Momin M
Bishop, Cindy
Ndulue, Emily
Aashka Dave
Clark, Justin
Etling, Bruce
Faris, Rob
Shah, Anushka
Rubinovitz, Jasmin
Hope, Alexis
D'Ignazio, Catherine
Bermejo, Fernando
Benkler, Yochai
Zuckerman, Ethan
description We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access as well as user-facing tools. We also highlight the strengths and limitations of the Media Cloud collection strategy compared to relevant alternatives. We give an overview two sample datasets generated using Media Cloud and discuss how researchers can use the platform to create their own datasets.
format Article
fullrecord <record><control><sourceid>proquest</sourceid><recordid>TN_cdi_proquest_journals_2510490256</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2510490256</sourcerecordid><originalsourceid>FETCH-proquest_journals_25104902563</originalsourceid><addsrcrecordid>eNqNjM0KgkAURocgSMp3uNBaGGcc-9lKPwTWoqCljHolZXDMq_X6zcIHaPXBOYdvxjwhZRhsIyEWzCdqOOci3gilpMcuKZa1hsTYsdxDqonqD8KtwxbuduwLhMQag8VQ2xZsBSdjc23gil8CR4bXFD8xX7F5pQ2hP-2SrY-HR3IOut6-R6Qha9xj61QmVMijHRcqlv9VP_-wOxw</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>2510490256</pqid></control><display><type>article</type><title>Media Cloud: Massive Open Source Collection of Global News on the Open Web</title><source>Free E- Journals</source><creator>Roberts, Hal ; Bhargava, Rahul ; Valiukas, Linas ; Jen, Dennis ; Malik, Momin M ; Bishop, Cindy ; Ndulue, Emily ; Aashka Dave ; Clark, Justin ; Etling, Bruce ; Faris, Rob ; Shah, Anushka ; Rubinovitz, Jasmin ; Hope, Alexis ; D'Ignazio, Catherine ; Bermejo, Fernando ; Benkler, Yochai ; Zuckerman, Ethan</creator><creatorcontrib>Roberts, Hal ; Bhargava, Rahul ; Valiukas, Linas ; Jen, Dennis ; Malik, Momin M ; Bishop, Cindy ; Ndulue, Emily ; Aashka Dave ; Clark, Justin ; Etling, Bruce ; Faris, Rob ; Shah, Anushka ; Rubinovitz, Jasmin ; Hope, Alexis ; D'Ignazio, Catherine ; Bermejo, Fernando ; Benkler, Yochai ; Zuckerman, Ethan</creatorcontrib><description>We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access as well as user-facing tools. We also highlight the strengths and limitations of the Media Cloud collection strategy compared to relevant alternatives. We give an overview two sample datasets generated using Media Cloud and discuss how researchers can use the platform to create their own datasets.</description><identifier>EISSN: 2331-8422</identifier><language>eng</language><publisher>Ithaca: Cornell University Library, arXiv.org</publisher><ispartof>arXiv.org, 2021-05</ispartof><rights>2021. This work is published under http://creativecommons.org/licenses/by-nc-nd/4.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>780,784</link.rule.ids></links><search><creatorcontrib>Roberts, Hal</creatorcontrib><creatorcontrib>Bhargava, Rahul</creatorcontrib><creatorcontrib>Valiukas, Linas</creatorcontrib><creatorcontrib>Jen, Dennis</creatorcontrib><creatorcontrib>Malik, Momin M</creatorcontrib><creatorcontrib>Bishop, Cindy</creatorcontrib><creatorcontrib>Ndulue, Emily</creatorcontrib><creatorcontrib>Aashka Dave</creatorcontrib><creatorcontrib>Clark, Justin</creatorcontrib><creatorcontrib>Etling, Bruce</creatorcontrib><creatorcontrib>Faris, Rob</creatorcontrib><creatorcontrib>Shah, Anushka</creatorcontrib><creatorcontrib>Rubinovitz, Jasmin</creatorcontrib><creatorcontrib>Hope, Alexis</creatorcontrib><creatorcontrib>D'Ignazio, Catherine</creatorcontrib><creatorcontrib>Bermejo, Fernando</creatorcontrib><creatorcontrib>Benkler, Yochai</creatorcontrib><creatorcontrib>Zuckerman, Ethan</creatorcontrib><title>Media Cloud: Massive Open Source Collection of Global News on the Open Web</title><title>arXiv.org</title><description>We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access as well as user-facing tools. We also highlight the strengths and limitations of the Media Cloud collection strategy compared to relevant alternatives. We give an overview two sample datasets generated using Media Cloud and discuss how researchers can use the platform to create their own datasets.</description><issn>2331-8422</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2021</creationdate><recordtype>article</recordtype><sourceid>ABUWG</sourceid><sourceid>AFKRA</sourceid><sourceid>AZQEC</sourceid><sourceid>BENPR</sourceid><sourceid>CCPQU</sourceid><sourceid>DWQXO</sourceid><recordid>eNqNjM0KgkAURocgSMp3uNBaGGcc-9lKPwTWoqCljHolZXDMq_X6zcIHaPXBOYdvxjwhZRhsIyEWzCdqOOci3gilpMcuKZa1hsTYsdxDqonqD8KtwxbuduwLhMQag8VQ2xZsBSdjc23gil8CR4bXFD8xX7F5pQ2hP-2SrY-HR3IOut6-R6Qha9xj61QmVMijHRcqlv9VP_-wOxw</recordid><startdate>20210501</startdate><enddate>20210501</enddate><creator>Roberts, Hal</creator><creator>Bhargava, Rahul</creator><creator>Valiukas, Linas</creator><creator>Jen, Dennis</creator><creator>Malik, Momin M</creator><creator>Bishop, Cindy</creator><creator>Ndulue, Emily</creator><creator>Aashka Dave</creator><creator>Clark, Justin</creator><creator>Etling, Bruce</creator><creator>Faris, Rob</creator><creator>Shah, Anushka</creator><creator>Rubinovitz, Jasmin</creator><creator>Hope, Alexis</creator><creator>D'Ignazio, Catherine</creator><creator>Bermejo, Fernando</creator><creator>Benkler, Yochai</creator><creator>Zuckerman, Ethan</creator><general>Cornell University Library, arXiv.org</general><scope>8FE</scope><scope>8FG</scope><scope>ABJCF</scope><scope>ABUWG</scope><scope>AFKRA</scope><scope>AZQEC</scope><scope>BENPR</scope><scope>BGLVJ</scope><scope>CCPQU</scope><scope>DWQXO</scope><scope>HCIFZ</scope><scope>L6V</scope><scope>M7S</scope><scope>PIMPY</scope><scope>PQEST</scope><scope>PQQKQ</scope><scope>PQUKI</scope><scope>PRINS</scope><scope>PTHSS</scope></search><sort><creationdate>20210501</creationdate><title>Media Cloud: Massive Open Source Collection of Global News on the Open Web</title><author>Roberts, Hal ; Bhargava, Rahul ; Valiukas, Linas ; Jen, Dennis ; Malik, Momin M ; Bishop, Cindy ; Ndulue, Emily ; Aashka Dave ; Clark, Justin ; Etling, Bruce ; Faris, Rob ; Shah, Anushka ; Rubinovitz, Jasmin ; Hope, Alexis ; D'Ignazio, Catherine ; Bermejo, Fernando ; Benkler, Yochai ; Zuckerman, Ethan</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-proquest_journals_25104902563</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2021</creationdate><toplevel>online_resources</toplevel><creatorcontrib>Roberts, Hal</creatorcontrib><creatorcontrib>Bhargava, Rahul</creatorcontrib><creatorcontrib>Valiukas, Linas</creatorcontrib><creatorcontrib>Jen, Dennis</creatorcontrib><creatorcontrib>Malik, Momin M</creatorcontrib><creatorcontrib>Bishop, Cindy</creatorcontrib><creatorcontrib>Ndulue, Emily</creatorcontrib><creatorcontrib>Aashka Dave</creatorcontrib><creatorcontrib>Clark, Justin</creatorcontrib><creatorcontrib>Etling, Bruce</creatorcontrib><creatorcontrib>Faris, Rob</creatorcontrib><creatorcontrib>Shah, Anushka</creatorcontrib><creatorcontrib>Rubinovitz, Jasmin</creatorcontrib><creatorcontrib>Hope, Alexis</creatorcontrib><creatorcontrib>D'Ignazio, Catherine</creatorcontrib><creatorcontrib>Bermejo, Fernando</creatorcontrib><creatorcontrib>Benkler, Yochai</creatorcontrib><creatorcontrib>Zuckerman, Ethan</creatorcontrib><collection>ProQuest SciTech Collection</collection><collection>ProQuest Technology Collection</collection><collection>Materials Science &amp; Engineering Collection</collection><collection>ProQuest Central (Alumni Edition)</collection><collection>ProQuest Central UK/Ireland</collection><collection>ProQuest Central Essentials</collection><collection>ProQuest Central</collection><collection>Technology Collection</collection><collection>ProQuest One Community College</collection><collection>ProQuest Central Korea</collection><collection>SciTech Premium Collection</collection><collection>ProQuest Engineering Collection</collection><collection>Engineering Database</collection><collection>Publicly Available Content Database</collection><collection>ProQuest One Academic Eastern Edition (DO NOT USE)</collection><collection>ProQuest One Academic</collection><collection>ProQuest One Academic UKI Edition</collection><collection>ProQuest Central China</collection><collection>Engineering Collection</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Roberts, Hal</au><au>Bhargava, Rahul</au><au>Valiukas, Linas</au><au>Jen, Dennis</au><au>Malik, Momin M</au><au>Bishop, Cindy</au><au>Ndulue, Emily</au><au>Aashka Dave</au><au>Clark, Justin</au><au>Etling, Bruce</au><au>Faris, Rob</au><au>Shah, Anushka</au><au>Rubinovitz, Jasmin</au><au>Hope, Alexis</au><au>D'Ignazio, Catherine</au><au>Bermejo, Fernando</au><au>Benkler, Yochai</au><au>Zuckerman, Ethan</au><format>book</format><genre>document</genre><ristype>GEN</ristype><atitle>Media Cloud: Massive Open Source Collection of Global News on the Open Web</atitle><jtitle>arXiv.org</jtitle><date>2021-05-01</date><risdate>2021</risdate><eissn>2331-8422</eissn><abstract>We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access as well as user-facing tools. We also highlight the strengths and limitations of the Media Cloud collection strategy compared to relevant alternatives. We give an overview two sample datasets generated using Media Cloud and discuss how researchers can use the platform to create their own datasets.</abstract><cop>Ithaca</cop><pub>Cornell University Library, arXiv.org</pub><oa>free_for_read</oa></addata></record>
fulltext fulltext
identifier EISSN: 2331-8422
ispartof arXiv.org, 2021-05
issn 2331-8422
language eng
recordid cdi_proquest_journals_2510490256
source Free E- Journals
title Media Cloud: Massive Open Source Collection of Global News on the Open Web
url https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-15T03%3A48%3A45IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest&rft_val_fmt=info:ofi/fmt:kev:mtx:book&rft.genre=document&rft.atitle=Media%20Cloud:%20Massive%20Open%20Source%20Collection%20of%20Global%20News%20on%20the%20Open%20Web&rft.jtitle=arXiv.org&rft.au=Roberts,%20Hal&rft.date=2021-05-01&rft.eissn=2331-8422&rft_id=info:doi/&rft_dat=%3Cproquest%3E2510490256%3C/proquest%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_pqid=2510490256&rft_id=info:pmid/&rfr_iscdi=true