Media Cloud: Massive Open Source Collection of Global News on the Open Web
We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data M...
Gespeichert in:
Hauptverfasser: | , , , , , , , , , , , , , , , , , |
---|---|
Format: | Artikel |
Sprache: | eng |
Schlagworte: | |
Online-Zugang: | Volltext bestellen |
Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Zusammenfassung: | We present the first full description of Media Cloud, an open source platform
based on crawling hyperlink structure in operation for over 10 years, that for
many uses will be the best way to collect data for studying the media ecosystem
on the open web. We document the key choices behind what data Media Cloud
collects and stores, how it processes and organizes these data, and its open
API access as well as user-facing tools. We also highlight the strengths and
limitations of the Media Cloud collection strategy compared to relevant
alternatives. We give an overview two sample datasets generated using Media
Cloud and discuss how researchers can use the platform to create their own
datasets. |
---|---|
DOI: | 10.48550/arxiv.2104.03702 |