Structural effects of network sampling coverage I: Nodes missing at random
•We examine the effect of missing data on measures of centrality, topology and homophily.•We examine measurement bias and variability across 12 empirical networks.•Measurement bias varies systematically across network measures and network types.•Larger, more centralized networks are generally more r...
Gespeichert in:
Veröffentlicht in: | Social networks 2013-10, Vol.35 (4), p.652-668 |
---|---|
Hauptverfasser: | , |
Format: | Artikel |
Sprache: | eng |
Schlagworte: | |
Online-Zugang: | Volltext |
Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Zusammenfassung: | •We examine the effect of missing data on measures of centrality, topology and homophily.•We examine measurement bias and variability across 12 empirical networks.•Measurement bias varies systematically across network measures and network types.•Larger, more centralized networks are generally more robust to missing data.•Cohesive networks are robust to missing data when measuring topology but not centralization.
Network measures assume a census of a well-bounded population. This level of coverage is rarely achieved in practice, however, and we have only limited information on the robustness of network measures to incomplete coverage. This paper examines the effect of node-level missingness on 4 classes of network measures: centrality, centralization, topology and homophily across a diverse sample of 12 empirical networks. We use a Monte Carlo simulation process to generate data with known levels of missingness and compare the resulting network scores to their known starting values. As with past studies (Borgatti et al., 2006; Kossinets, 2006), we find that measurement bias generally increases with more missing data. The exact rate and nature of this increase, however, varies systematically across network measures. For example, betweenness and Bonacich centralization are quite sensitive to missing data while closeness and in-degree are robust. Similarly, while the tau statistic and distance are difficult to capture with missing data, transitivity shows little bias even with very high levels of missingness. The results are also clearly dependent on the features of the network. Larger, more centralized networks are generally more robust to missing data, but this is especially true for centrality and centralization measures. More cohesive networks are robust to missing data when measuring topological features but not when measuring centralization. Overall, the results suggest that missing data may have quite large or quite small effects on network measurement, depending on the type of network and the question being posed. |
---|---|
ISSN: | 0378-8733 1879-2111 |
DOI: | 10.1016/j.socnet.2013.09.003 |