Published September 6, 2024 | Version v1
Journal article

Network compression with configuration models and the minimum description length

  • 1. Vermont Complex Systems Center, University of Vermont, Burlington, Vermont 05405, USA
  • 2. Department of Computer Science, University of Vermont, Burlington, Vermont 05405, USA
  • 3. Department of Mathematics & Statistics, University of Vermont, Burlington, Vermont 05405, USA
  • 4. Département de physique, de génie physique et d'optique, Université Laval, Québec (Québec), Canada G1V 0A6
  • 5. Institute of Data Science, University of Hong Kong, Hong Kong
  • 6. Department of Urban Planning and Design, University of Hong Kong, Hong Kong
  • 7. Urban Systems Institute, University of Hong Kong, Hong Kong
  • 8. Centre interdisciplinaire en modélisation mathématique, Université Laval, Québec (Québec), Canada G1V 0A6

Description

Random network models, constrained to reproduce specific statistical features, are often used to represent and analyze network data and their mathematical descriptions. Chief among them, the configuration model constrains random networks by their degree distribution and is foundational to many areas of network science. However, configuration models and their variants are often selected based on intuition or mathematical and computational simplicity rather than on statistical evidence. To evaluate the quality of a network representation, we need to consider both the amount of information required to specify a random network model and the probability of recovering the original data when using the model as a generative process. To this end, we calculate the approximate size of network ensembles generated by the popular configuration model and its generalizations, including versions accounting for degree correlations and centrality layers. We then apply the minimum description length principle as a model selection criterion over the resulting nested family of configuration models. Using a dataset of over 100 networks from various domains, we find that the classic configuration model is generally preferred on networks with an average degree above 10, while a layered configuration model constrained by a centrality metric offers the most compact representation of the majority of sparse networks.

Additional details

Identifiers

DOI
10.1103/PhysRevE.110.034305;
Crossref Funder ID
10.13039/100000001; 10.13039/501100000038;

Publishing Information

Journal Title
Physical Review E
Journal Volume
110
Journal Issue
3
Journal Page Range
11 pgs.
ISSN
1089-3787

Optional Information

Copyright
©2024 American Physical Society
Contract/Grant/Project number
DMS-1829826; 2019-05183
Notes
Record automatically processed
Funding organization
National Science Foundation; Natural Sciences and Engineering Research Council of Canada