Network compression with configuration models and the minimum description length
Creators
- 1. Vermont Complex Systems Center, University of Vermont, Burlington, Vermont 05405, USA
- 2. Department of Computer Science, University of Vermont, Burlington, Vermont 05405, USA
- 3. Department of Mathematics & Statistics, University of Vermont, Burlington, Vermont 05405, USA
- 4. Département de physique, de génie physique et d'optique, Université Laval, Québec (Québec), Canada G1V 0A6
- 5. Institute of Data Science, University of Hong Kong, Hong Kong
- 6. Department of Urban Planning and Design, University of Hong Kong, Hong Kong
- 7. Urban Systems Institute, University of Hong Kong, Hong Kong
- 8. Centre interdisciplinaire en modélisation mathématique, Université Laval, Québec (Québec), Canada G1V 0A6
Description
Random network models, constrained to reproduce specific statistical features, are often used to represent and analyze network data and their mathematical descriptions. Chief among them, the configuration model constrains random networks by their degree distribution and is foundational to many areas of network science. However, configuration models and their variants are often selected based on intuition or mathematical and computational simplicity rather than on statistical evidence. To evaluate the quality of a network representation, we need to consider both the amount of information required to specify a random network model and the probability of recovering the original data when using the model as a generative process. To this end, we calculate the approximate size of network ensembles generated by the popular configuration model and its generalizations, including versions accounting for degree correlations and centrality layers. We then apply the minimum description length principle as a model selection criterion over the resulting nested family of configuration models. Using a dataset of over 100 networks from various domains, we find that the classic configuration model is generally preferred on networks with an average degree above 10, while a layered configuration model constrained by a centrality metric offers the most compact representation of the majority of sparse networks.
Additional details
Identifiers
- DOI
- 10.1103/PhysRevE.110.034305;
- Crossref Funder ID
- 10.13039/100000001; 10.13039/501100000038;
Publishing Information
- Journal Title
- Physical Review E
- Journal Volume
- 110
- Journal Issue
- 3
- Journal Page Range
- 11 pgs.
- ISSN
- 1089-3787
INIS
- Country of Publication
- United States
- Country of Input or Organization
- International Atomic Energy Agency (IAEA)
- Subject category
- S97: MATHEMATICAL METHODS AND COMPUTING; S71: CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSICS;
- Descriptors DEI
- APPROXIMATIONS; COMPRESSION; CONFIGURATION; CORRELATIONS; DATA-FLOW PROCESSING; DYNAMICAL SYSTEMS; LAYERS; LENGTH; MATHEMATICAL MODELS; METRICS; PROBABILITY; PROBABILITY DENSITY FUNCTIONS; RANDOMNESS; STATISTICAL MECHANICS; STATISTICAL MODELS; STATISTICS
- Descriptors DEC
- CALCULATION METHODS; DIMENSIONS; FUNCTIONS; MATHEMATICAL MODELS; MATHEMATICS; MECHANICS; PROGRAMMING
Optional Information
- Copyright
- ©2024 American Physical Society
- Contract/Grant/Project number
- DMS-1829826; 2019-05183
- Notes
- Record automatically processed
- Funding organization
- National Science Foundation; Natural Sciences and Engineering Research Council of Canada