Comparative Analysis of Similarity-Based Edge Construction Methods for Village Welfare Index Networks
DOI:
https://doi.org/10.63158/journalisi.v8i4.1769Keywords:
Data Desa Presisi, Edge Construction, Household Similarity Network, Jaccard Index, Louvain ModularityAbstract
This study compares three similarity-based edge construction methods for village welfare networks from household welfare attributes in Data Desa Presisi. Households are represented as nodes, while edges denote computational similarity between household welfare profiles rather than social relationships. Five Village Welfare Index dimensions were discretized and transformed into binary one-hot representations. Pairwise similarities were calculated using Cosine Similarity, Jaccard Index, and Pearson Correlation, followed by automatic thresholding to retain strong relationships. Network evaluation focused on the Largest Connected Component, while community structure was assessed using Louvain modularity across 14 villages. All methods produced analyzable networks, achieving a 100% success rate and 97.17% average node coverage. Jaccard achieved the highest mean modularity (0.5074) and win rate (71.43%; 10 of 14 villages). A Friedman test confirmed differences among methods (χ² = 8.71, p = 0.013). Holm-corrected Wilcoxon tests showed that Jaccard significantly outperformed Cosine and Pearson, whereas Cosine and Pearson did not differ significantly. These findings indicate that attribute-overlap-based edge construction provides most consistent representation of household welfare-profile proximity under the tested binary representation and automatic-thresholding scheme, while emphasizing that this advantage is context-dependent rather than universal.
Downloads
References
[1] United Nations Development Programme (UNDP), Human Development Report 1990: Concept and Measurement of Human Development. New York, NY, USA: Oxford Univ. Press, 1990.
[2] S. Alkire, “Dimensions of human development,” World Dev., vol. 30, no. 2, pp. 181–205, 2002, doi: 10.1016/S0305-750X(01)00109-7.
[3] Badan Pusat Statistik, Indeks Pembangunan Desa 2018. Jakarta, Indonesia: Badan Pusat Statistik, 2019.
[4] S. Sjaf, Sampean, A. A. Arsyad, L. Elson, A. R. Mahardika, L. Hakim, S. A. Amongjati, R. Gandi, Z. A. Barlan, I. M. G. Aditya, S. A. B. Maulana, and M. R. Rangkuti, “Data Desa Presisi: A new method of rural data collection,” MethodsX, vol. 9, Art. no. 101868, 2022, doi: 10.1016/j.mex.2022.101868.
[5] N. Couldry and U. A. Mejias, The Costs of Connection: How Data Is Colonizing Human Life and Appropriating It for Capitalism. Stanford, CA, USA: Stanford Univ. Press, 2019, doi: 10.1515/9781503609754.
[6] N. Couldry and A. Powell, “Big Data from the bottom up,” Big Data Soc., vol. 1, no. 2, 2014, doi: 10.1177/2053951714539277.
[7] A. Majeed and I. Rauf, “Graph theory: A comprehensive survey about graph theory applications in computer science and social networks,” Inventions, vol. 5, no. 1, Art. no. 10, 2020, doi: 10.3390/inventions5010010.
[8] F. Cacheda, D. Fernandez, F. J. Novoa, and V. Carneiro, “Early detection of depression: Social network analysis and random forest techniques,” J. Med. Internet Res., vol. 21, no. 6, Art. no. e12554, 2019, doi: 10.2196/12554.
[9] T. Adali and A. Ortega, “Applications of graph theory [Scanning the Issue],” Proc. IEEE, vol. 106, no. 5, pp. 784–786, May 2018, doi: 10.1109/JPROC.2018.2820300.
[10] S. B. Guerrero-Ocampo and J. M. Díaz-Puente, “Social network analysis uses and contributions to innovation initiatives in rural areas: A review,” Sustainability, vol. 15, no. 18, Art. no. 14018, 2023, doi: 10.3390/su151814018.
[11] A. Levy, B. R. Shalom, and M. Chalamish, “A guide to similarity measures and their data science applications,” J. Big Data, vol. 12, no. 1, Art. no. 188, 2025, doi: 10.1186/s40537-025-01227-1.
[12] S. Ghosh, N. Das, T. Gonçalves, P. Quaresma, and M. Kundu, “The journey of graph kernels through two decades,” Comput. Sci. Rev., vol. 27, pp. 88–111, 2018, doi: 10.1016/j.cosrev.2017.11.002.
[13] V. D. Blondel, J.-L. Guillaume, R. Lambiotte, and E. Lefebvre, “Fast unfolding of communities in large networks,” J. Stat. Mech. Theory Exp., vol. 2008, no. 10, Art. no. P10008, 2008, doi: 10.1088/1742-5468/2008/10/P10008.
[14] Y. Lee, Y. Lee, J. C. Seong, A. Stanescu, and C. S. Hwang, “A comparison of network clustering algorithms in keyword network analysis: A case study with geography conference presentations,” Int. J. Geospat. Environ. Res., vol. 7, no. 3, pp. 1–16, 2020.
[15] M. E. J. Newman and M. Girvan, “Finding and evaluating community structure in networks,” Phys. Rev. E, vol. 69, no. 2, Art. no. 026113, 2004, doi: 10.1103/PhysRevE.69.026113.
[16] X. Cheng, C. Yang, Y. Zhao, Y. Wang, H. Karimi, and T. Derr, “A comprehensive analysis of social tie strength: Definitions, prediction methods, and future directions,” arXiv preprint arXiv:2410.19214, 2024, doi: 10.48550/arXiv.2410.19214.
[17] F. Montes, R. C. Jimenez, and J.-P. Onnela, “Connected but segregated: Social networks in rural villages,” J. Complex Netw., vol. 6, no. 5, pp. 693–705, 2018, doi: 10.1093/comnet/cnx054.
[18] P. Li, J. Yu, J. Liu, D. Zhou, and B. Cao, “Generating weighted social networks using multigraph,” Physica A, vol. 539, Art. no. 122894, 2020, doi: 10.1016/j.physa.2019.122894.
[19] M. McPherson, L. Smith-Lovin, and J. M. Cook, “Birds of a feather: Homophily in social networks,” Annu. Rev. Sociol., vol. 27, pp. 415–444, 2001, doi: 10.1146/annurev.soc.27.1.415.
[20] M. T. Rivera, S. B. Soderstrom, and B. Uzzi, “Dynamics of dyads in social networks: Assortative, relational, and proximity mechanisms,” Annu. Rev. Sociol., vol. 36, pp. 91–115, 2010, doi: 10.1146/annurev.soc.34.040507.134743.
[21] M. Suyudi, Sukono, and A. T. Bon, “Theoretical and simulation approaches centrality in social networks,” in Proc. 11th Annu. Int. Conf. Ind. Eng. Oper. Manag. (IEOM), Singapore, Mar. 7–11, 2021, pp. 4093–4101, doi: 10.46254/AN11.20210734.
[22] M. E. J. Newman, “Assortative mixing in networks,” Phys. Rev. Lett., vol. 89, no. 20, Art. no. 208701, 2002, doi: 10.1103/PhysRevLett.89.208701.
[23] U. von Luxburg, “A tutorial on spectral clustering,” Stat. Comput., vol. 17, no. 4, pp. 395–416, 2007, doi: 10.1007/s11222-007-9033-z.
[24] M. Maier, M. Hein, and U. von Luxburg, “Optimal construction of k-nearest-neighbor graphs for identifying noisy clusters,” Theor. Comput. Sci., vol. 410, no. 19, pp. 1749–1764, 2009, doi: 10.1016/j.tcs.2009.01.009.
[25] M. Á. Serrano, M. Boguñá, and A. Vespignani, “Extracting the multiscale backbone of complex weighted networks,” Proc. Natl. Acad. Sci. U.S.A., vol. 106, no. 16, pp. 6483–6488, 2009, doi: 10.1073/pnas.0808904106.
[26] J. Olivier, W. D. Johnson, and G. D. Marshall, “The logarithmic transformation and the geometric mean in reporting experimental IgE results: What are they and when and why to use them?,” Ann. Allergy Asthma Immunol., vol. 100, no. 4, pp. 333–337, 2008, doi: 10.1016/S1081-1206(10)60595-9.
[27] A. Mahmoudi, M. R. Yaakub, and A. Abu Bakar, “A new method to discretize time to identify the milestones of online social networks,” Soc. Netw. Anal. Min., vol. 8, no. 1, Art. no. 34, pp. 1–20, 2018, doi: 10.1007/s13278-018-0511-4.
[28] M. Fernández, H. Kirchner, and B. Pinaud, “Labelled port graph—A formal structure for models and computations,” Electron. Notes Theor. Comput. Sci., vol. 338, pp. 3–21, 2018, doi: 10.1016/j.entcs.2018.10.002.
[29] G. Salton, A. Wong, and C. S. Yang, “A vector space model for automatic indexing,” Commun. ACM, vol. 18, no. 11, pp. 613–620, 1975, doi: 10.1145/361219.361220.
[30] P. Jaccard, “The distribution of the flora in the alpine zone,” New Phytol., vol. 11, no. 2, pp. 37–50, 1912, doi: 10.1111/j.1469-8137.1912.tb05611.x.
[31] K. Pearson, “Note on regression and inheritance in the case of two parents,” Proc. R. Soc. Lond., vol. 58, pp. 240–242, 1895, doi: 10.1098/rspl.1895.0041.
[32] V. A. Traag, L. Waltman, and N. J. van Eck, “From Louvain to Leiden: Guaranteeing well-connected communities,” Sci. Rep., vol. 9, Art. no. 5233, 2019, doi: 10.1038/s41598-019-41695-z.
[33] S. Fortunato and M. Barthélemy, “Resolution limit in community detection,” Proc. Natl. Acad. Sci. U.S.A., vol. 104, no. 1, pp. 36–41, 2007, doi: 10.1073/pnas.0605965104.
[34] F. Radicchi, J. J. Ramasco, and S. Fortunato, “Information filtering in complex weighted networks,” Phys. Rev. E, vol. 83, no. 4, Art. no. 046101, 2011, doi: 10.1103/PhysRevE.83.046101.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information Systems and Informatics

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors Declaration
- The Authors certify that they have read, understood, and agreed to the Journal of Information Systems and Informatics (JournalISI) submission guidelines, policies, and submission declaration. The submission has been prepared using the provided template.
- The Authors certify that all authors have approved the publication of this manuscript and that there is no conflict of interest.
- The Authors confirm that the manuscript is their original work, has not received prior publication, is not under consideration for publication elsewhere, and has not been previously published.
- The Authors confirm that all authors listed on the title page have contributed significantly to the work, have read the manuscript, attest to the validity and legitimacy of the data and its interpretation, and agree to its submission.
- The Authors confirm that the manuscript is not copied from or plagiarized from any other published work.
- The Authors declare that the manuscript will not be submitted for publication in any other journal or magazine until a decision is made by the journal editors.
- If the manuscript is finally accepted for publication, the Authors confirm that they will either proceed with publication immediately or withdraw the manuscript in accordance with the journal’s withdrawal policies.
- The Authors agree that, upon publication of the manuscript in this journal, they transfer copyright or assign exclusive rights to the publisher, including commercial rights














