Publication: A multivariate approach to the symmetrical uncertainty measure: Application to feature selection problem
| dc.contributor.author | Sosa Cabrera, Gustavo | |
| dc.contributor.author | García Torres, Miguel | |
| dc.contributor.author | Gómez Guerrero, Santiago | |
| dc.contributor.author | E. Schaerer, Christian | |
| dc.contributor.author | Divina, Federico | |
| dc.date.accessioned | 2024-02-01T09:31:48Z | |
| dc.date.available | 2024-02-01T09:31:48Z | |
| dc.date.issued | 2019 | |
| dc.description.abstract | In this work we propose an extension of the Symmetrical Uncertainty (SU) measure in order to address the multivariate case, simultaneously acquiring the capability to detect possible correlations and interactions among features. This generalization, denoted Multivariate Symmetrical Uncertainty (MSU), is based on the concepts of Total Correlation (TC) and Mutual Information (MI) extended to the multivariate case. The generalized measure accounts for the total amount of dependency within a set of variables as a single monolithic quantity. Multivariate measures are usually biased due to several factors. To overcome this problem, a mathematical expression is proposed, based on the cardinality of all features, which can be used to calculate the number of samples needed to estimate the MSU without bias at a pre-specified significance level. Theoretical and experimental results on synthetic data show that the proposed sample size expression properly controls the bias. In addition, when the MSU is applied to feature selection on synthetic and real-world data, it has the advantage of adequately capturing linear and nonlinear correlations and interactions, and it can therefore be used as a new feature subset evaluation method. | |
| dc.description.sponsorship | Deporte e Informática | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.citation | Information Sciences, vol. 494, p. 1-20 | |
| dc.identifier.doi | 10.1016/j.ins.2019.04.046 | |
| dc.identifier.uri | https://hdl.handle.net/10433/19571 | |
| dc.language.iso | en | |
| dc.publisher | Elsevier | |
| dc.relation.projectID | info:eu-repo/grantAgreement/MINECO//TIN2015-64776-C3-2-R/ES/DIFFERENTIAL@UPO: MASSIVE DATA MANAGEMENT, FILTERING AND EXPLORATORY ANALYSIS/ | |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | en |
| dc.rights.accessRights | open access | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject | Multivariate symmetrical uncertainty | |
| dc.subject | Mutual information | |
| dc.subject | Entropy | |
| dc.subject | Feature selection | |
| dc.title | A multivariate approach to the symmetrical uncertainty measure: Application to feature selection problem | |
| dc.type | journal article | |
| dc.type.hasVersion | AM | |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | 4ce19614-9553-49b0-9b6e-09817f551658 | |
| relation.isAuthorOfPublication | 82e2c456-c4b8-494e-b3d9-f6c84c8cf9a5 | |
| relation.isAuthorOfPublication.latestForDiscovery | 4ce19614-9553-49b0-9b6e-09817f551658 |
Files
Original bundle
1 - 1 of 1

