Regularized optimal transport of covariates and outcomes in data recoding
Autor: | Valérie Garès, Jérémy Omer |
---|---|
Přispěvatelé: | Institut de Recherche Mathématique de Rennes (IRMAR), AGROCAMPUS OUEST, Institut national d'enseignement supérieur pour l'agriculture, l'alimentation et l'environnement (Institut Agro)-Institut national d'enseignement supérieur pour l'agriculture, l'alimentation et l'environnement (Institut Agro)-Université de Rennes 1 (UR1), Université de Rennes (UNIV-RENNES)-Université de Rennes (UNIV-RENNES)-Université de Rennes 2 (UR2), Université de Rennes (UNIV-RENNES)-École normale supérieure - Rennes (ENS Rennes)-Centre National de la Recherche Scientifique (CNRS)-Institut National des Sciences Appliquées - Rennes (INSA Rennes), Institut National des Sciences Appliquées (INSA)-Université de Rennes (UNIV-RENNES)-Institut National des Sciences Appliquées (INSA), IRMAR-STAT, Université de Rennes (UR)-Institut National des Sciences Appliquées - Rennes (INSA Rennes), Institut National des Sciences Appliquées (INSA)-Institut National des Sciences Appliquées (INSA)-École normale supérieure - Rennes (ENS Rennes)-Université de Rennes 2 (UR2)-Centre National de la Recherche Scientifique (CNRS)-INSTITUT AGRO Agrocampus Ouest, Institut national d'enseignement supérieur pour l'agriculture, l'alimentation et l'environnement (Institut Agro)-Institut national d'enseignement supérieur pour l'agriculture, l'alimentation et l'environnement (Institut Agro) |
Jazyk: | angličtina |
Rok vydání: | 2020 |
Předmět: |
Statistics and Probability
Domain adaptation Computer science Epidemiology Variable recoding 01 natural sciences Outcome (game theory) 010104 statistics & probability Outcome variable [MATH.MATH-ST]Mathematics [math]/Statistics [math.ST] 0502 economics and business Statistics Covariate 0101 mathematics 050205 econometrics Linkage (software) Linkage 05 social sciences Merging databases [INFO.INFO-RO]Computer Science [cs]/Operations Research [cs.RO] Statistical matching Optimal Transportation Heterogeneous sources [MATH.MATH-OC]Mathematics [math]/Optimization and Control [math.OC] Statistics Probability and Uncertainty [STAT.ME]Statistics [stat]/Methodology [stat.ME] |
Zdroj: | Journal of the American Statistical Association Journal of the American Statistical Association, Taylor & Francis, In press, pp.1-14. ⟨10.1080/01621459.2020.1775615⟩ Journal of the American Statistical Association, In press, pp.1-14. ⟨10.1080/01621459.2020.1775615⟩ |
ISSN: | 0162-1459 1537-274X |
Popis: | International audience; When databases are constructed from heterogeneous sources, it is not unusual that different encodings are used for the same outcome. In such case, it is necessary to recode the outcome variable before merging two databases. The method proposed for the recoding is an application of optimal transportation where we search for a bijective mapping between the distributions of such variable in two databases. In this article, we build upon the work by Garès et al. [9], where they transport the distributions of categorical outcomes assuming that they are distributed equally in the two databases. Here, we extend the scope of the model to treat all the situations where the covariates explain the outcomes similarly in the two databases. In particular, we do not require that the outcomes be distributed equally. For this, we propose a model where joint distributions of outcomes and covariates are transported. We also propose to enrich the model by relaxing the constraints on marginal distributions and adding an L1 regularization term. The performances of the models are evaluated in a simulation study, and they are applied to a real dataset. |
Databáze: | OpenAIRE |
Externí odkaz: |