Abstract
Abstract: Recruitment for clinical trials is undergoing a major transition, from trial centers concentrated in North America and Western Europe to a much greater emphasis on Eastern Europe, South and East Asia, and Latin America. This shift challenges traditional approaches to genetic association studies, in which studies have been designed around populations of relatively common ancestry. There is particular concern about the potential impact of genetic substructure, stratification and heterogeneity on the analysis of diverse samples. Although methods such as those implemented in STRUCTURE and EIGENSTRAT have been developed to assess and address the problems arising from the analysis of samples containing individuals from multiple populations, we have an insufficient understanding of how population-dependent prevalence, allele frequency, penetrance, and LD affect study type I error and power. Through simulation we explore the effects of heterogeneity in these population parameters on the analysis with a variety of statistical models. We find that type 1 error can be controlled when combining samples. Methods using combined samples outperform methods that rely on analyzing samples separately in ~70% of the conditions considered. The models with the highest overall power were those that assumed a main effect for population with no difference in genetic effect between populations. We conclude that the same information that has been used to exclude data to minimize diversity would be better employed as a statistical covariate in an inclusive analysis.