TY - JOUR
T1 - Measuring linkage disequilibrium and improvement of pruning and clumping in structured populations
AU - Bercovich, Ulises
AU - Rasmussen, Malthe Sebro
AU - Li, Zilong
AU - Wiuf, Carsten
AU - Albrechtsen, Anders
N1 - © The Author(s) 2025. Published by Oxford University Press on behalf of The Genetics Society of America. All rights reserved. For commercial re-use, please contact [email protected] for reprints and translation rights for reprints. All other permissions can be obtained through our RightsLink service via the Permissions link on the article page on our site—for further information please contact [email protected].
PY - 2025/2/5
Y1 - 2025/2/5
N2 - Standard measures of linkage disequilibrium (LD) are affected by admixture and population structure, such that loci that are not in LD within each ancestral population appear linked when considered jointly across the populations. The influence of population structure on LD can cause problems for downstream analysis methods, in particular those that rely on LD pruning or clumping. To address this issue, we propose a measure of LD that accommodates population structure using the top inferred principal components. We estimate LD from the correlation of genotype residuals and prove that this LD measure remains unaffected by population structure when analyzing multiple populations jointly, even with admixed individuals. Based on this adjusted measure of LD, we can perform LD pruning to remove the correlation between markers for downstream analysis. Traditional LD pruning is more likely to remove markers with high differences in allele frequencies between populations, which biases measures for genetic differentiation and removes markers that are not in LD in the ancestral populations. Using data from moderately differentiated human populations and highly differentiated giraffe populations we show that traditional LD pruning biases FST and principal component analysis (PCA), which can be alleviated with the adjusted LD measure. In addition, we show that the adjusted LD leads to better PCA when pruning and that LD clumping retains more sites with the retained sites having stronger associations.
AB - Standard measures of linkage disequilibrium (LD) are affected by admixture and population structure, such that loci that are not in LD within each ancestral population appear linked when considered jointly across the populations. The influence of population structure on LD can cause problems for downstream analysis methods, in particular those that rely on LD pruning or clumping. To address this issue, we propose a measure of LD that accommodates population structure using the top inferred principal components. We estimate LD from the correlation of genotype residuals and prove that this LD measure remains unaffected by population structure when analyzing multiple populations jointly, even with admixed individuals. Based on this adjusted measure of LD, we can perform LD pruning to remove the correlation between markers for downstream analysis. Traditional LD pruning is more likely to remove markers with high differences in allele frequencies between populations, which biases measures for genetic differentiation and removes markers that are not in LD in the ancestral populations. Using data from moderately differentiated human populations and highly differentiated giraffe populations we show that traditional LD pruning biases FST and principal component analysis (PCA), which can be alleviated with the adjusted LD measure. In addition, we show that the adjusted LD leads to better PCA when pruning and that LD clumping retains more sites with the retained sites having stronger associations.
U2 - 10.1093/genetics/iyaf009
DO - 10.1093/genetics/iyaf009
M3 - Journal article
C2 - 39907701
SN - 1943-2631
VL - 229
JO - Genetics
JF - Genetics
IS - 3
M1 - iyaf009
ER -