TY - JOUR
T1 - Long-read, high-throughput sequencing of 10,000 human mitogenomes using the PacBio Sequel IIe
AU - Holland, Mitchell M.
AU - McElhoe, Jennifer A.
AU - Holland, Charity A.
AU - Khosa, Jasmeen K.
AU - Alcaide, Claudia Prieto
AU - Hannon, Daniel
AU - Brownfield, Erin D.
AU - Korber, Jade T.
AU - Foster, Miriam G.
AU - Cundey, Rachel T.
AU - Claessens, Floor
AU - Higginbotham, Jennifer
AU - Anderson, Elise
AU - Sturk-Andreaggi, Kimberly
AU - Marshall, Charla
AU - Parson, Walther
N1 - Publisher Copyright:
© 2026 Elsevier B.V. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2027
Y1 - 2027
N2 - Sequence analysis of the human mitochondrial genome (mitogenome) is of interest to the molecular anthropology, medical, and forensic communities. Quality mitogenome data is an essential component of haplotype search databases, serving as an important element of forensic investigations to ensure that weight estimates are reflective of accurate coincidental match probabilities. The European DNA Profiling group (EDNAP) Mitochondrial DNA Population database (EMPOP) is considered the gold standard for this purpose, serving as a reference database and a quality-control tool. The current study reports on the development of a sequencing pipeline for mitogenomes that is user friendly, robust, and cost effective for uploading mitogenome sequences to EMPOP. Whole blood or buffy coat samples were extracted using the Zymo Research Quick-DNA Miniprep Plus kit. Amplification of the mitogenome was performed using two overlapping long-range amplicons of approximately 8.5 kb. Batches of amplicons from 372 samples, plus eight DNA extraction reagent blanks and four amplification negative controls, were normalized and pooled using SequalPrep plates. A library of amplicons was prepared by ligation of SMRT bells (single molecule, real time adaptors), and prepared libraries were run on the PacBio Sequel IIe instrument using a high-fidelity (HiFi) approach. The total time for laboratory processing of 384 samples, prior to SMRT bell ligation, was up to 62 working hours. Total cost of reagents and supplies for all steps was approximately 20 U.S. dollars (USD) per sample. Including labor, the cost was approximately 30 USD. The success rate for 10,394 total samples tested was ∼98.2%, with only one of the two target amplicons failing to produce suitable sequence data. Therefore, on a per amplicon basis, the success rate was ∼99.2%. Concordance studies using two short-read sequencing methods confirmed the reliability of the long-read approach. The long-read pipeline can be easily adopted by laboratories and used in high-throughput studies involving quality biological samples to generate large mitogenome databases, including those for upload to EMPOP.
AB - Sequence analysis of the human mitochondrial genome (mitogenome) is of interest to the molecular anthropology, medical, and forensic communities. Quality mitogenome data is an essential component of haplotype search databases, serving as an important element of forensic investigations to ensure that weight estimates are reflective of accurate coincidental match probabilities. The European DNA Profiling group (EDNAP) Mitochondrial DNA Population database (EMPOP) is considered the gold standard for this purpose, serving as a reference database and a quality-control tool. The current study reports on the development of a sequencing pipeline for mitogenomes that is user friendly, robust, and cost effective for uploading mitogenome sequences to EMPOP. Whole blood or buffy coat samples were extracted using the Zymo Research Quick-DNA Miniprep Plus kit. Amplification of the mitogenome was performed using two overlapping long-range amplicons of approximately 8.5 kb. Batches of amplicons from 372 samples, plus eight DNA extraction reagent blanks and four amplification negative controls, were normalized and pooled using SequalPrep plates. A library of amplicons was prepared by ligation of SMRT bells (single molecule, real time adaptors), and prepared libraries were run on the PacBio Sequel IIe instrument using a high-fidelity (HiFi) approach. The total time for laboratory processing of 384 samples, prior to SMRT bell ligation, was up to 62 working hours. Total cost of reagents and supplies for all steps was approximately 20 U.S. dollars (USD) per sample. Including labor, the cost was approximately 30 USD. The success rate for 10,394 total samples tested was ∼98.2%, with only one of the two target amplicons failing to produce suitable sequence data. Therefore, on a per amplicon basis, the success rate was ∼99.2%. Concordance studies using two short-read sequencing methods confirmed the reliability of the long-read approach. The long-read pipeline can be easily adopted by laboratories and used in high-throughput studies involving quality biological samples to generate large mitogenome databases, including those for upload to EMPOP.
KW - Bioinformatics
KW - Databasing
KW - EMPOP
KW - Forensic
KW - Massively parallel sequencing
KW - MtDNA
KW - PacBio
U2 - 10.1016/j.fsigen.2026.103588
DO - 10.1016/j.fsigen.2026.103588
M3 - Journal article
C2 - 42570395
AN - SCOPUS:105046700015
SN - 1872-4973
VL - 87
JO - Forensic Science International: Genetics
JF - Forensic Science International: Genetics
M1 - 103588
ER -