My work connects paralogous genes, copy-number variation and transposable elements: finding which copies are present, what they carry, and how they evolved.
Resolving genes that look alike
MRJD / TruPath Genome
Reconstructing gene copies and haplotypes where paralogs make short reads ambiguous. I designed and evaluated algorithms that combine joint analysis with proximity evidence.
Known pathogenic findings recovered in 31 evaluated samples across eight loci; an evaluation set, not diagnostic sensitivity.
Application note Slides Abstract
Method, contribution & evaluation
- Problem
- Genes such as SMN1, PMS2 and CYP21A2 share 98–99.9% of their sequence with a neighbouring paralog or pseudogene, so short reads map ambiguously and labs fall back on cascades of MLPA and long-range PCR.
- Method
- Multi-region joint detection (MRJD) analyses a gene and its paralogs together instead of one region at a time: it estimates total copy number, reconstructs each individual copy, and — with TruPath Genome — uses on-flowcell proximity linkage to support copy assignment.
- My role
- Algorithm design and evaluation. First-named inventor on the multi-region joint detection patent application; presented the TruPath evaluation as a platform talk at ACMG 2026 and wrote the Illumina application note.
- Evidence
- Across 31 samples spanning eight paralogous loci (18–357 kb, 98.2–99.9% homology), the evaluation recovered the expected pathogenic finding in every sample, agreeing with orthogonal methods including long-read sequencing. Research use only; this is recovery of known findings in an evaluation set, not a measure of diagnostic sensitivity.
Short reads can fit a gene and its paralog equally well. The highly similar subregions of pairs such as SMN1/SMN2 and PMS2/PMS2CL leave few distinguishing bases to anchor them. Explore that ambiguity below.
Explore the mapping problem
Why short reads become ambiguous
Choose a gene pair, then change the read length. Where copies look alike, simulated reads have fewer clues to their origin.
Illustrative simulation; not an aligner benchmark.
30× per copy · error-free reads · 450 bp mean fragment size
Read-pair results use all … simulated pairs across both copies. Callability refers to the displayed interval, not the whole gene.
Assumptions & reference data
Two different kinds of input. The difference map is derived from the reference: each preset comes from aligning that gene pair in the hg38 reference (NCBI sequences). The reads are entirely simulated. Nothing here is real sequencing data.
The difference map is binned, not exact. The presets store the number of differing positions per 500 bp bin, not the positions themselves. The demo re-draws positions at random within each bin, so bin-scale structure — the long identical stretches, the dense clusters — reflects the reference alignment, but individual tick positions do not correspond to specific coordinates. This is an approximation, not per-site realism.
Interval identity is not the same as the identity you see quoted for a gene. Each preset reports the average identity across the whole interval it displays. Published figures often describe the high-homology subregion instead, which is shorter and more similar. The PMS2 preset shows 95.3% across its ~19.3 kb interval, while the commonly quoted ~98% refers to the shorter high-homology region PMS2 shares with PMS2CL; averaging across the full interval pulls the number down because it includes diverged sequence. The two numbers describe different spans and neither has been adjusted to match the other.
The reads are simulated. Error-free paired-end mates at 30× per copy, 450 bp mean fragment size jittered per pair. A pair is anchored by any differing position either mate covers, plus any bases spilling into the unique sequence flanking the duplication — which is why the edges usually survive when the interior does not.
The MAPQ transformation is conventional; the probability model behind it is simplified. MAPQ = −10 log10 P(wrong copy) is the standard definition, but this demo feeds it a deliberately simple model: a pair covering no distinguishing position fits both copies equally well and is placed at random with MAPQ 0, and otherwise the probability assumes an independent 1% chance per site that a sequencing error masks a true difference, capped at 60. A real aligner also models base qualities, indels, mismatches elsewhere in the read, and its own scoring heuristics. "Confidently placed" here means MAPQ ≥ 20; "callable" means at least 8 such reads cover the position, a typical variant-calling filter.
Denominators. Both percentages above are taken over every simulated pair from both copies; the "assigned to the other copy" figure is the subset of those pairs placed on the copy they did not come from. The pileup draws only pairs assigned to the first copy, and only a sample of them — the row cap keeps the panel readable — while every reported number uses all simulated pairs.
From conventional MRJD to proximity-enabled MRJD
Conventional MRJD analyses paralogous regions together, retaining reads that map ambiguously. In high-sensitivity mode it can detect an allele even when its paralog assignment remains unresolved. TruPath Genome adds physical linkage evidence, supporting reconstruction and assignment of individual copies and haplotypes.
The additional evidence
A short read can borrow a longer view
Conceptual illustration of physical linkage, not an MRJD run. The same seven reads appear in each step.
-
01 Sequence alone
ABCDEFUUninformativeA and D carry distinguishing sequence. The other reads could fit either copy.
-
02 Physical linkage
ABCDEFUNo linkageLinkage connects ambiguous reads to informative sequence on the same molecule.
-
03 Copy assignment
ABCCopy 1DEFCopy 2UUnresolvedThe linked groups support separate copies. Without an anchor, U stays unresolved.
Detection, assignment & release support
Conventional MRJD, available in DRAGEN from v4.3, analyses a gene and its paralogs jointly and keeps the ambiguously mapping reads a conventional caller discards. On PMS2 it reports recall of approximately 99.7% for SNVs and 97.1% for indels against a long-range PCR truth set. Those figures measure detection — whether the allele was found at all. They are not a measure of assignment: in high-sensitivity mode an allele can be detected while its placement stays ambiguous, and the Illumina article describes exactly this, with such alleles reported in both PMS2 and PMS2CL. Calls from that workflow are not phased. Read the conventional MRJD article.
Proximity-enabled MRJD adds evidence the short reads never carried. TruPath Genome (formerly Constellation) is an Illumina on-flowcell chemistry — I work on the analysis side, not the sequencing technology — that captures and tagments DNA in spatially proximal nanowells, so reads descended from the same original molecule stay physically linked. MRJD uses that linkage to estimate total copy number, reconstruct each individual copy, and then assign the reconstructed copies. It is this linkage, not the earlier recall statistics, that supports copy reconstruction and assignment; the older numbers were measured on a workflow that did not attempt either.
What “assignment” means depends on the locus. For non-tandem paralogs such as SMN1 and SMN2, which sit far apart in the reference, proximity information assigns each reconstructed copy to the locus it most likely came from. For tandem paralogs such as CYP21A2 and CYP21A1P, which sit side by side, it instead groups copies onto the same chromosome — that is phasing. Calling one of those two chromosome haplotypes maternal and the other paternal is a further step that requires parental genotypes; linkage alone does not supply it.
The application note records that CFC1, IKBKG and HYDIN were evaluated but were not supported in TruPath Genome with MRJD in DRAGEN v4.5, This describes the evaluated release, not current availability.
A worked case: PMS2 and PMS2CL
PMS2 and PMS2CL are about as tangled as it gets, and PMS2CL sequence can be copied into PMS2 by gene conversion. In the figure below, MRJD separates the sample’s two PMS2 haplotypes: one carries a PMS2CL-to-PMS2 gene conversion tract, and PMS2CL itself is deleted on both haplotypes. The application note reports this configuration as consistent with the sample’s prior characterization.
Gene conversion can transfer damaging sequence, but its effect depends on the changes transferred. The source classifies this particular haplotype as non-functional without identifying a causal variant.
Counting alpha-globin copies
Alpha-globin genotyping
Resolving common alpha-globin copy-number configurations from whole-genome sequencing. I led development of the targeted HBA caller.
Compared with long-read data in 246 samples; separately, genotype calls were Mendelian-consistent in 575 trios.
Method, contribution & evaluation
- Problem
- The alpha-globin locus is structurally complex and its two near-identical genes defeat standard short-read genotyping, obscuring the copy-number configurations behind alpha-thalassemia.
- Method
- A targeted caller that reads several non-homologous regions within and near the HBA locus to estimate copy-number genotypes, covering 14 configurations spanning common molecular subtypes.
- My role
- Led development of the caller; first-listed author on the Illumina technical article and first-named inventor on the associated copy-number genotyping patent application.
- Evidence
- Evaluation against PacBio HiFi data in 246 diverse 1000 Genomes samples showed high concordance (four no-calls), and calls were consistent with Mendelian inheritance in 575 trios. Research use only.
Related work
Benchmarking germline CNV callers for clinical applications
Paper ↗A multi-author benchmark of short-read WGS copy-number callers against reference cell lines, quantifying how far current tools fall short of clinical sensitivity. I am one of twelve authors. Bioinformatics Advances, 2025.
Reconstructing insertions and genome history
TELR
Reconstructing transposable element insertions and estimating their allele frequencies in complex genomes. I developed the software and contributed analysis; joint first author with Guilherme B. Dias.
Insertion sequences reconstructed in Drosophila S2R+ supported family-level phylogenetic analysis.
Method, contribution & evaluation
- Problem
- Many transposable elements are longer than a short read, and even a long read spanning an insertion does not by itself give its exact boundaries, its family, or how many of a sample's genome copies carry it — which matters in aneuploid cell lines.
- Method
- TELR takes candidate insertion sites from structural-variant calls, locally assembles and polishes a contig spanning each insertion, identifies the TE family and reference coordinates, then estimates intra-sample TE allele frequency from read depth.
- My role
- Software development and analysis. Joint first author with Guilherme B. Dias — the paper states that the first two authors should be regarded as joint first authors.
- Evidence
- Applied to the Drosophila S2R+ cell line, TELR reconstructed insertion sequences for expanded TE families and estimated their intra-sample allele frequencies, supporting family-level phylogenetic analysis.
What the family trees actually show
Because TELR reconstructs the sequence of each insertion rather than only its location, we could build family trees for the most active TE families and ask how the extra copies in the S2R+ cell line are related to one another.
The answer turned out to depend on the family. The paper identifies a single expansion clade for 1731, gypsy, gypsy1, mdg3 and Stalker2, consistent with proliferation from one source lineage. It identifies multiple expansion clades for jockey, Juan and 3S18. So the pattern is family-dependent, not uniform.
These sequence phylogenies reveal relationships among copies, not the number of separate introductions into the culture.
Why the biology matters
TE insertions are individually rare and densely scattered, which makes them unusually good markers for telling cell lines apart. In Drosophila cell culture, genome-wide TE insertion profiles cluster replicate samples of the same line with 100% bootstrap support, and a panel of just six LTR retrotransposon families is enough to assign a sample correctly. That is a practical basis for cell-line authentication, and applying it surfaced lines whose recorded identity did not match their genome.
The same data connect TE activity to a second process. Large copy-neutral tracts of loss of heterozygosity (LOH) had erased SNP heterozygosity across whole chromosome arms. Ongoing transposition restores TE heterozygosity but not SNP heterozygosity, and that asymmetry is what makes a later, secondary LOH event recognisable — so the interaction between the two processes, rather than either alone, is what reveals the order of events in a line’s history.
Evidence and inferred genome histories
What is measured and what is inferred. The clustering and its bootstrap support, the six-family authentication panel, and the copy-number and B-allele-frequency evidence for copy-neutral LOH are direct observations. Mitotic recombination as the mechanism behind those copy-neutral tracts, the specific histories proposed for the misidentified lines, and the reading of one line as an ancestral state of another are inferences drawn from those patterns. In the S2 subline study, likewise, the phylogeny and the copy-number differences are measured, while “ongoing episodic transposition rather than a single early burst” is a model that the insertion-site occupancy and ancestral-state analyses support rather than something observed directly.
Related papers
TE profiles reveal cell line identity and loss of heterozygosity
Paper ↗Introduced ngs_te_mapper2 and used genome-wide TE insertion profiles to authenticate Drosophila cell lines, and identified LOH as a mechanism shaping those genomes. First author. Genetics, 2021.
Ongoing transposition reveals the phylogeny of Drosophila S2 sublines
Paper ↗Sequenced the genomes of 25 S2 sublines and used TE insertions as markers to reconstruct their relationships, finding that the S2 designation is paraphyletic and that copy-number evolution supports the same topology. First author. Genetics, 2022.
Reproducible evaluation of TE detectors with McClintock 2
Paper ↗A benchmarking meta-pipeline for twelve short-read TE detectors. Per the paper's contribution statement, the pipeline was developed by P. J. Basting with contributions from J. Chen, myself and C. M. Bergman. Mobile DNA, 2023.
P element target site prediction
Machine learning models trained on engineered features to predict P element insertion site preferences. Unpublished Ph.D. work.
Research software
Research software from my Ph.D., developed in the Bergman lab and released under open-source licences. My role differs between them, so it is stated on each.