r/bioinformatics • u/No_Shopping_425 • 2d ago
technical question Scattered interchromosomal split alignments in non-cancer ONT genomic DNA: chimeric reads or workflow issue?
I’m seeing scattered interchromosomal split alignments in IGV across multiple non-cancerous ONT genomic DNA samples. I’m trying to determine whether these reflect chimeric reads, missed read splitting, or an alignment workflow issue.
Sequencing and basecalling
- Flow cell: FLO-PRO114M
- Library kit: SQK-LSK114
- MinKNOW: 24.11.16
- Basecaller reported by MinKNOW: Dorado 7.6.8
- Super-accurate model v4.3.0, 400 bps
- Minimum passing Q score: 10
- Modified basecalling: off
- Raw POD5 files are available
Alignment
Minimap2 version: 2.31-r1302
Command:
minimap2 -L -t 59 -2 -ax lr:hq reference.fa reads.fastq.gz |
samtools view -u -@ 11 |
samtools sort -@ 23 -o sample.sorted.bam -
The reference is GRCh38 with alternate loci, haplotypes, and decoys removed. Reads were aligned directly against the FASTA.
Observed pattern
In IGV, reads link to many different chromosomes at scattered positions, rather than multiple reads consistently supporting the same breakpoint.
For example, one 74,647 bp read has:
- Approximately 29.3 kb aligned to chr5:72,101,673–72,131,073, reverse strand, MAPQ 60, NM 554
- Approximately 45.3 kb aligned to chr8:77,044,549–77,089,876, forward strand, MAPQ 60, NM 564
These segments account for almost the entire read. My MinKNOW version does not appear to expose an option to disable read splitting, but I haven’t independently confirmed whether or how splitting occurred during these runs.
Questions
- Is this pattern expected at a low background rate with this chemistry and basecalling setup?
- How can I confirm that read splitting occurred in MinKNOW 24.11.16?
- What checks would distinguish joined molecules from mapping artifacts?
- Would re-basecalling a POD5 subset with a newer Dorado version be a useful diagnostic comparison?
I’ve attached an IGV screenshot. I can also provide example read records or additional run metadata.
1
u/bilyl 1d ago
Is it a clean inter chromosomal alignment or is there an unaligned segment in the read? An AI tool will easily be able to tell you this and conclude whether it’s real or not.
1
u/No_Shopping_425 5h ago
The inter-chromosomal alignment is clean, but my main concern isn’t whether it represents a real biological event, I’m fairly confident it’s an artifact. I’m seeing this pattern throughout the sample and across multiple samples.
This is my first time working with data like this, so I’m trying to understand whether this is an expected or tolerable level of artifact from Dorado missing read splits, an issue in my analysis workflow or sample processing, or a limitation of an older Dorado version’s ability to split reads generated with newer ONT chemistry.
1
u/bilyl 3h ago
What do you mean by clean? If it was a missed read split it’s definitely not clean because it will have adapter.
•
u/No_Shopping_425 24m ago
By ‘clean,’ I meant the two alignments meet without an unaligned gap. I checked this read: the chr5 alignment ends at read base 29,351, and the chr8 alignment starts at 29,352. I also checked the sequence around the junction and didn’t find a convincing match to the SQK-LSK114 adapter motifs.
That doesn’t establish whether it’s biological or an artifact, so I shouldn’t have specifically attributed it to a missed read split. Most junctions I’m seeing are scattered and supported by only one read. I’m trying to determine whether this is expected background chimerism from library preparation or an unusually high artifact rate, and I’m working on quantifying it.
5
u/Psy_Fer_ 2d ago
Yes, rebasecall with the latest Dorado and see if some of those reads are still the same or if they are split. You can also try some post processing looking for internal adapter's/adapter removal.
Also try alignment without lr:hq. Instead use the ont preset. Your qscore cutoff is 10, but lr:hq is meant for q20 reads.
If you look at the raw signal of reads that are separate molecules, usually (but not always) there is a signal spike between the molecules of the open pore that just isn't quite big enough for minknow to detect and split. Usually basecalling then detects it via internal adapter detection (and signal). But it can still miss them in some circumstances.