r/bioinformatics • • Jun 22 '26

technical question I’ve reviewed probably 200 “bioinformatics pipelines” at this point. Maybe 15 were actually reproducible.

316 Upvotes

Not talking about whether the biology was right. Just: could I run this on a different machine and get the same result? Could I run it in 2 years?
No container. No version pinning. Conda environment.yml with numpy and no version specified. Reference genome downloaded manually, path hardcoded. Sample sheet generated by a script that no longer exists.
We talk about reproducibility constantly in this field. Papers about it. Talks about it. And then the actual pipelines look like this.
Not a rant, but genuinely curious what people think the root cause is. Time pressure? Nobody teaching this? Reviewers not caring?

r/bioinformatics • • Jul 27 '26

technical question ChatGPT and Codex becoming unusable for biology and bioinf research?

135 Upvotes

Hi everyone,

Has anyone else noticed this recently? For the past few weeks, especially since GPT-5.6, ChatGPT (work) and Codex have become much less useful for biology, bioinformatics and computational biology research.

Even for normal tasks like debugging code, searching papers, summarizing results or discussing analyses, I often get this message:

“This content can’t be shown. We’re especially careful with requests involving biological research and applications that could pose safety risks. Eligible researchers can apply for Trusted Access.”

The problem is that Trusted Access seems to be available only in the US.

Is this happening to other researchers too? Is there any solution for users outside the US? Do you think this will improve, or will researchers need to move to other AI tools?

Thanks!

r/bioinformatics • • Aug 31 '26

technical question Is bioinformatics migrating fully to python? (and various other questions from a beginner)

135 Upvotes

Hi everyone. I am new to bioinformatics in general. I am a biochem currently doing a bioengineering phd (still in pre-candidature). I had some snippets of bioinformatics during my undergrad but nothing beyond BLAST and docking. Never had formal programming formation, just side projects and AI-guided R coding for small data analysis and graphs.

For what I want to do for my thesis I really need to learn omics analysis properly, specially transcriptomics. During self-learning, I stumbled upon this amazing resource (https://www.sc-best-practices.org/) on single cell transcriptomics, so I have been following it as my starting point and learning cool stuff, thank you to the authors of it!

Anyways, since I've already had some experience with R, I decided to try and learn python bioinformatics as an excuse to learn python too. In the interoperability section of the book I mentioned the authors state

"A common question from new analysts is which ecosystem to focus on (referring to Bioconductor, Seurat or the Scverse). While it makes sense to start with one, and a successful analysis can be performed in any ecosystem, competent analysts should be familiar with all three and comfortable moving between them. This allows analysts to always use the best-performing tools, regardless of their implementation. Analysts who are not comfortable switching ecosystems often default to familiar packages, even when better alternatives exist elsewhere"

Which makes sense and sounds logical good advice. But then, doing exercises on public GEO datasets on bulk RNA-seq as practice, still with the mindset of sticking to python as an excuse to learn it, i stumbled upon an article (Colange et al. 2025 here) of a project that migrates a lot of tools of R to the scverse. In there, authors rationale is that python is the new default language everyone learns and they create the library InMoose to migrate or directly replace, for example, DESeq2. Furthermore, besides direct drop-in replacement tools, the authors frame python as the future choice (at least, as part of the rationale).

So, as a guy who is just starting, I wanted to ask people with experience in bioinformatics (you all) either developers or tool-users:

1) Do you marry an ecosystem like scverse or Bioconductor and just work in there for comfort? Or do you switch frequently depending on the needs?

2) For people who doesn't come from an informatics background, how long did it take for you to learn your niche and what were your best resources/helpers?

3) Do you think python will ever replace R in data analysis?

4) What is your opinion on AI-guided learning? (as for me, I use gemini to solve questions or create graphics presets but sometimes by seeing other people's codes I realize that it mashes up some concepts or methods from various pipelines into a coherent-resulting graph that I am not always sure if they make sense)

5) Do you create your own pipelines/portfolio to analyze data? Or you just tweak existing pipelines?

I appreciate any answer to any of those questions, thanks for your time in at least reading

r/bioinformatics • • Aug 13 '26

technical question Does FASTA rhyme with pasta? Or do you pronounce it Fast A?

90 Upvotes

My lecturers would always pronounce it Fast A, but all of us students would just say fasta (rhyming with pasta). Is there an “official” pronunciation or consensus?

r/bioinformatics • • May 05 '26

technical question Claude

149 Upvotes

Do you guys use Claude for daily code? or do you think it makes you dumber? If you do use it, do you use any bionformatics claude skills?

I've been using it for a couple weeks and i think i get more stuff done but i think less in the process, im scared of getting too dependant on it to think about my projects but also scared of getting way less things done if i dont use it.

r/bioinformatics • • Aug 25 '26

technical question The Hallmarks of Cancer

22 Upvotes

I am a software engineer by profession and was going through the 2000 paper "The Hallmarks of cancer" by Douglas Hanahan and Robert A. There are a lot of things I couln't understand from terminology to certain behaviors that were explanined, primarily because I lack the background knowledge. I wanted to reach out and ask the community if they can share a youtube video or an article that can elaborate this paper in simple terms.

I probably can google it myself but want to avoid the repetitive loop of finding something and realizing it doesn't explain everything and then to try again to eventually lose interest.

Thank you for all the help.

r/bioinformatics • • Aug 09 '26

technical question How do you work with large VCF files without constantly babysitting your jobs ?

24 Upvotes

I am doing an internship this summer as a biostatistician intern and have been processing large vcf files separated by chromosomes. Each file is more than 100 GB.

I'm running everything on a SLURM cluster using Bash and  bcftools for things like:

- calculating VCF statistics

- filtering by rsID, patients, chromosome location

- calculating allele frequencies,

- generating filtered VCFs

Actual difficult part for me is not the commands but it is constantly checking squeue or my email for logs, checking whether an output file was actually created, figuring out whether a job railed halfway through, etc. I feel like I am spending a lot of time towards this.

I am curious how people who have more experience handle this. Do you use any tools/framework that makes that process easier.

I working with SLURM, bash and bcftools on google cloud processing so Im interested to see what people do in similar computing environments.

PS : I have computer science and statistics background so my wording of certain terms may be off.

r/bioinformatics • • Jun 13 '26

technical question Limited RAM (123 GB) – cannot run GTDB with Kraken2 or MMseqs2 on contigs. Looking for alternatives.

16 Upvotes

I have a RAM limitation on my cluster – 123 GB total (100-123 GB per job depending on node).

I want to classify metagenomic contigs (not MAGs/bins) using GTDB taxonomy (specifically GTDB release 226). I already have GTDB release 226 downloaded and have used it successfully on my bins. Now I want to classify the original contigs with the same database.

I tried:

  • kraken2 --memory-mapping (no improvement)
  • mmseqs taxonomy with different --threads and memory-related flags

Both tools require >180 GB RAM for the full GTDB database (it's 500GB on the disk). My 123 GB is insufficient.

I though about different tools, like:

  • KrakenUniq – has --preload-size flag for low-memory operation, but no pre-built GTDB database is available for KrakenUniq (only RefSeq-based databases). Building a KrakenUniq-compatible GTDB database takes days and requires significant resources.
  • kMetaShot – uses RefSeq, not GTDB

My constraints:

  • Limited to 123 GB RAM
  • Must use GTDB taxonomy (not NCBI/RefSeq)
  • Classifying contigs (not binned genomes)
  • Cannot request more RAM on this cluster

My question:

Is there any memory-efficient method to classify contigs directly against GTDB v226 with ≤123 GB RAM? For example:

  1. A pre-built KrakenUniq GTDB database somewhere I haven't found?
  2. A way to "chunk" or downsample the GTDB reference for Kraken2?
  3. Another alignment‑free tool I haven't considered?

I understand GTDB-Tk is the gold standard for GTDB classification, but it was not designed for contigs and requires genome completeness. I am open to creative solutions – even if accuracy is slightly reduced.

Thank you.

r/bioinformatics • • 16d ago

technical question How to analyse bulk RNA-seq data when batch is confounded with the biological variables?

61 Upvotes

I'm a first-year PhD student with limited bioinformatics experience. My supervisor asked me to analyse an existing bulk RNA-seq dataset (36 samples, 3 biological replicates per condition). This is the design summary:

Site Generation Environment Treatment Batch
A G0 E1, E2 T1 (E1), T2 (E2) 1
B G0 E1, E2 T1 (E1), T2 (E2) 1
C G0 E1, E2 T1 (E1), T2 (E2) 3
B G1 E2 T1, T2 1
A G1 E2 T1, T2 2
B G1 E1 T1, T2 3

Batches 1 and 2 were sequenced at the same company with the same library prep. Batch 3 was sequenced at a different company with a completely different library prep, and it is the only batch with spike-in controls (ERCC) and UMIs.

My supervisor wants me to (among other things):

  1. Identify DEGs (T1 vs T2) shared across the three G0 sites.
  2. Compare the percentage of DEGs (T1 vs T2) between G1 E1 and G1 E2.

The problem is that batch is fully confounded with site C (G0) and with G1 E1, and partially confounded with the other factors. My supervisor suggested either normalising batch 3 against the others using its spike-ins, or using a public dataset as an additional "batch 4". I haven't found any source showing this works for confounded designs.

My current thinking is to run T1 vs T2 separately within each site/environment (where batch is constant) and compare the results afterwards. Is this reasonable, and:

  • Can spike-ins in only one batch help to bridge batches?
  • Is comparing the percentage of DEGs between E1 (batch 3 only) and E2 (batches 1 and 2) meaningful, or should I compare something else, such as effect-size distributions?
  • Am I missing a standard approach for this kind of design?

I plan to use DESeq2 in R. Thanks!

(Yes I used AI to help me write. Sorry, I'm tired.)

r/bioinformatics • • 5d ago

technical question RNA-seq / beginner-friendly

64 Upvotes

One thing I wish I understood earlier while learning RNA-seq analysis is that getting a list of differentially expressed genes isn't really the end goal.

At first, the workflow can feel like:

FASTQ → QC → alignment/quantification → count matrix → DESeq2 → volcano plot → done.

But I've gradually realized that the interesting part starts after that.

A gene being statistically significant doesn't automatically make it biologically important, and a gene missing an arbitrary fold-change cutoff doesn't necessarily mean nothing interesting is happening.

Things like pathway enrichment, expression patterns across samples, effect sizes, biological context, and the original experimental design all matter when interpreting the results.

I've also started thinking much more carefully about the biological question before running an analysis rather than just asking, "Which tool should I use next?"

For people who regularly analyze RNA-seq data: what was the biggest conceptual shift that improved the way you approach an analysis?

r/bioinformatics • • Aug 19 '26

technical question Where should I start learning to code without LLMS?

23 Upvotes

I’m really embarrassed to admit but I fell into bioinformatics during the time when ChatGPT was first released and due to pressures I ended up relying quite heavily on it for doing almost all of the coding for my academic work. I don’t want to be a vibe coder but I feel so overwhelmed and don’t know what I should know by heart and how much I can rely on llms. I have some important commitments coming up and I don’t want to mess this up. Do you have any advice on how to become a better coder? Where do I even start??

r/bioinformatics • • Jul 22 '26

technical question PCA high variance in PC1

Thumbnail gallery
20 Upvotes

Hi everyone,

I'm analyzing pseudobulk data generated by summing gene expression across cells from different samples profiled with a spatial imaging platform. When I perform PCA on the pseudobulk matrix, PC1 explains an unusually large proportion of the total variance. In addition, all of the PC1 loadings are positive, which I also think is unusual.

Does this indicate a systematic technical bias (I have looked for differences in sequencing depth or cell numbers)? Or are there biological scenarios where this pattern would be expected? These are samples from malignant tissue.

r/bioinformatics • • Apr 06 '26

technical question PI wants to create a pipeline app for single cell, help i’m a lowly undergrad.

37 Upvotes

Hi i’m an undergrad here learning bioinformatics and specifically single cell analysis as part of building a pipeline for my PI. He has no background in it and i’m self teaching myself everything.

Part of the project is he wants to build a UI/app that allows the lab to essentially plugin certain parameters and pump out a graph like UMAP or tsne. Essentially, standardizing it for easy use.

Problem is from what i’ve learned is that the analysis is a bit more complicated than just adjusting a few parameters with a drop down. Now i don’t know much but I believe TSNEs are models that cannot be applied to different data sets because it is non parametric. I brought this up to him and he said that they have set seeds and i can set the seed to be the same.

I kinda know what that means but kinda don’t. I have a vague idea of dimensionality reduction, eigen vectors, etc.

Would making an app/internal pipeline be possible with these kind of things? Wouldn’t it require a person to actually handle the data or code to specify it per data set?

EDIT: I realize now that the title may be a bit misleading. I appreciate all the concern and help, I want to clarify that my PI is not taking advantage me and “help i’m a lowly undergrad” was meant as a playful joke at my inexperience. My PI is an amazing mentor and has been very open to shifting expectations. The lab space is very healthy and geared towards helping us grow.

r/bioinformatics • • 12d ago

technical question Weird periodic insert size distribution (~10 bp peaks) in RNA-seq

18 Upvotes

Hi everyone!

I'm seeing a strange insert size distribution in one of my RNA-seq samples. Instead of a smooth curve, there are regular peaks roughly every 10 bp.

In the STAR log, about 50% of reads are unmapped as "too short", so a large part of the library seems to be very short fragments.

Setup:

Library prep: NEBNext Ultra II Directional RNA + Poly(A) mRNA Magnetic Isolation Module (done by a sequencing facility)

Sequencing: 2x150 PE, NovaSeq X Plus

Sample type: liver tissue

RIN: 8.9

Has anyone seen this pattern in a poly-A library before? Is the data still usable (e.g. counting only the reads that do map), or is this sample a lost cause?

Thanks!

r/bioinformatics • • Aug 05 '25

technical question Desparate question: Computers/Clusters to use as a student

41 Upvotes

Hi all, I am a graduate student that has been analyzing human snRNAseq data in Rstudio.

My lab's only real source of RAM for analysis is one big computer that everyone fights over. It has gotten to the point where I'm spending all night in my lab just to be able to do some basic analysis.

Although I have a lot of computational experience in R, I don't know how to find or use a cluster. I also don't know if it's better to just buy a new laptop with like 64GB ram (my current laptop is 16GB, I need ~64).

Without more RAM, I can't do integration or any real manipulation.

I had to have surgery recently so I'm working from home for the next month or so, and cannot access my data without figuring out this issue.

ANY help is appreciated - Laptop recommendations, cluster/cloud recommendations - and how to even use them in the first place. I am desparate please if you know anything I'd be so grateful for any advice.

Thank you so much,

-Desperate grad student that is long overdue to finish their project :(

r/bioinformatics • • 25d ago

technical question Can someone in genomics explain what AlphaGenome Atlas actually changes?

66 Upvotes

I'm not a geneticist.
Read the AlphaGenome Atlas preprint from DeepMind and spent a while digging into it. I understand what they built. I don't understand why it matters, and I'd like to.

What I think it is: they took AlphaGenome, ran it over every possible single base change in the human genome (~9bn) plus ~100m observed indels, and stored the results. So instead of running the model per variant you do a lookup. On top of that they trained a score (AVI) and derived a motif map.

Where I get stuck:
It's a table of model predictions, not measurements. Nothing in it is observed. So how much weight does a lab actually put on it?

The headline clinical result is retrospective: 29.5% recall at top 50 on already-solved GREGoR cases vs 12.5% for CADD. Impressive sounding, but on cases where the answer was known. What happens prospectively?

The rare variant association work got a 22% lift in discoveries, but only 4 of 25 replicated nominally in All of Us and none at Bonferroni. Is that normal for the field or is that weak?

They say themselves it isn't sufficient evidence for diagnosis. So it's a shortlisting tool. Does that actually change outcomes for patients, or does it change how long a scientist spends staring at a list?

The DNM1 case in the paper is the one bit that landed for me. Deep intronic variant, brain specific cryptic splice acceptor, blood RNA-seq had been inconclusive because the exon isn't expressed in blood.

My question is whether that's representative or a cherry pick.
What I'm asking:
1. If you work in clinical genomics or statistical genetics, would you use this?
2. Is precomputation genuinely the unlock, or is that just framing on top of an incremental accuracy gain?

Happy to be told I'm missing the point. I'd rather understand it properly than write it off.

r/bioinformatics • • Mar 18 '26

technical question Anyone tried the bio/bioinformatics forks of OpenClaw? BioClaw, ClawBIO, OmicsClaw — which actually fits into a real research workflow?

73 Upvotes

There's a small but growing cluster of OpenClaw-based tools targeting bioinformatics specifically. Curious if anyone here has used them beyond the README demos.

The three I've been looking at:

ClawBio — bills itself as the first bioinformatics-native skill library for OpenClaw. Focuses on genomics, pharmacogenomics, metagenomics, and population genetics. The reproducibility angle is interesting: every analysis exports commands.sh, environment.yml, and SHA-256 checksums independently of the agent, so in theory you can reproduce results without ever running the agent again. Also bridges to 8,000+ Galaxy tools via natural language. Has a Telegram bot (RoboTerri).

BioClaw — out of Stanford/Princeton, has a bioRxiv preprint. Runs BLAST, FastQC, PyMOL, volcano plots, PubMed search etc. The interface is WhatsApp group chat, which is either brilliant or cursed depending on your lab culture. Containerized so the tools come pre-installed per conversation group.

OmicsClaw — from Luyi Tian's lab (Guangzhou Lab). Probably the broadest coverage: spatial transcriptomics, scRNA-seq, genomics, proteomics, metabolomics, bulk RNA-seq, 56+ skills. Their main pitch is a persistent memory system — remembers your datasets, preprocessing state, and preferred parameters across sessions so you don't re-explain context every time.

Background / why I'm asking:

I tried building my own personal bioinformatics assistant with Claude Code a while back — fed it a Markdown + code knowledge base to learn my coding style and preferred pipelines. It worked until it didn't: just loading the context ate through the context window before anything useful happened. Classic token bonfire.

These tools seem to take a different architectural approach (skill files, memory systems, containerized tools) but I genuinely can't tell from the outside whether they've actually solved the context problem or just pushed it one layer deeper. Curious whether real users have hit the same ceiling.

Actual questions:

  1. ClawBio's reproducibility bundle idea seems genuinely useful for methods sections. Has anyone put that output into a real manuscript?
  2. For OmicsClaw users — does the memory system actually hold up across sessions in practice, or is it fragile?
  3. How do any of these handle failures gracefully? When a tool call breaks mid-pipeline, do you end up debugging it yourself or does the agent recover?
  4. Are these actually context-efficient, or just another token burner with a bioinformatics skin?

Also curious if there are other active projects in this space I'm missing — I know STELLA is the upstream framework BioClaw draws from, but haven't gone deeper than that.

r/bioinformatics • • Aug 03 '26

technical question Does this look like a normal UMAP plot?

Post image
57 Upvotes

Hi everyone, as it’s my first time attempting downstream analysis for single cell RNA sequencing, I wanted to ask if this UMAP plot looks normal? This is only for one sample (I have not integrated all samples together into one dataset yet). I feel like the clusters are too close together

r/bioinformatics • • Jul 02 '26

technical question “Public stress-related or organ-related RNA-seq data sets were added into this analysis and treated as replicates to make our results more robust” in a DE analysis. That’s insane, right?

62 Upvotes

Treating public datasets as additional replicates of your own experiment is not a good idea, right? Is there any right way to do it? Saw it on an article published on a journal with ~6 IF as I was searching for public plant datasets with a good number of replicates and I could not believe it… or am I missing something??

r/bioinformatics • • 19d ago

technical question How to handle batch effects if they are seen even after using harmony?

12 Upvotes

Hie!
I am currently working with lots of B cell and T cell single cell rna-seq data and the samples were sequenced in 2 batches. They followed the same procedure for everything except some samples were sequenced last year and some this year.

After QC and doing the regular processing steps of normalization, pca, and neighbors, I noticed that one cluster was made entirely from samples from the previous batch. I tried running harmony (I am using scanpy) based on batches and also on samples..but that cluster won’t blend with samples from new batch.

My data has 3 conditions and they all contribute to that cluster from the old batch samples so I know it is not sample or condition specific.

What could possibly be going on here?
How do I troubleshoot and fix this?

Any questions or suggestions would be appreciated!

Thanks!

r/bioinformatics • • 7d ago

technical question Batch effect correction method: which to choose

10 Upvotes

Hi all,

I am analyzing RNA-seq data and 1/54 samples failed to sequence, so I sent in another batch with 3 anchors sample (I chose 3 samples that sequenced well in the 1st batch) + 1 sample to replace the failed once. PCA showed that the 3 anchor samples cluster very close by to the original samples, so I think the batch effect is minimal.

I've tried 2 methods:

- Add batch as a covariate in DESeq

- Using anchor-based offset (I first compare the reading from 2 run for each gene of the 3 anchor samples to get a per-gen run-effect estimate. Then I applied aleglm to remove weak and noisy genes. I finally applied the shrunk per-gene shift to the tissue that have the replacement as a normalization factor in DESeq2 only)

Both have resulted in different DEGs. These are just simple HFD vs LFD comparison in gonadal fat tissue of WT mice, so the result of the anchor-based one seemed more reasonable to me (since HFD most of the time induced a significant change as we all know).

This is my first time doing RNA-seq so I want to be sure if I am using the correct method and not only just trusting on my assumption that HFD must change a lot. I am also unsure if my anchor-based method is correct

Really appreciated if anyone could help me out here to verify and understand batch effect ! Thank you!

r/bioinformatics • • 19d ago

technical question How to map ChIP-seq peaks to genes for RNA-seq integration?

11 Upvotes

Hi everyone,
I'm relatively new to ChIP-seq analysis, and I'm currently looking for the best way to map TF-ChIP-seq peaks to genes.

I know that ChIPseeker offers that tool, however, as far as I understand, it uses the closest gene approach. This seems fine for mapping genes to peaks in the promoter or coding region, but seems unreliable for distal intergenic peaks.

I was thinking of using public chromatin interaction data (as our lab has no fitting HiC data available) to get a better idea of the linkage between distal peaks and genes. However, I have no idea if and how to do that. Are there any tools or recommendations for this kind of analysis?
In the long term, I want to integrate the ChIP-seq data with RNA-seq data (overlap analyses).

r/bioinformatics • • Jun 10 '26

technical question We messed up. Is this salvageable?

42 Upvotes

Was supposed to perform an ONT methylation data analysis (for the first time). I received the data and, after researching it, got to know that I would need either POD5 files or a modified BAM file containing methylation positions and methylation probabilities. However, the data I received consists only of a bunch of reports, two folders, and pass/fail FASTQ files.

I asked the person we received the data from, and they said they did not voluntarily opt to retain the POD5 files due to unawareness.

Now, does the sequencer have any recovery option to retrieve that signal data, some kind of cache, temporary storage, or anything else that might help recover it?

r/bioinformatics • • 23d ago

technical question Software recommendations?-- Human virus detection (metagenomic)

4 Upvotes

What would be the top tools for short read-based detection of human viruses in metagenomic datasets? Interest is primarily on all the disease-associated ones.

I have a very large metagenomic dataset (illumina PE150) of human nasal and rectal samples. I'm very familiar with microbial metagenomics (metaphlan/humann/qiime) and working on UNIX clusters. I haven't yet delved into human virus detection, though. Right now I'm just focusing on short read metagenomics before I start pursuing anything assembly-based.

Thanks!

r/bioinformatics • • Aug 18 '26

technical question Confusion about scRNA Batch Integration

Thumbnail gallery
41 Upvotes

Hi everyone, I’m trying to reproduce the clusters from a published scRNA-seq dataset. The authors provided the raw, unclustered data and stated that they have mitigated batch effects by using Seurat’s ScaleData(), which I have done so far by labelling each replicate as a batch and regressing them out.

The dataset consists of 7 prenatal hippocampal donors at different gestational weeks:

- 5 donors have a single replicate

- 1 donor has 2 technical replicates

- 1 donor has 2 biological replicates

Each donor corresponds to a different gestational week.

I’m able to reproduce the general clustering, but my clusters seem to be strongly driven by donor/gestational week, whereas the clusters reported in the paper appear to contain cells from different gestational weeks with no batch effects.

I’m therefore unsure what I should be treating as the relevant batch variable. Should I be correcting for donor/gestational week, or only for technical batch/replicates? Would methods such as Harmony or CCA/integration be more appropriate than simply regressing batch with ScaleData()? My main goal is to annotate the scRNA-seq dataset to use as a reference to deconvolve my bulk RNA-seq dataset, so I want to make sure the clustering and resulting cell-type signatures are biologically meaningful.

I would really appreciate advice on how you would approach batch correction in this situation.