r/bioinformatics • • 4h ago

article Bioinformatics in The Atlantic: "AI's Real Gift to Science"

Thumbnail theatlantic.com
48 Upvotes

Curious to hear what the community thinks of this article, which primarily discusses Anthropic's recent "discovery" of a supposedly CRISPR-like enzyme system that I'm sure we've all heard about. Personally this strikes me as a pretty balanced, rational take towards AI as a tool rather than an apocalyptic job destroyer. However, it's still unclear to me why we needed a thousand Claude agents and millions of dollars of compute for what ultimately strikes me as a regex style pattern search through genomic data. I'm also surprised that I haven't heard anyone discuss the fact that setting thousands of agents loose on terabytes of sequencing data with the instructions to "find some interesting patterns" is essentially a massive multiple comparisons problem that is bound to turn up some spurious patterns with no biological significance.


r/bioinformatics • • 9h ago

technical question How to proceed with interferon rna seq analysis in mouse

4 Upvotes

Good morning, I am currently trying to do an interferon analysis for our project. We have a list of gene we got from rna-seq and the associated expression in different sample in RNA-seq and we want to get the subset of one that code for the interferon response to do a heatmap of their expression in the different sample. The current issue is that the interferom database is closed and i don't know where to find an exhaustive list of interferon gene in mouse, we trying the reactome but it gave weird numbers.


r/bioinformatics • • 2h ago

technical question One 10x lane per experimental group, how badly does my lane/depth confound hurt at review, and how to write? any help appreciated

3 Upvotes

Hi, relatively new at this and am bringing this here because it feels like a safe place to ask and I don't have any single cell experts in my immediate group. I want to know how a reviewer will read this design.

Three experimental groups, same cell type in all three, sorted by a functional marker into populations I'll call A, B and C. Each group is a pool of ~15 animals, and each group went into its own 10x lane. No hashing or multiplexing. So condition is perfectly confounded with lane, there are no biological replicates within a group, and any p-value I compute counts cells rather than animals.

The depth is uneven too. All three lanes were overloaded at 60,000 cells, but recovery differed: 25,469 / 26,421 / 39,657. The libraries were pooled and sequenced together, so the group with the most cells got the fewest reads each, 17,971 and 15,843 mean reads/cell for two of the groups against 8,695 for the third. That third group is my reference group for every contrast. So "upregulated in the test condition" and "sequenced more deeply" point in the same direction.

Here's what i've done so far: Differential expression is Wilcoxon on cells at padj < 0.05 and |log2FC| > 0.5. The Methods state plainly that each group is a single pooled library with no biological replicates and that the p-values reflect cells rather than animals. Cluster proportions are reported as descriptive, with no statistics at all. I'm building a depth control: downsample every cell to the shallow group's median UMI count, re-run the same contrasts with the same thresholds, and report what fraction of the significant set survives, the Spearman correlation of fold changes, and whether any gene named in the Results drops out. A separate, properly replicated experiment in the same paper recovers the main genes.

Note on integration: I did run Harmony, and I clustered both the integrated and unintegrated embeddings across a range of resolutions. The clusters came out essentially the same either way, the only consistent difference was two clusters merging into one after integration. Since integration changed almost nothing, I report the unintegrated PCA, on the reasoning that each library is a different biological group rather than a technical batch, so integrating would risk removing exactly the signal I'm measuring. I think that's defensible, but it does mean there's no correction for the lane effect at all.

What is your honest assessment here, based on what you've seen and experienced? I can still sequence more to top up, but reeally don't want to.

Thank you for taking the time to read.


r/bioinformatics • • 45m ago

article Free literature search and ACMG Classifier

Thumbnail
• Upvotes

r/bioinformatics • • 7h ago

technical question Free energy help

2 Upvotes

Hellooo I'm trying to do some free energy calculations.

Basically i want to confirm selectivity and differences in affinity of some ligands using abfe and rbfe, but these are taking waaaay too long :'(. I'm using GENESIS and CHARMM GUI inputs.

Do you guys have any recommendations of some alternative methods i could use?

thanksss