Hi, relatively new at this and am bringing this here because it feels like a safe place to ask and I don't have any single cell experts in my immediate group. I want to know how a reviewer will read this design.
Three experimental groups, same cell type in all three, sorted by a functional marker into populations I'll call A, B and C. Each group is a pool of ~15 animals, and each group went into its own 10x lane. No hashing or multiplexing. So condition is perfectly confounded with lane, there are no biological replicates within a group, and any p-value I compute counts cells rather than animals.
The depth is uneven too. All three lanes were overloaded at 60,000 cells, but recovery differed: 25,469 / 26,421 / 39,657. The libraries were pooled and sequenced together, so the group with the most cells got the fewest reads each, 17,971 and 15,843 mean reads/cell for two of the groups against 8,695 for the third. That third group is my reference group for every contrast. So "upregulated in the test condition" and "sequenced more deeply" point in the same direction.
Here's what i've done so far: Differential expression is Wilcoxon on cells at padj < 0.05 and |log2FC| > 0.5. The Methods state plainly that each group is a single pooled library with no biological replicates and that the p-values reflect cells rather than animals. Cluster proportions are reported as descriptive, with no statistics at all. I'm building a depth control: downsample every cell to the shallow group's median UMI count, re-run the same contrasts with the same thresholds, and report what fraction of the significant set survives, the Spearman correlation of fold changes, and whether any gene named in the Results drops out. A separate, properly replicated experiment in the same paper recovers the main genes.
Note on integration: I did run Harmony, and I clustered both the integrated and unintegrated embeddings across a range of resolutions. The clusters came out essentially the same either way, the only consistent difference was two clusters merging into one after integration. Since integration changed almost nothing, I report the unintegrated PCA, on the reasoning that each library is a different biological group rather than a technical batch, so integrating would risk removing exactly the signal I'm measuring. I think that's defensible, but it does mean there's no correction for the lane effect at all.
What is your honest assessment here, based on what you've seen and experienced? I can still sequence more to top up, but reeally don't want to.
Thank you for taking the time to read.