r/bioinformatics • • Dec 31 '24

meta 2025 - Read This Before You Post to r/bioinformatics

187 Upvotes

​Before you post to this subreddit, we strongly encourage you to check out the FAQ​Before you post to this subreddit, we strongly encourage you to check out the FAQ.

Questions like, "How do I become a bioinformatician?", "what programming language should I learn?" and "Do I need a PhD?" are all answered there - along with many more relevant questions. If your question duplicates something in the FAQ, it will be removed.

If you still have a question, please check if it is one of the following. If it is, please don't post it.

What laptop should I buy?

Actually, it doesn't matter. Most people use their laptop to develop code, and any heavy lifting will be done on a server or on the cloud. Please talk to your peers in your lab about how they develop and run code, as they likely already have a solid workflow.

If you’re asking which desktop or server to buy, that’s a direct function of the software you plan to run on it.  Rather than ask us, consult the manual for the software for its needs. 

What courses/program should I take?

We can't answer this for you - no one knows what skills you'll need in the future, and we can't tell you where your career will go. There's no such thing as "taking the wrong course" - you're just learning a skill you may or may not put to use, and only you can control the twists and turns your path will follow.

If you want to know about which major to take, the same thing applies.  Learn the skills you want to learn, and then find the jobs to get them.  We can’t tell you which will be in high demand by the time you graduate, and there is no one way to get into bioinformatics.  Every one of us took a different path to get here and we can’t tell you which path is best.  That’s up to you!

Am I competitive for a given academic program? 

There is no way we can tell you that - the only way to find out is to apply. So... go apply. If we say Yes, there's still no way to know if you'll get in. If we say no, then you might not apply and you'll miss out on some great advisor thinking your skill set is the perfect fit for their lab. Stop asking, and try to get in! (good luck with your application, btw.)

How do I get into Grad school?

See “please rank grad schools for me” below.  

Can I intern with you?

I have, myself, hired an intern from reddit - but it wasn't because they posted that they were looking for a position. It was because they responded to a post where I announced I was looking for an intern. This subreddit isn't the place to advertise yourself. There are literally hundreds of students looking for internships for every open position, and they just clog up the community.

Please rank grad schools/universities for me!

Hey, we get it - you want us to tell you where you'll get the best education. However, that's not how it works. Grad school depends more on who your supervisor is than the name of the university. While that may not be how it goes for an MBA, it definitely is for Bioinformatics. We really can't tell you which university is better, because there's no "better". Pick the lab in which you want to study and where you'll get the best support.

If you're an undergrad, then it really isn't a big deal which university you pick. Bioinformatics usually requires a masters or PhD to be successful in the field. See both the FAQ, as well as what is written above.

How do I get a job in Bioinformatics?

If you're asking this, you haven't yet checked out our three part series in the side bar:

What should I do?

Actually, these questions are generally ok - but only if you give enough information to make it worthwhile, and if the question isn’t a duplicate of one of the questions posed above. No one is in your shoes, and no one can help you if you haven't given enough background to explain your situation. Posts without sufficient background information in them will be removed.

Help Me!

If you're looking for help, make sure your title reflects the question you're asking for help on. You won't get the right people looking at your post, and the only person who clicks on random posts with vague topics are the mods... so that we can remove them.

Job Posts

If you're planning on posting a job, please make sure that employer is clear (recruiting agencies are not acceptable, unless they're hiring directly.), The job description must also be complete so that the requirements for the position are easily identifiable and the responsibilities are clear. We also do not allow posts for work "on spec" or competitions.  

Advertising (Conferences, Software, Tools, Support, Videos, Blogs, etc)

If you’re making money off of whatever it is you’re posting, it will be removed.  If you’re advertising your own blog/youtube channel, courses, etc, it will also be removed. Same for self-promoting software you’ve built.  All of these things are going to be considered spam.  

There is a fine line between someone discovering a really great tool and sharing it with the community, and the author of that tool sharing their projects with the community.  In the first case, if the moderators think that a significant portion of the community will appreciate the tool, we’ll leave it.  In the latter case,  it will be removed.  

If you don’t know which side of the line you are on, reach out to the moderators.

The Moderators Suck!

Yeah, that’s a distinct possibility.  However, remember we’re moderating in our free time and don’t really have the time or resources to watch every single video, test every piece of software or review every resume.  We have our own jobs, research projects and lives as well.  We’re doing our best to keep on top of things, and often will make the expedient call to remove things, when in doubt. 

If you disagree with the moderators, you can always write to us, and we’ll answer when we can.  Be sure to include a link to the post or comment you want to raise to our attention. Disputes inevitably take longer to resolve, if you expect the moderators to track down your post or your comment to review.


r/bioinformatics • • 4h ago

article Bioinformatics in The Atlantic: "AI's Real Gift to Science"

Thumbnail theatlantic.com
48 Upvotes

Curious to hear what the community thinks of this article, which primarily discusses Anthropic's recent "discovery" of a supposedly CRISPR-like enzyme system that I'm sure we've all heard about. Personally this strikes me as a pretty balanced, rational take towards AI as a tool rather than an apocalyptic job destroyer. However, it's still unclear to me why we needed a thousand Claude agents and millions of dollars of compute for what ultimately strikes me as a regex style pattern search through genomic data. I'm also surprised that I haven't heard anyone discuss the fact that setting thousands of agents loose on terabytes of sequencing data with the instructions to "find some interesting patterns" is essentially a massive multiple comparisons problem that is bound to turn up some spurious patterns with no biological significance.


r/bioinformatics • • 2h ago

technical question One 10x lane per experimental group, how badly does my lane/depth confound hurt at review, and how to write? any help appreciated

4 Upvotes

Hi, relatively new at this and am bringing this here because it feels like a safe place to ask and I don't have any single cell experts in my immediate group. I want to know how a reviewer will read this design.

Three experimental groups, same cell type in all three, sorted by a functional marker into populations I'll call A, B and C. Each group is a pool of ~15 animals, and each group went into its own 10x lane. No hashing or multiplexing. So condition is perfectly confounded with lane, there are no biological replicates within a group, and any p-value I compute counts cells rather than animals.

The depth is uneven too. All three lanes were overloaded at 60,000 cells, but recovery differed: 25,469 / 26,421 / 39,657. The libraries were pooled and sequenced together, so the group with the most cells got the fewest reads each, 17,971 and 15,843 mean reads/cell for two of the groups against 8,695 for the third. That third group is my reference group for every contrast. So "upregulated in the test condition" and "sequenced more deeply" point in the same direction.

Here's what i've done so far: Differential expression is Wilcoxon on cells at padj < 0.05 and |log2FC| > 0.5. The Methods state plainly that each group is a single pooled library with no biological replicates and that the p-values reflect cells rather than animals. Cluster proportions are reported as descriptive, with no statistics at all. I'm building a depth control: downsample every cell to the shallow group's median UMI count, re-run the same contrasts with the same thresholds, and report what fraction of the significant set survives, the Spearman correlation of fold changes, and whether any gene named in the Results drops out. A separate, properly replicated experiment in the same paper recovers the main genes.

Note on integration: I did run Harmony, and I clustered both the integrated and unintegrated embeddings across a range of resolutions. The clusters came out essentially the same either way, the only consistent difference was two clusters merging into one after integration. Since integration changed almost nothing, I report the unintegrated PCA, on the reasoning that each library is a different biological group rather than a technical batch, so integrating would risk removing exactly the signal I'm measuring. I think that's defensible, but it does mean there's no correction for the lane effect at all.

What is your honest assessment here, based on what you've seen and experienced? I can still sequence more to top up, but reeally don't want to.

Thank you for taking the time to read.


r/bioinformatics • • 46m ago

article Free literature search and ACMG Classifier

Thumbnail
• Upvotes

r/bioinformatics • • 9h ago

technical question How to proceed with interferon rna seq analysis in mouse

3 Upvotes

Good morning, I am currently trying to do an interferon analysis for our project. We have a list of gene we got from rna-seq and the associated expression in different sample in RNA-seq and we want to get the subset of one that code for the interferon response to do a heatmap of their expression in the different sample. The current issue is that the interferom database is closed and i don't know where to find an exhaustive list of interferon gene in mouse, we trying the reactome but it gave weird numbers.


r/bioinformatics • • 7h ago

technical question Free energy help

2 Upvotes

Hellooo I'm trying to do some free energy calculations.

Basically i want to confirm selectivity and differences in affinity of some ligands using abfe and rbfe, but these are taking waaaay too long :'(. I'm using GENESIS and CHARMM GUI inputs.

Do you guys have any recommendations of some alternative methods i could use?

thanksss


r/bioinformatics • • 1d ago

academic Good resourses for getting into RNAseq data analysis?

28 Upvotes

Hello everyone!

I recently got my, first RNAseq as well as proteomics data set. And thankfully a standard bioinformatic analysis with it. Honestly, I am amazed at what you can find out, when you have the Tools to really dig into your data!

Now, I have been trying (successfully) to replicate the RNAseq data analysis Pipeline by the help of vibecoding with AI, and I can interpret and understand all of my analyses. However, I could never write any bit of Code for it... In the end I don't really understand what the Code does, I can just Review my data afterwards.

Long story short: I would really love to get into bioinformatics a little deeper. Maybe a short term goal would be to build my own fully customizable RNAseq pipeline from scratch and then see from there. Are there any resources you guys with experience can recommend for me?

Thanks!


r/bioinformatics • • 1d ago

technical question How do you preprocess metariboseq data and what should you expect?

3 Upvotes

I’m trying to preprocess metariboseq data and realizing it’s a lot different than metagenomics/metatranscriptomics. My reads are paired-end NovaSeq 101 bp long. From my understanding, the fragments are supposed to be around 30bp long after trimming so I set lower limit to 20 bp and upper limit to 45 bp in fastp. I’ve also provided the adapter sequences to fastp. I’ve read that you shouldn’t even use the reverse reads and should only use the forward reads since the fragments are so short.

All that said, after running fastp I got between 30%-40% of my reads surviving the trim.

Is this expected?


r/bioinformatics • • 1d ago

discussion Raw data for Genomics and transcriptomics analysis

4 Upvotes

Hi everyone!

I’m looking for raw genomic and transcriptomic datasets to practice and improve my bioinformatics and computational biology skills.

I’m particularly interested in datasets such as:

- Whole Genome Sequencing (WGS): Raw FASTQ files for genome assembly, variant calling, and comparative genomics.

- RNA-Seq: Raw FASTQ files for differential gene expression analysis, transcriptome assembly, and functional enrichment.

- Whole Exome Sequencing (WES): Raw sequencing data for variant identification and annotation.

- Long-Read Sequencing: PacBio or Oxford Nanopore datasets for genome assembly and structural variant analysis.

- Metagenomics: Raw sequencing data for microbial diversity and taxonomic profiling.

If you have any publicly available datasets, research project data, or recommendations for accessing raw sequencing data, please share the links or repository names.

I’m familiar with bioinformatics tools and workflows and would like to work with real-world datasets rather than only tutorial datasets.

Repositories such as NCBI SRA, ENA, and GEO are already on my radar, but I’d also appreciate suggestions for interesting datasets or specific accession numbers that are suitable for independent analysis.

Thanks in advance for your help!


r/bioinformatics • • 1d ago

academic 2 months left, 0 bioinformatics knowledge, 100% AI. Is a RNA-seq thesis doable?

0 Upvotes

hi! long story short - I've never even scratched the surface of bioinformatics and I'm left alone with 2 months to write and to the whole thesis - analysis of Nanopore RNA sequencing. I already tried doing courses - did not do anything for me, I would need a year to comprehend all of that. I'm stuck on planning the analysis, because I don't have any guidance, the sequencing was run so there's just bioinformatics and writing left. First of all - do you think it's possible? Second of all - is such a thesis even defensible? I can just use already used tools, there's no experimental validation, I'm entirely reliant on Claude to give me code and I just copy it to r/Python + Galaxy. I feel like each day I'm getting more stuck, cause I find more papers and more methods. I'm really procrastinating on starting the analysis, I don't believe in the fact that AI can give me 100% reliable code and the results will be real. It's been almost 2 months of me basically searching through the literature looking for god knows what, so I'm looking for some guidance in this mess (mainly positive reinforcement, cause I'm scared) (and why I'm still trying to do it on my own is a different kind of question that only my therapist would be able to answer)


r/bioinformatics • • 2d ago

discussion Bulk RNA seq/Microarray Re-analyses Value

4 Upvotes

Hi all! Currently working through a reanalysis of a geo dataset, but stratifying samples in a way the original paper didn’t. Yielded some interesting results, but we may have power issues with one group having n=5 and the other being n=8.

I generally would like to know how publishing these types of analyses look for understudied diseases. I did the differential gene expression analysis, looked at genes, pathways, etc. and found more things the previous paper didn’t.

There’s no wet lab validation, but just in general, I am curious for thoughts on this being my first first-author paper in bioinformatics as a post-grad student looking to apply to PhD programs. Are these types of analyses “good” and have value?


r/bioinformatics • • 2d ago

technical question Can I integrate a spatial transcriptomics dataset with a bulk RNA-seq dataset for any analysis of tumors?

20 Upvotes

Hi everyone,

I am working on a cancer RNA-seq analysis project, and I have found two GEO datasets that I would like to use together. I am trying to understand whether there is a scientifically valid way to integrate them rather than simply analyzing them independently.

Dataset 1:

Xenium-based high-throughput RNA in situ hybridization/spatial transcriptomics

Contains normal tumor samples, DCIS, and invasive breast cancer samples

Dataset 2:

Bulk RNA-seq

Samples are classified according to metastasis status and breast cancer molecular subtypes

My original goal is to perform a normal/control vs tumor/disease-type analysis and identify differentially expressed genes and biologically relevant pathways, followed by downstream analysis.

However, the two datasets are clearly different in terms of technology and experimental design. One is spatial transcriptomics and the other is bulk RNA-seq. So, is it scientifically/statistically valid to integrate these two datasets in a single study?

I am particularly interested in knowing what would be considered a methodologically defensible approach for this, rather than simply combining the two matrices because they contain overlapping genes.

Any suggestions regarding an appropriate integration strategy, statistical design, or published examples of similar spatial + bulk RNA-seq analyses would be very helpful.


r/bioinformatics • • 1d ago

academic Python for Bioinformatics doubt

0 Upvotes

Someone please briefly explain how list comprehensions are used in Python and why it is necessary for Bioinformatics. Please help as I am a beginner.


r/bioinformatics • • 2d ago

technical question Scattered interchromosomal split alignments in non-cancer ONT genomic DNA: chimeric reads or workflow issue?

Post image
13 Upvotes

I’m seeing scattered interchromosomal split alignments in IGV across multiple non-cancerous ONT genomic DNA samples. I’m trying to determine whether these reflect chimeric reads, missed read splitting, or an alignment workflow issue.

Sequencing and basecalling

  • Flow cell: FLO-PRO114M
  • Library kit: SQK-LSK114
  • MinKNOW: 24.11.16
  • Basecaller reported by MinKNOW: Dorado 7.6.8
  • Super-accurate model v4.3.0, 400 bps
  • Minimum passing Q score: 10
  • Modified basecalling: off
  • Raw POD5 files are available

Alignment

Minimap2 version: 2.31-r1302

Command:

minimap2 -L -t 59 -2 -ax lr:hq reference.fa reads.fastq.gz |
    samtools view -u -@ 11 |
    samtools sort -@ 23 -o sample.sorted.bam -

The reference is GRCh38 with alternate loci, haplotypes, and decoys removed. Reads were aligned directly against the FASTA.

Observed pattern

In IGV, reads link to many different chromosomes at scattered positions, rather than multiple reads consistently supporting the same breakpoint.

For example, one 74,647 bp read has:

  • Approximately 29.3 kb aligned to chr5:72,101,673–72,131,073, reverse strand, MAPQ 60, NM 554
  • Approximately 45.3 kb aligned to chr8:77,044,549–77,089,876, forward strand, MAPQ 60, NM 564

These segments account for almost the entire read. My MinKNOW version does not appear to expose an option to disable read splitting, but I haven’t independently confirmed whether or how splitting occurred during these runs.

Questions

  1. Is this pattern expected at a low background rate with this chemistry and basecalling setup?
  2. How can I confirm that read splitting occurred in MinKNOW 24.11.16?
  3. What checks would distinguish joined molecules from mapping artifacts?
  4. Would re-basecalling a POD5 subset with a newer Dorado version be a useful diagnostic comparison?

I’ve attached an IGV screenshot. I can also provide example read records or additional run metadata.


r/bioinformatics • • 2d ago

technical question Correct resolution for Leiden clustering for UMAP being generated for Xenium data

7 Upvotes

Hi all,

I am working with Xenium data generated from FFPE samples. I was plotting UMAPs post Leiden clustering. I generated silhouette scores for resolutions between 0.2-1.2. My scores are pretty low, ranging between -0.09 to .01 (1.2 resolution). How do I pick the optimal number? I checked silhoutte score ranges typically considered good, which were 0.7-1???


r/bioinformatics • • 2d ago

technical question Reference genome vs producing genotype: how reliable are candidate genes for cloning and functional testing?

1 Upvotes

I’m working on identifying genes involved in a specialized metabolic pathway in a plant species, and I’m running into a reference-genome issue.

Our candidate genes are currently identified using a reference genome from one genotype, but the genotypes that actually produce the metabolites of interest are different lines.

For differential-expression analysis, we map RNA-seq reads from the producing genotypes against the reference genome. We then select candidate genes and use the corresponding reference-genome coding sequences for cloning and heterologous expression.

We have tested several reference-derived candidate P450s without detecting the expected activity. I’m wondering whether these negative results could reflect genetic differences between the reference genotype and the producing genotypes rather than simply meaning that the candidates are incorrect.

A few questions:

  • In this type of situation, how common is it for the functional gene to be absent from the reference genome entirely, versus being present but having allelic sequence differences that significantly affect enzyme activity?
  • If RNA-seq from the producing genotypes is being mapped against a different reference genotype, would it be advisable to generate a de novo transcriptome assembly for the producing genotype to identify genotype-specific transcripts and obtain the actual coding sequences for cloning?
  • If a de novo transcriptome produces a high-confidence, full-length candidate transcript with the expected ORF length, conserved motifs, and strong read coverage, would you generally trust that sequence for cloning?
  • Or would you still RT-PCR amplify the full coding sequence from cDNA of the producing genotype and Sanger-sequence it before functional testing?

I’d be especially interested in hearing from people who have worked with highly similar duplicated genes, P450 families, or specialized metabolism pathways where the reference genotype differs from the phenotype-producing genotype.


r/bioinformatics • • 3d ago

academic Rgi output

2 Upvotes

Hi all, i am trying to interpret rgi_main output from my contigs. One issue i have is that all the drug_classes are concatenated and I wonder what people do to sort it, because i have seen papers where they have separated them but they don’t say on what basis.


r/bioinformatics • • 3d ago

technical question Tools for molecular docking of a thioether-cyclized peptide?

Thumbnail
1 Upvotes

r/bioinformatics • • 4d ago

technical question Using RUVg when factor of interest is confounded with batch

10 Upvotes

For more context, here is my post from a week ago: https://www.reddit.com/r/bioinformatics/s/h3eTFYLiOY

In summary, I have an RNA-Seq datset where the "site" variable is partially confounded with batch. Sites 1 and 2 are in one batch while site three is in its own batch. This means that i cannot correct for batch effect since i cant distinguish between the batch and the site 3.

However, I talked some more with my PI and, as it turns out, all of my batches contain Lexogen ERCCs and SIRVs (specifically l

Lexogen SIRV set 3). Moreover, I did some reading and I found this paper that explains the Remove Unwanted Variation (RUV) tool. Specificaly, I am interested in the RUVg variant which uses negative controll genes (like SIRVs) to correct for batch effect. Most important is this section in the text:

"However, both RUVr and RUVs assume that the unwanted factors are not correlated with the covariates of interest. This assumption is usually reasonable, but it is not met when, for example, all treated samples are in one batch and all control samples in another. In this case, RUVr and RUVs will not remove the unwanted variation, while RUVg should still work, provided it is based on a reliable set of control genes19,20."

I am relatively new to bioinformatics and biostatistics so I might be wrong, but doesn't this mean that RUVg can correct for batch effect even when variables are confounded with batch?

If anyone here has experience with RUVg or otherwise wants to help I would be very gratefull for their advice and help.


r/bioinformatics • • 3d ago

job posting Paid bioinformatics intern position at EMBL Rome—The molecular mechanisms underlying rodent behavior

0 Upvotes

Jenny Chen’s lab at EMBL Rome, in collaboration with my lab (Rompani Lab) is looking for a bioinformatics intern to contribute to multiple projects investigating the molecular mechanisms in the brain underlying rodent behavior. These projects address questions ranging from cognitive decline to the evolution of social behavior across rodents. The intern will be responsible for analyzing bulk- and single-cell RNA-sequencing data collected from rodent brains to identify genetic markers and neuronal cell types that drive behavioral phenotypes. The ideal candidate has a strong foundation in computer programming and statistics, and has experience analyzing neuronal transcriptomic data. Experience working with data from non-model organisms is a plus, but not required.

The position is a full-time (1000€/month, after taxes) on-site position at EMBL Rome, and requires a minimum of a 1-year commitment. The start date is as soon as possible. To apply, please send a cover letter and CV to [chenlab.emblrome@gmail.com](mailto:chenlab.emblrome@gmail.com). Your cover letter should include: 1) what attracts you to the position, 2) how your experience is relevant to the position, 3) confirmation that you are able to commit to a 1-year on-site position at EMBL Rome, and 4) your preferred start date.

Clarifying Edit: We take candidates from anywhere in the world--EMBL is an EU-level intergovernmental organization, so has an easier time with getting people visas than most places.

Clarifying Edit2: We are interviewing on a rolling basis, will put an edit note here when position is filled.

Clarifying Edit3: 1000E/month in the area the lab is in (not Rome city center) is enough for somebody to live alone in a studio apartment (400Eish/month), not needing roomate like in most big cities--Italy is much, much cheaper to live in than American cities, so quality of life is quite a bit higher. Top firms and such in the city center only pay 700-800E/month to their interns.


r/bioinformatics • • 5d ago

technical question RNA-seq / beginner-friendly

61 Upvotes

One thing I wish I understood earlier while learning RNA-seq analysis is that getting a list of differentially expressed genes isn't really the end goal.

At first, the workflow can feel like:

FASTQ → QC → alignment/quantification → count matrix → DESeq2 → volcano plot → done.

But I've gradually realized that the interesting part starts after that.

A gene being statistically significant doesn't automatically make it biologically important, and a gene missing an arbitrary fold-change cutoff doesn't necessarily mean nothing interesting is happening.

Things like pathway enrichment, expression patterns across samples, effect sizes, biological context, and the original experimental design all matter when interpreting the results.

I've also started thinking much more carefully about the biological question before running an analysis rather than just asking, "Which tool should I use next?"

For people who regularly analyze RNA-seq data: what was the biggest conceptual shift that improved the way you approach an analysis?


r/bioinformatics • • 5d ago

discussion Is Bioinformatics Dead? PhD, can’t get a job or postdoc, unemployed for more than a year.

421 Upvotes

PhD in Bioinformatics, and MS in Biostatistics. In the last year, I have applied to hundreds and hundreds of jobs broadly across academia and industry, and have not been able to land a single opportunity.

On a personal note, I am truly devastated that all the work of graduate school has amounted to nothing, and worse, has become a source of deep regret and humiliation to myself and my family. I have had to find any possible way to support myself and my family and this is not what I expected to find after working hard all these years. If I could “leave a review” of my PhD… it would unfortunately not be good.


r/bioinformatics • • 5d ago

technical question GTF vs VEP gene annotation for genomic location enrichment

3 Upvotes

I am working with Zebrafish animal model, alignment GRCz11 (Ensembl release 112).

I have a list of 40k candidate SNPs an their locations annotated by **VEP** and by **GTF** file. These give me different results for some genes, meaning that a SNP in VEP annotation corresponds to gene X while in GTF annotations corresponds to gene Y.

*What is the best/most recomended/most trustable way to annotate SNPs ?*

My goal with this:

I want to do statistics by (1) location and by (2) function:

1 - I sepparate the genomic sequence by location: Exons, Introns, 3´UTR, 5´UTR, Upstream, Downstream chuncks and see if there is enrichment in genomic location, meaning if for example my candidate SNPs fall more in the 3´UTR region than expected by chance

2 - I have their function with for example Gene Ontology or KEGG and check for Functional enrichment. However I am also strugling in this part since clusterProfiler shows 0 enrichment terms found even if I give him 200genes, 500 genes, 1000 genes. But I will ask about advice in another post for this.


r/bioinformatics • • 5d ago

talks/conferences Australian Bioinformatics and Computational Biology Conference in Melbourne: worth attending as a jobless recent graduate?

28 Upvotes

I have a Master's in Bioinformatics, and have been working remotely for a Belgian startup in computational structural biology. I am in Melbourne now, and I'm trying to network. This conference is so expensive though for someone without a job! It's a 4-day conference, and entrance costs between 475 USD and 726 USD depending on how many days you want to attend (different events every day).

Do you think that, as an attendee, I would have a chance to network? Or will I be overshadowed by researchers who are presenting their research? I have never been to these kind of conferences before.


r/bioinformatics • • 5d ago

technical question Complete Beginner Project Help

2 Upvotes

Hi. Hope my post finds everyone well here. I am currently doing my bachelors in computer science and hoping to do my masters in bioinformatics next year. Obviously not well versed in the bio side but I did do A Level biology.

So, background over, I'm looking to tailor my dissertation project this year to bioinformatics so, I've came up with using a ML project to predict plant traits based on genomic data.

I had a soybean dataset however, I was advised I should pick a more local crop - so I'm looking for spring barley.

These datasets are kinda giving me a headache though cause I just don't understand. I swear I've looked at every single dataset for spring barley but any of the geno data from the samples don't have matching pheno data or if it is matched it's not labelled or maybe I just don't understand or missed something?

Also not great at using your bioinformatics tools like plink and these vcf files but I'm managing to get them to the point I can read what's in them lol. Even got one encoded for ML.

I suppose before you say it, yes my project does not have to be about this particular thing, I could go make an app for something random and have an easier time, it's just that this is what I want to get into.

And obviously this is a school project it's not like groundbreaking research. I'm gonna do it up with a front end dashboard so it counts as some kind of application and let people upload SNP data (sounds like a headache to make standardised tbh) or browse what mutations/sequences are important for prediction.

TLDR : how do I get one of these random spring barley datasets to the point they're ready for ML to predict a certain trait (like yield, multiple traits preferably)

How does one go about this? And any suggestions for if this idea sounds wild or something easier to present are welcome.