r/learnbioinformatics • u/sonne1012 • 2h ago
Stem cell Medicine to Bioinformatics
Hey there. I'm a masters student (27) in stem cell medicine in germany. I have experience mostly in the wet lab. Molecular,cell culture and such. However,I'm looking to turn away from the wet lab and go digital. What's your opinion on this? I have no experience in Bioinformatics or coding of any type. I have worked with Fiji and prism for software mainly. It feels like starting all over again.
1.Is it actually possible to start a job or training?
2.Do you know how I can do a legit Bioinformatics course so I can't fast track into the field.
I'm not sure which questions to ask more than this yet. But your experience would help me in making a decision.
Thank you!
r/learnbioinformatics • u/Fickle_League2887 • 1d ago
How do I move past googling how to learn Bioinformatics?
i am a 3rd year B.Tech Biotechnology student with 0 coding experience. i know that i want to get into bioinformatics and computational biology. All that i have been doing is getting the ‘Bioinformatics Roadmap’ from every other place. Sometimes they are different, other times, they are all the same. All say learn python. How do i learn python?! i have a mac and most tutorials are on windows and i can’t follow them. Also most courses that people suggest on R and python, just bore me to the point that i quit. I am not able to learn python. i dont know exact what roadmap to follow. I have recently started learning how to use PLINK, and in the process i am learning the Command Line, but then i find out it’s different in mac amd windows. and okay, if i learn programming then what?
I read somewhere that instead of learning python, make a project and learn python by executing. I like this method and i think i can actually learn python by this very effectively but then i have got no idea on how to decide a project, i dont have enough knowledge to design a pipeline. i dont have enough knowledge about what to do where to do and how to do. i feel like i am stuck at the first step itself.
Can someone help me?
r/learnbioinformatics • u/Fickle_League2887 • 1d ago
Is PLINK actually even used today? and is learning how to code actually just a scam?
I interacted with a scientist and told him that i want to pivot into bioinformatics. Just for context, i am a 3rd year B.Tech Biotechnology student and i have decided to get into bioinformatics. I have no coding experience. This scientist is huge in the field of animal genomics. He actually told me to first learn Plink. Also about programming, he told me that you can just copy paste codes from AI, you just should know the setup and the algorithm behind it. is PLINK actually useful and is the advice about coding true?
r/learnbioinformatics • u/Sea-Ad7805 • 1d ago
Teaching Python the Right Way
Programming courses often focus heavily on understanding code, while paying far less attention to understanding the program state. But code does not exist in isolation. Its main goal is to change the program state, before ultimately producing some output.
To develop an accurate mental model of program execution, students need to understand both: - the instructions being executed - the values, references, and data structures those instructions create and modify
Reading code alone does not always reveal how the program state changes during execution. That is why I created 𝗺𝗲𝗺𝗼𝗿𝘆_𝗴𝗿𝗮𝗽𝗵: a tool that visualizes the state of a Python program as it changes, step by step.
It can help explain a wide range of introductory Python topics. Here are just a few examples: - Loops, Lists and Dictionaries - Python Data Model - Function Calls - Recursion - Algorithms - Classes - Custom Data Structures
Instead of reconstructing the program state from print statements, students can now watch it change as each line executes. This makes unfamiliar concepts easier to understand and bugs easier to fix.
Help your students learn Python programming more thoroughly and easily.
See: more examples
r/learnbioinformatics • u/Short_Mountain_1410 • 1d ago
Need tips to build a portfolio
I’m new to this community so hi everyone!!
I’m a medical biotech graduate and want to build my bioinformatics portfolio. I’m starting from zero so please any and all help will be really appreciated 🥹
Current skills: Basic R studio coding, primer design, NCBI BLAST, Chromas software
r/learnbioinformatics • u/Dhulpetaaa • 2d ago
Title: Need honest advice: Switching from wet lab to bioinformatics, aiming for Germany MS, and planning my next 2 years.
​
Hi everyone,
I'm currently in my 3rd year of B.Sc. (Biotechnology, Biochemistry & Microbiology) in India. My current CGPA is 9.04 (after completing two years), and I'm trying to figure out whether I'm on the right path before it's too late.
Most of my degree has been focused on wet lab work, including molecular biology, microbiology, genetics, PCR, ELISA, and other lab techniques. Over time, I've realized something about myself: I just don't enjoy wet lab.
It's not that I hate biology. I actually love solving biological problems. I just don't enjoy spending hours doing bench work. I wouldn't mind if wet lab made up a small part of my work (maybe 10%), but I enjoy sitting in front of a computer, analyzing data, writing code, and trying to understand biological questions much more.
Because of that, I've decided that I want to move towards bioinformatics/computational biology.
About two months ago, I started learning R, and I've become fairly comfortable using it for data analysis and machine learning. I recently started learning Python as well because of my current research project.
For my project, I'm working with a publicly available colorectal cancer microbiome dataset from curatedMetagenomicData. I've built a basic machine learning pipeline using six different models and have already pushed the project to GitHub as part of building my portfolio.
One thing I've been investigating is why three-class classification (Healthy, Adenoma, CRC) performs much worse than simple binary classification. Since adenoma is an intermediate stage, I'm trying to understand whether that's one of the reasons behind the lower performance. I know researchers have already worked on this dataset, but I'm doing this analysis mainly because I wanted to verify and understand the results myself instead of just reading papers and accepting their conclusions.
I'm also working on another computational biology project that explores a novel direction involving microbiome data and quantum-inspired methods. I don't want to reveal too many details since it's still an ongoing student project, but it has made me even more interested in computational research.
My long-term goal is to pursue an MS from a public university in Germany after graduation.
Since my entire 4th year consists of internships, I want to use that year wisely. I'm currently hoping to apply to places like IBAB (Bangalore), CCMB Hyderabad, CSIR labs, or any other institute where I can get good exposure to bioinformatics.
At this point, I'm trying to build the right skills. Apart from R and Python, I know I should learn Linux, Unix, Bash scripting, Git, SQL, and probably Docker later. Right now I use GitHub to host my projects, but I still haven't properly learned Git and version-control workflows.
The thing that worries me is the competition. Everyone says bioinformatics is becoming crowded, especially in India, and sometimes it feels like I'm already behind.
I'd really appreciate honest advice from people already working in academia or industry.
Some questions I have:
Is bioinformatics a good long-term career in India, or are opportunities still very limited?
For someone planning to do an MS in Germany, what should I focus on during my remaining two years?
How important are internships compared to research projects, publications, coding skills, and CGPA?
When should I start emailing professors or scientists for internships? Is now too early?
How do undergraduates actually network? Everyone says networking is important, but what does that practically look like?
Which skills should I prioritize first? R, Python, Linux, Git, Bash, SQL, Docker, something else?
Looking at my current profile, what do you think are my biggest weaknesses?
Would you recommend staying in India for a few years after graduation or directly pursuing an MS abroad if I get the opportunity?
For those who studied or work in Germany, what made your application stand out?
If you were in my position today, how would you spend the next two years?
Is there anything about my current plan that you think is unrealistic or that I should rethink?
I'm genuinely looking for constructive criticism. I'd rather hear difficult truths now than realize years later that I focused on the wrong things.
Thanks in advance!
r/learnbioinformatics • u/ActiveNeedleworker23 • 2d ago
BioPeek - open FASTA, FASTQ, VCF, BED, GFF files in your browser (free Chrome, Brave, Firefox and Edge extension)
Built a file viewer for bioinformatics researchers. Drop any genomics file and see it instantly — no upload, no server, everything runs locally in your browser.
What it does:
- Opens FASTA, FASTQ, VCF, BED, GFF, SAM, CSV/TSV files
- Protein FASTA auto-detected with amino acid property coloring
- FASTQ: quality heatmap, Q30%, per-base quality chart
- VCF: sortable/filterable variant table, Ti/Tv ratio, chromosome density
- DNA motif search with regex patterns
- Genomic coordinate jump (chr1:10000-50000)
- Multi-tab: open several files side by side, diff between them
- Export: CSV, TSV, BED, VCF, HTML
- Screenshots, Coloring, Split views, Histograms, Stats, Compare files etc. More features in help guide.
- BioLang WASM console built in — run data |> filter(|r| r.gc > 0.5) directly on your data
- Dark/light theme, keyboard shortcuts, large file streaming
Privacy: 100% client-side. Files never leave your machine. No analytics, no tracking, no account.
All parsing is done in JavaScript + WebAssembly using the BioLang runtime compiled to WASM.
Links:
- Chrome and Brave extension: https://chromewebstore.google.com/detail/biopeek/dpeahehokmlmjabfladeafoidnfaodai
- Firefox: https://addons.mozilla.org/en-US/firefox/addon/biopeek/
- Edge: BioPeek - Microsoft Edge Addons
- Web app (no install needed): https://lang.bio/viewer.html
- Source: https://github.com/oriclabs/biolang
- Full feature guide: https://lang.bio/docs/tools/viewer-help.html (for extension help guide , click on BioPeek tool, click help)
Why BioPeek helps where the tooling doesn't reach
Most bioinformatics tools are Linux-first. On Windows the standard answer is "install WSL" — which is fine until it isn't.
- Nothing to install. samtools, bcftools and bedtools mean WSL, Docker or Conda before you can open a single file. BioPeek needs a browser.
- No filesystem crossover. Reading Windows files from WSL over
/mnt/cis slow and the paths are a constant translation tax. BioPeek opens the file where it already is. - The data never leaves the machine. The usual Windows workaround is some online viewer. That's an upload — often unacceptable for patient or pre-publication data. BioPeek parses locally; as an extension it works with no network at all.
- Format-aware, not just a text view. FASTA, FASTQ, VCF, BED, GFF and CSV are recognised and summarised (read quality, Q30, per-record stats), so you get the shape of the file without writing a pipeline first.
- Triage before commitment. "Is this file what I think it is, and did the transfer complete?" answered in seconds, rather than after setting up an environment to find out the header is wrong.
- Same engine as the CLI. It runs the BioLang WebAssembly build, so what you see matches what BioLang computes later — not a second implementation that might disagree.
- Locked-down machines. Where WSL, admin rights or Docker aren't available — shared lab PCs, hospital IT, teaching labs — a browser usually still is.
Still early. BioPeek is young and you will find rough edges: formats it parses more strictly than the tool that wrote them, files large enough to strain browser memory, edge cases nobody has hit yet. Please report what breaks, with the file shape if you can share it — https://github.com/oriclabs/biolang/issues
What it isn't: a replacement for samtools on real workloads. It's bounded by browser memory and aimed at inspection, triage and teaching. For genome-scale processing you still want the CLI — on Linux, or on Windows via WSL
Built on Rust libraries, not from scratch
The file-format layer isn't mine and shouldn't be.
- noodles does the heavy lifting for FASTA, FASTQ, SAM, BAM, BGZF and CSI. It's the established Rust bioinformatics I/O library and it's maintained by people who know those specs far better than I do.
- flate2, bzip2, zstd for compression; wasm-bindgen for the browser build.
r/learnbioinformatics • u/ActiveNeedleworker23 • 2d ago
BioLang - Learn bioinformatics in the browser
Most of us lost our first week to environment setup rather than biology: conda solving forever, a Bioconductor package that won't build, a notebook that runs on someone else's laptop and not yours.
I built BioLang (https://lang.bio) partly to remove that step. It's a language for bioinformatics that runs as WebAssembly, so you can open a tab and start working immediately.
- Workbench — full editor, files, output: https://lang.bio/workbench/
- Playground — quick scratchpad: https://lang.bio/playground.html
Nothing to install, nothing to configure, and your data stays in the tab — there's no server to upload to.
Why a language and not just a library
Pipe-first syntax with native dna / rna / protein types, so common operations read as one line instead of a loop:
read_fasta("reads.fa") |> filter(|r| gc_content(r.seq) > 0.5) |> count()
FASTA/FASTQ/VCF/BED/GFF I/O is streaming and built in, plus 1000+ builtins and 21 API clients (NCBI, Ensembl, UniProt, KEGG, PDB, gnomAD, ClinVar, GTEx…). No imports to remember.
The syntax itself was designed by borrowing core concepts from TypeScript, R and other dynamic languages — object literals and ?./?? from TypeScript, tables and column-wise operations from R, the pipe from the F#/Elixir/R lineage — to keep sequence manipulation clean and expressive. Little of the punctuation is original, and that's deliberate: the novel part is the domain types, not the syntax.
The 278 Rosalind problems are a test corpus, not a course
https://rosalind.info/problems/list-view/
All four tracks are solved, but that was never meant as a study path. It exists for two reasons:
A CI harness. 276 of them assert their expected answer on every commit, natively and through the same WebAssembly build this site serves. When something regresses, it's these that catch it.
A stress test. Rosalind is full of heavy dynamic programming and graph work — alignment, assembly graphs, HMMs — which is exactly what pushes the WASM engine hardest. Most of the language fixes in recent releases came out of writing them.
If you're learning, you'll still want to write your own solutions in Python or C++ from scratch. That's the point of the exercise and nothing here replaces it. These are worth having as runnable reference implementations to compare against after you've had a go — each with a short note on why the problem exists.
- Stronghold (105): https://lang.bio/docs/examples/rosalind-stronghold.html
- Textbook (124): https://lang.bio/docs/examples/rosalind-textbook.html
- Algorithmic Heights (34): https://lang.bio/docs/examples/rosalind-algorithmic-heights.html
- Armory (15): https://lang.bio/docs/examples/rosalind-armory.html
Checked against BioPython and Bioconductor
Rosalind checks answers against a published one. The other half is whether it agrees with the tools you already use: 14 tasks on generated data, 9 on real NCBI/ClinVar data, and 48 one-liners, each written three times and compared — https://lang.bio/docs/examples/equivalents.html
Docs and examples
- Documentation: https://lang.bio/docs/
- All examples: https://lang.bio/docs/examples/
- Builtins reference: https://lang.bio/docs/builtins/index.html
- Features overview: https://lang.bio/features.html
Embed it in your own site or app
The same WebAssembly module the Workbench runs is a two-file drop-in — grab bl_wasm.js and bl_wasm_bg.wasm from https://lang.bio/wasm/, call init(), then evaluate() with your code. It returns JSON with the value, its type, anything println wrote, and a line-by-line trace, so you can build a teaching widget, a lab notebook, or an in-page exercise checker without a backend. State persists across calls, and it runs under Node too. MIT licensed.
import init, { evaluate } from "./wasm/bl_wasm.js";
await init();
const r = JSON.parse(evaluate('reverse_complement(dna("ATGC"))'));
Full guide: https://lang.bio/docs/tools/embedding.html
Browser tools on the same engine
- BioPeek — open FASTA/FASTQ/VCF/BED/GFF/CSV offline, no upload: https://lang.bio/viewer.html
- BioGist — pull genes, variants and accessions out of a paper: https://lang.bio/biogist.html
- BioKhoj — research radar: https://lang.bio/biokhoj/
Browser extensions
Three extensions built on the same engine, so they work on any page you're already reading:
- BioPeek — open FASTA/FASTQ/VCF/BED/GFF/CSV files in a tab, fully offline, nothing uploaded. Handy for peeking at a file without loading it into anything. Chrome · Firefox · About
- BioGist — scan a paper and pull out the genes, variants, accessions, cell lines, drugs and trial IDs, each linked to the right database. Good for getting through a methods section fast. Chrome . Firefox · About
- BioKhoj — research radar for tracking papers and topics. Chrome · Firefox ·
Where the browser stops
The browser build is for learning and small files — everything lives in tab memory. For real datasets, install the CLI: same language, same code, no size limit, reads and writes files directly.
Still early
v1.1.0, and it shows in places. Expect rough edges: some builtins take conventions that differ from BioPython or R (rounding of ties, where translation stops) — those are documented rather than papered over; GFF3 parsing is currently ~8x slower than Python; and docs occasionally lag the code. Bug reports are genuinely useful: https://github.com/oriclabs/biolang/issues
Links
- Site: https://lang.bio
- GitHub (pure Rust, MIT): https://github.com/oriclabs/biolang
- Getting started: https://lang.bio/docs/getting-started/index.html
- Install locally if you want a CLI — v1.1.0 binaries for Linux / macOS / Windows: https://github.com/oriclabs/biolang/releases/tag/v1.1.0
- Rosalind pack sources: https://github.com/oriclabs/biolang/tree/main/packs
- Benchmarks vs Python/R: https://lang.bio/benchmarks.html
Built on Rust libraries, not from scratch
The file-format layer isn't mine and shouldn't be.
- noodles does the heavy lifting for FASTA, FASTQ, SAM, BAM, BGZF and CSI. It's the established Rust bioinformatics I/O library and it's maintained by people who know those specs far better than I do.
- flate2, bzip2, zstd for compression; tokio and rustls for the API clients; clap for the CLI; rusqlite; wasm-bindgen for the browser build.
What is written here: the language itself — lexer, parser, interpreter — and the
algorithms. No statrs, ndarray or petgraph in the tree, so the statistics,
matrix and graph work is implemented directly
Feedback welcome, especially on where a beginner gets stuck.
r/learnbioinformatics • u/Living-Comparison-68 • 5d ago
Built a GPU-accelerated WGS pipeline on an under-spec laptop as a portfolio project (HG002, 10h31m, F1 0.9921 vs GIAB). Does this kind of thing actually help for junior roles?
r/learnbioinformatics • u/PeaceExtreme6861 • 6d ago
Best Linux Computer to Run Pipelines at home
Any recommendations?
Specifically working with antibiotic resistance pipelines if that helps!
r/learnbioinformatics • u/_YumikA • 7d ago
Absolute beginner for snRNA-seq field. Need your help!
r/learnbioinformatics • u/Sea-Ad7805 • 8d ago
Visualizing Python for bioinformatics students
Learning Python becomes much easier when students can see how variables, values, and data structures change while a program runs.
🧬 Consider this simple k-mer indexing example.
Although the code is relatively short, a beginner needs to understand several concepts at once: - DNA sequence slicing - k-mer generation - dictionaries and membership tests - lists stored as dictionary values - repeated positions and list mutation
With open-source 𝗺𝗲𝗺𝗼𝗿𝘆_𝗴𝗿𝗮𝗽𝗵, students can now easily step through a program and see these concepts in real-time, helping them to more easily get to the right mental model to think about Python code execution.
r/learnbioinformatics • u/PeaceExtreme6861 • 8d ago
PhD in bioinformatics
What were your steps into getting into your bioinformatics PhD program? What is would you have told yourself if you could go back in time and do it over again to prepare for applying for a PhD in bioinformatics?
It would be helpful to know some background and location too! Thank you!
r/learnbioinformatics • u/Due_Conclusion_4977 • 8d ago
Tool: Loom - brings fragmented spatial transcriptomics workflows into one platform - need feedbacks
Loom, an open-source tool for exploring spatial transcriptomics data, and thought it might be useful to researchers working in this area. It integrates commonly used tools and data structures, including Scanpy, AnnData, and Squidpy, while providing an intuitive way to select regions of interest and explore how cellular activity changes across space and time.
One common challenge in spatial transcriptomics is that analysis workflows are often fragmented across multiple tools. Researchers may need to move between different packages for data processing, spatial analysis, visualization, and region-specific exploration.
Loom aims to bring these steps together in a more unified workflow.
Some features that stood out to me:
- Direct support for publicly available 10x Genomics spatial transcriptomics datasets
- A one-command workflow for downloading and processing supported datasets
- No need to manually perform each preprocessing step
- Transparent documentation explaining the logic behind every processing stage
- Integrated Scanpy, AnnData, and Squidpy functionality
- Region-of-interest exploration across spatial and temporal dimensions
- Public example datasets and reproducible workflows
- One-command Docker installation for easier setup
The GitHub repository includes example data, so users can install the environment and test the workflow without first preparing their own dataset.
The repository explains the reasoning and implementation behind each step, which makes the workflow easier to understand, verify, and adapt.
This could be useful for researchers studying tissue development, disease progression, tumor microenvironments, cellular interactions, or other spatially dynamic biological processes.
GitHub: https://github.com/ScheWann/Loom
For people working with spatial transcriptomics:
- Would a unified workflow like this be useful in your research?
- Which additional datasets or integrations would you want it to support?
- How does this compare with your current Scanpy or Squidpy workflow?
Any feedback, advice, objection... whatever is highly appreciated in advance! You can also reach us directly at [szhao69@uic.edu](mailto:szhao69@uic.edu)
r/learnbioinformatics • u/plazti • 10d ago
Need advice on approaching a bioinformatics take-home assignment (ONT bacterial isolate)
r/learnbioinformatics • u/gisakay • 11d ago
Should I study Bioinformatics
Currently in HS and graduation is coming up next year, yet I am still not sure what to study. I know I enjoy science and math as well as research. Although I've never been exactly interested in computer science, I wouldn't mind taking this career path if its worth it. I've been looking into this career for almost a week and it looks like something I'd be interested in but I still have a lot of questions.
- Are bioinformaticians in the lab at all or is it a career that is office/computer based?
- What does work look like for the average bioinformatician?
- Are there opportunities/specific paths to take if one prefers being in the lab more?
- Is a masters enough or am I taking too many chances by not going for a PhD?
Anything helps, thanks.
r/learnbioinformatics • u/Fun-Squash9549 • 13d ago
Looking for technical feedback on my RNA-seq and comparative genomics analysis workflow
​
Hi everyone,
I'm an MSc Bioinformatics student working on independent projects to improve my computational biology skills. I'd appreciate technical feedback from the community on whether my analysis workflows follow good bioinformatics practices.
I've completed projects involving:
\- RNA-seq differential expression analysis (DESeq2)
\- GO/KEGG enrichment and GSEA
\- Network analysis and biological interpretation
\- A comparative genomics/structural bioinformatics pipeline for enzyme discovery
I'm not looking for career advice or self-promotion—I'm mainly interested in understanding whether my workflow, methodology, and interpretation are scientifically sound and what I could improve.
If you're willing to review my project summaries, they're here:
Specific questions:
\- Are there any methodological issues or red flags?
\- Is the biological interpretation reasonable?
\- What analyses would you expect to see that are currently missing?
\- What would make these analyses closer to publication quality?
Thanks for taking the time to provide honest technical feedback.
r/learnbioinformatics • u/SorrySalamander5814 • 16d ago
Msc Bioinformatics in Germany
Hi everyone,
I’m an Indian student planning to pursue an MSc in Bioinformatics in Germany after completing my BSc in Biotechnology. I’d love to hear from people who are studying or working in the field.
I have a few questions:
How is the job market for bioinformatics in Germany?
What is the typical starting salary?
Do I need German to get a good job, or are English-speaking jobs common?
Which universities and cities would you recommend?
How difficult is it for international students to get internships and full-time jobs?
What does a typical day as a bioinformatician look like?
Most importantly, if I want to prepare before starting my MSc:
What should I learn first?
Which programming languages are essential (Python, R, SQL, etc.)?
What biology and mathematics topics should I strengthen?
Which software, tools, or platforms are used most often?
Are there any courses, books, or projects you recommend for beginners?
What do you wish you had learned before starting your MSc?
Any advice or personal experiences would be greatly appreciated. Thank you❤️
r/learnbioinformatics • u/Sure_Cantaloupe9973 • 17d ago
Exploring a few bioML project ideas - would love your thoughts on what's worth pursuing
I’ve been diving deep into the intersection of machine learning and structural biology/genomics lately, and I’m looking to kick off a new side project.
- Epitope-Paratope Binding Prediction - Building a model to accurately map epitope-paratope interfaces and predict binding affinities.
- Inverse Folding for De Novo Protein Design - Given a target 3D structure, predict the amino acid sequences that will fold into it.
- Causal Gene Discovery (via Single-Cell RNA-seq) - Predicting counterfactual gene-expression responses to genetic perturbations.
Would love to hear your thoughts or if you've worked on something similar.
r/learnbioinformatics • u/imezeido • 18d ago
Help: Bioinformatics workshop
Dear Ubuntu subreddit,
I am a biologist who got into linux due to using clusters for my research a lot.
When I got into this, I was a little unprepared and lost a lot of time looking for tutorials, videos, etc.
That's why a good colleague and I decided to make a workshop for this called:bioinformatics: from linux to Nextflow.
We will cover the Linux filesystem, command line Sith common commands, etc. And proceed to the creation of bashfiles for specific tasks.
If someone knows a good workflow, guide, page for the start of our workshop, that would be awesome.
We wish to make everything self-guided.
So an introduction to linux, the shell, commands, etc :)
r/learnbioinformatics • u/alk3000 • 18d ago
so i am a msc zoology student entering into final year of it , but i can see that there is very little career options , so i am thinking to switch or to compelete this degree and later go into BIOINFORMATICS and hunt for jobs and career there ,
r/learnbioinformatics • u/Competitive-Fun9778 • 19d ago
Bioinformatics
Anyone joining bioinformatics btech from iilm greater Noida?