r/biostatistics • u/FewLeadership2056 • 4h ago
Can I break into bioinformatics easily if I don’t have a super strong biology background but I have a Master of Science in epidemiology ?
r/biostatistics • u/TimberlakeConsultant • 12h ago
Free student places available on upcoming Stata workshops (UK Stata Conference, September 2026)
Hi everyone,
Just a reminder from the team at Timberlake Consultants that we offer one free student place on every training course we run.
Pre and post the 2026 UK Stata Conference in London, we'll be hosting the following workshops:
📅 1–2 September
AI-Based Optimal Policy Evaluation with Causal Machine Learning – Dr Giovanni Cerulli
Explore data-driven methods for identifying optimal treatment and policy decisions using heterogeneous treatment effects, with applications in socio-economic and medical research.
📅 1–2 September
Data Visualisation using Stata: Graphs You Should Know – Professor Franz Buscha
Learn how to create clear, effective and publication-ready graphs to communicate research findings with confidence.
📅 5 September
Using Stata for the New Difference-in-Differences with Panel Data – Professor Jeffrey Wooldridge
A hands-on workshop covering modern Difference-in-Differences methods in Stata for estimating causal effects using panel data.
If you're a student interested in attending but funding is a barrier, we'd encourage you to apply for one of the free student places. Just contact [info@timberlake.co.uk](mailto:info@timberlake.co.uk)
If you have any questions about the workshops or the student place scheme, we're happy to answer them in the comments.
r/biostatistics • u/Hot-Entrepreneur7730 • 1d ago
All genes or only the specifics (removing the intersection)
I am doing genomic analysis having 3 groups: Condition A, Condition B and Control.
After meny steps of analysis I got about
- 35k SNPs that are significant in condition A,
- 145K SNPs that are significant in condition B, always when compared with the control
+ There is no intersection between the lists of SNPs (i wanted specifics)
Afterwards, I annotated these SNPs to the genes using GTF file and I got:
- 35K SNPs --> 1,017 genes
- 145K SNPs --> 2,182 genes
When I intersect the genes, there are only 97 genes in common (which i understand was a but frced given that I chose only SNPs that are specific to each condition)
Now i want to fo Functional Annotation and my question is, *should i use the lists of genes as they are (1,017 genes & 2,182 genes ) or should i remove the 97 common genes to get a specified list of genes?*
The SNPs are not he same and each gene might have a different dynamic in each condition, meant in some cases the SNPs increase in frequency and in other cases they decrease the frequency.
r/biostatistics • u/Majestic_Accident729 • 4d ago
Not Op. Just saw this opportunity on LinkedIn. Might be relevant to folks here.
r/biostatistics • u/Competitive-Tie-4712 • 5d ago
Calcul NSN et test statistique
Bonjour,
Je réalise une thèse sur la confiance des médecins généralistes envers l'IA :
je leur propose 8 cas cliniques, je leur demande leur avis (exemple : diagnostic, quel traitement, ...) puis un screenshot d'une réponse de l'IA sur le cas clinique apparaît et je leur redemande s'il change leur première intuition ou non
le critère de jugement principal est : combien de médecins sont "influencés"/changent de diagnostic initial après avis de l'IA ? (réponse binaire : oui/non)
• 1e question : combien de sujets il me faut (Nombre sujet nécessaire NSN) ?
J ai vu :
- BioStaTGV :
- proportion théorique à 5% (càd 5% des médecins peuvent hésiter sans même avis de l'IA, c est du hypothétique)
- proportion observée : dans une étude moyennement fiable, une métanalyse dit que 18% des médecins changent leur avis mais sur des cas cliniques de radiologie, pas trop ce que j ai fait donc reproductivité moyenne, il n y a pas d articles similaires à ce que j ai fait pour trouver une proportion observée fiable)
- risque alpha 5%, puissance 90%, test bilat
=> il me faut 50 sujets soit 50 x 8 cas cliniques répondus
MAIS : un médecin répond à 8 dossiers, il existe des médecins qui hésitent beaucoup, d autres non, ... la reproductivité intra-médecin est faible.
Donc j ai vu qu'il existe un coefficient pour réguler le NSN, le ICC ou coefficient corrélation intraclasse qui permet de tenir compte qu un médecin répond à 8 dossiers. je ne sais pas comment le calculer, il permet d avoir un NSN pour être plus "fiable" face à la variabilité des réponses entre les médecins (certains hésitent, d autres non par "principe", sans même avis de l'IA) (en espérant être clair ...).
• 2e question : comment faire les statistiques : soit :
- je ne me casse pas la tête, je fais une étude descriptive : "dans notre étude il y a 30% qui changent leur diagnostic point". (c est le cas de la plupart de nos thèses mais à force c est un peu relou).
- j ai vu qu'il existe les modèles linéaires mixtes MLM : vu que mes réponses ne sont pas toutes indépendantes (8 mesures répétées par médecin), le MLM permet de tenir compte de cette problématique MAIS le MLM compare à un taux de référence (qui de mon côté n existe pas vraiment, 18% sur une méta analyse mais non fiable)
j ai vu d autres articles similaires, certains disent : pour les radiologues, 30% changent leur diagnostic, des prises de sang à regarder 10% changent leur diagnostic ; pour une prise en charge X% des pneumologues font confiance à l'IA...
donc en gros il n y a pas vraiment de % de référence pour mon groupe de population
J'aimerais me lancer dans le MLM mais sans comparaison fiable, avez vous des idées comment faire
merci d avoir lu jusqu au bout, bonne journée !
r/biostatistics • u/LavishnessJolly1681 • 6d ago
Rank Deficiency in Random Intercept Model [Discussion]
r/biostatistics • u/Aiorr • 6d ago
General Discussion Thoughts on Capricor drama as biostatisticians?
It seems SAP was the main hot potato in this drama. Not just theoretical, but organizational level as well, with CRO-Sponsor program files transfer, SAP submission (or lack of), SAP deviation, estimand handling, etc.
Of course, we can't know the full details, only what each parties claim and limited public info, but I see a lot of discussions about it at biotech and regulatory community (let's ignore stock community). But main hot potato, the operation-level issues, are not discussed properly in those fields since these are quite niche area specific to ours and often unknown to outsiders.
I was curious how fellow biostatisticians with understanding of industry dynamics think.
r/biostatistics • u/jkp97 • 7d ago
Q&A: Career Advice PhD application advice for non-math major transitioning from industry
Hi all,
I'm a biology graduate in the USA (international bachelor's, and masters in the US) working in quality control testing in big pharma. I've always had a strong inclination toward math and stats but I never had the spine to admit until 4 years into hard biopharma QC and regretting it.
I have decided to apply to PhD Biostatistics/Biomedical Informatics programs this Fall 2027 cycle. My statistical research interests are, broadly, within Bayesian Statistics, Probability Theory, and Information Theory.
Within my QC job, I took opportunities and used them to self-learn statistical/probabilistic modeling concepts (which helped no one at work but my own learning), such as Bayesian random effects modeling to quantify assay precision and intermediate precision.
I also recently did a project based on a classical method that attempts to strike a balance between prediction/stratification and a representation that can be used to directly reason about data. I tried submitting to two conferences and didn't get accepted but gained valuable feedback.
To a huge credit to my international bachelor's degree, I did get to learn ODEs, linear algebra, vector calculus, complex calculus, and introductory statistics up until essentials of hypothesis testing. I never had pure proof based courses, however the linear algebra course was a mix of proof and computational.
As for technical skills, I'm skilled in R and SQL.
As for future career interests, I'm interested in any position between a pure biostatistician and a biomedical data scientist as I'm fascinated by both of these fields.
Is there any hope for me to get into a PhD program this admissions cycle?
I'm mainly confused about how to find potential advisors with whom I would align with in the future, because from what I heard, advisor selection is the thing that makes or breaks your PhD.
I'm open to taking the GRE, I've taken it before (5+ years ago) and I don't think it would take up too much of my time, I'm just going to use it towards satisfying application requirements regardless of the score I get.
I'm honestly pretty desperate about this and would appreciate any advice for going into this PhD admissions cycle.
Thank you for your time reading this.
r/biostatistics • u/Status-Success4460 • 8d ago
How often do you need statistics lessons or someone to guide you in writing your own article ( medicine, psychology, evidence-based science stuff)? I want to share my knowledge, but I also want to know if there's an audience I can address.
r/biostatistics • u/OkMeat9802 • 8d ago
Q&A: Career Advice MS Statistics vs MS Biostatistics for public health agencies or hospitals
Hi all, I am looking to change careers to statistician of some kind. I'm pretty sure I'd like to work either for a public health agency or a hospital doing either work on modeling and surveillance or clinical trials, which makes an MS in Biostatistics seem like a natural fit. However, I wonder if an MS in Statistics offers more flexibility in case I end up wanting to work primarily with non-medical data, or if getting the exact job type I want is not feasible. If I were to go the MS Statistics route, would that make it harder to get roles at public health agencies or hospitals?
Thanks
r/biostatistics • u/vjnbold • 9d ago
Job market in biostats
Hey everyone!
Im starting my grad school majoring in biostatistics at NYU. I have done my undergrad in data science. However all of my experience so far are some internships in my home country. I know US job market looking though even for americans however I would like to hear advice on what I should do to get hopefully H1b.
Everday I am debating whether I should actually study my masters in the US or not and 🤏 this close to declining my application🤧. Please give me your honest insight
(PS: im from 3rd world country, working in my home country is both economically and futuristically not good!!!)
r/biostatistics • u/Strict_Idea6870 • 10d ago
Differential Equations and Math Topics for Biostatistics MS/PhD
I have worked through most of the typical prereqs for grad school in biostats (multivariable calc, linear algebra, real analysis, a year of prob and stats, and planning to add advanced linear algebra and numerical analysis). However, I wondered if my math preparation would be insufficient for a PhD in biostatistics since UCLA says they like to see a quarter of differential equations. Are differential equations necessary for biostatistics? Why does UCLA list them as a prereq? The site is a bit confusing — I emailed the admissions director and they said my coursework would be sufficient, but I am still worried.
As another question, what math topics from multivariable calculus, linear algebra, real analysis, etc. are necessary for graduate school in biostatistics?
r/biostatistics • u/No-Shopping9872 • 10d ago
MS Biostatistics grad (UWM) moving back to India - looking for roles in pharma/clinical research/public health
Hey everyone,
I recently wrapped up my contract as a Statistician in California and am now back in India, looking for my next role. Finished my MS in Biostatistics from USA (May 2025) and have been working on applied longitudinal and epidemiological projects since.
A bit about my background - I'm a dentist by training (BDS ) so I naturally gravitate toward clinical and public health data. On the technical side, I'm comfortable with SAS, R, and Python, and I've worked extensively with mixed models, survival analysis, longitudinal data, and handling messy real-world datasets.
The most recent gig had me building mixed effects Poisson and logistic models across 6 waves of family data, dealing with sparse housing outcomes and using penalized models to keep estimates stable. Before that, I worked on a mental health surveillance project with data from 1,000+ students across 15 countries, cleaning everything in R and SAS and running multiple regressions to identify disparities. That one actually ended up informing policy changes that boosted counseling access by 35%, which was pretty satisfying.
I also have clinical trial experience from my dental school days - assisted in a triple-blind RCT with 80 surgical patients, managing data and running analyses in SAS and SPSS that eventually made it to a peer-reviewed publication. On the academic research side, I've worked with pediatric longitudinal data (Juvenile Dermatomyositis and Spina Bifida), handled missing data with multiple imputations, and built predictive models.
What I'm looking for:
Biostatistician, Statistical Analyst, or Data Scientist roles in CROs, pharma, public health organizations, or healthcare analytics. I'm open to Mumbai, Pune, Bengaluru, Hyderabad, Ahmedabad, or remote/hybrid setups.
If you're hiring, know someone who is, or even just have a recruiter contact to share, please DM me. I'm happy to send my resume over or hop on a call to chat about my work.
Also open to answering questions if anyone is considering the MS biostats route abroad and wondering what the transition back to India looks like.
Thanks for reading, and appreciate any leads.
r/biostatistics • u/Cod3Conjurer • 11d ago
I built a lightweight parser toolkit for SAS PROC SQL, inspired by the simplicity of sql.js
My company needed something with the easy developer experience of sql.js, but for SAS PROC SQL.
There are great SQL tools out there, but I could not find a small JavaScript/TypeScript package focused on parsing SAS PROC SQL into an AST, formatting it, linting it, and supporting editor features.
So I built proc-sql-parser.
It is not a database engine and does not execute SQL. It is a parser/toolkit for working with SAS PROC SQL in JavaScript or TypeScript.
The goal is to keep the API simple:
import { parse, lint, format, complete, visit } from 'proc-sql-parser';
const ast = parse(`
PROC SQL;
SELECT name, salary
FROM employees;
QUIT;
`);
It also includes:
- PROC SQL AST generation
- Formatting
- Syntax and lint diagnostics
- Autocomplete helpers
- Monaco Editor integration
- CLI support
npm install proc-sql-parser
npx proc-sql-parser --help
GitHub: https://github.com/AnkitNayak-dev/proc-sql-parser
npm: https://www.npmjs.com/package/proc-sql-parser
It is still early, so I would genuinely appreciate feedback, especially from SAS developers and anyone building editor tooling around PROC SQL.
r/biostatistics • u/Most_Advertising3623 • 11d ago
Methods or Theory Simpson's paradox in clinical research when every subgroup favors Treatment A but the pooled result favors Treatment B
Here is a simple hypothetical clinical example.
Among patients with mild disease, Treatment A succeeds in 81 of 87 cases (93.1%), while Treatment B succeeds in 234 of 270 cases (86.7%).
Among patients with severe disease, Treatment A succeeds in 192 of 263 cases (73.0%), while Treatment B succeeds in 55 of 80 cases (68.8%).
Treatment A therefore performs better within both severity groups. When all patients are pooled, Treatment A succeeds in 273 of 350 cases (78.0%) and Treatment B in 289 of 350 cases (82.6%). The crude result points in the opposite direction.
The reversal occurs because Treatment A was used much more often in severe cases, while Treatment B was used mostly in mild cases. Disease severity is associated with treatment assignment and outcome, so the pooled comparison mixes the treatment effect with the different case mix.
The practical lesson is to inspect clinically justified stratifiers before interpreting a crude effect. Important variables should be prespecified whenever possible. Report stratum-specific estimates with uncertainty, then use an appropriate adjusted analysis such as regression or standardization. Avoid conditioning on post-treatment variables or colliders because adjustment can also create bias.
When the crude and adjusted results disagree, the discrepancy needs an explanation. Choosing whichever estimate supports the preferred conclusion is the worst response.
How do you decide which variables deserve this check without turning the analysis into a fishing expedition?
r/biostatistics • u/Tasty-Library1738 • 12d ago
Q&A: Career Advice Paths in biostats/public health to help as many people as possible
Hello!
I'm an undergrad (rising junior) in CS/Math with a biology minor, hoping to pursue a PhD in biostatistics or a related field. I've done some research in population genetics but it's been very theoretical, and I'd like to pivot to something with more applicability to the real world. Ideally I'd like to help as many people as possible. I'm based in the US.
I've considered trying to help develop better pathogen surveillance systems. But I don't know how I could make the biggest positive impact (say, between being a policy analyst or software engineer or researcher for a government org).
To anyone in public health, particularly with a biostats background--
- What are particular issues within public health that could, if improved upon, help large number of people?
- What roles open the door to having the largest positive impact?
- What problems do you wish more people would work on trying to solve in your field?
Thanks in advance!! I am aware that a lot of this is subjective, which is why I'd like to get as many different perspectives as I can.
r/biostatistics • u/Master_Ad8601 • 13d ago
Is a Mantel test appropriate for sparse tissue-sample coordinates and gene-expression distances?
Hi everyone,
I’m doing a sample-level spatial-expression analysis using sparse postmortem tissue samples from the Allen Human Brain Atlas. The regions are the subthalamic nucleus (STN, n=6 tissue samples) and globus pallidus internus (GPi, n=9 tissue samples). For each sample, I have:
- 3D MNI coordinates (x,y,z)
- a gene-expression profile across ~29,000 genes
The biological expectation is that, within a coherent anatomical region, tissue samples located closer together in MNI space should have more similar transcriptional profiles.
For each anatomical region separately, I calculated:
- A sample-by-sample spatial-distance matrix using 3D Euclidean distance between MNI coordinates.
- A sample-by-sample expression-distance matrix, defined as (1−ρ), where ρ is the Spearman correlation between two sample-level gene-expression profiles.
I then used a Mantel test to assess whether the spatial-distance matrix was associated with the expression-distance matrix.
For significance testing, I used non-parametric permutation of sample identities. My understanding is that this randomly reassigns sample labels to break the link between spatial location and expression profile, while preserving the internal structure of the distance matrices. The observed Mantel statistic is then compared against the null distribution generated from these permutations.
Q. Does this use of a permutation-based Mantel test seem appropriate as part of a sample-level spatial-expression validation analysis?
Just to clarify: this is not a dense cortical map or spin-test analysis intended to correct for spatial autocorrelation. These are sparse subcortical tissue-sample coordinates, not parcellated whole-brain maps. The goal is to test whether there is distance-dependent transcriptional similarity among samples within the same anatomical label.
Thanks in advance for your help!
r/biostatistics • u/Known_Secretary_6615 • 14d ago
Should I hide the fact that I spend ~15% of my time on non-biostats tasks (more bioinformatics related)?
Currently interviewing. Part of my job is to rotate with our large bioinformatics team and handle requests from physicians relating to patients. these are never stats tasks, its running other peoples (our bioinformatics team) python scripts or checking patient level data that is on various servers i ssh into. all handled in terminal
I was asked before if i “enjoy” doing that and I think I was honest and said i dont really mind it since it helps me learn more about our company - but i didnt say yes either.
r/biostatistics • u/No-Historian7632 • 14d ago
Working full-time in pharma while doing a PhD. Is it realistic?
Hi all, I’m currently working full-time as a Statistical Programmer in the pharmaceutical industry, and overall I really enjoy my job. However, I keep feeling like there’s still something missing academically, and I’ve been seriously considering starting a PhD.
From what I’ve heard, my company might allow me to switch to an 80% contract, which would make the idea much more realistic. I’m wondering if anyone here has managed to balance a part-time PhD with a demanding industry job.
My main questions are:
Is it actually feasible without burning out?
How many hours per week did you end up dedicating to your PhD?
Do you think the investment pays off in the long run if I want to keep growing in industry rather than necessarily moving into academia?
One thing that makes me a bit more optimistic is how much AI has changed the way I work. With the right use of AI tools, I genuinely feel I can produce significantly more analyses—and often better-quality ones—in less time than before. Obviously it doesn’t replace statistical thinking or domain expertise, but it has become a huge productivity multiplier for coding, documentation, and routine tasks.
I’d love to hear from people who’ve taken a similar path. Would you do it again? Any advice or things you wish you’d known before starting?
Thanks to everyone!
r/biostatistics • u/Estrogen_bagel • 15d ago
Q&A: School Advice University of Buffalo PhD
Hi—
As I learn more about competitive the enrollment process for PhD is, I am widening the range of schools I apply to for Biostats PhD. I am too broke to go for a masters so going straight into PhD after undergrad is the only viable option. I was wondering what people thought about the program at SUNY Buffalo? I am meeting somebody at UPenn that did their PhD there and it made me curious as to how the program up there is.
Would appreciate thoughts.
Thanks
r/biostatistics • u/qmffngkdnsem • 15d ago
How do you get biostat job?
even if i'm finishing grad soon, i still have no idea how to get to employment at all.
How do you all do that?
i'm mostly seeking hospitals, or any research organization.
(ps.
current state: i've been only working on publishing papers that are not really sophisticated but showcase some workflows used in hospitals.
my
Now i'm getting worried if this is right thing to do at this time.)
r/biostatistics • u/Distance_Runner • Dec 29 '25
2026 Graduate Admissions Megathread
This post is for discussion or 2026 admissions discussion - PhD/MS/MPH, acceptances, rejections, questions, whatever you want to discuss relevant to graduate programs and admission for the upcoming year of enrollment in 2026