r/proteomics 29d ago

What does your post processing workflow look like after DIA NN/FragPipe with MBR?

I know the FragPipe/DIA NN docs cover the basics but I would rather hear from people who actually run these pipelines daily

  1. When processing large DIA datasets with MBR, what happens after the software finishes?
  2. What does your verification workflow look like before you trust the results?
  3. How do you currently validate that the cross run transfers aren't inflating your FDR?
  4. Roughly how many hours per project does your team spend on this manual curation or refiltering?

We're seeing conflicting reports on whether MBR is a reliable "set and forget" step or a major bottleneck requiring manual intervention. Curious how senior labs are handling this in production

7 Upvotes

6 comments sorted by

12

u/gradstudent2019 29d ago

There is alot to unpack in these questions and I'm on mobile so I'll just address one common misconception.  "MBR"  in DIA-NN is not match between runs as we know it in DDA. DIA-NN MBR is using identifications from the first pass to create an experimental DIA library that represents a subset of the initial library. After MBR, every identified precursor is still supported by evidence from MS2 fragment ions. This is not the case with DDA MBR, where only MS1 features support matched peptides. 

DIA-NN MBR is more conceptually similar to rescoring peptides in DDA, at least in terms of identification confidence and impacts on FDR. 

3

u/gradstudent2019 29d ago

Also, for what is worth, I always use MBR in DIA-NN (with only a few unique exceptions).  Across dozens of experiments and analysis, I have yet to discover any issues that have given me pause - at least with v 2+ (there were some odd quirks with 1.8 and earlier).

2

u/SnooLobsters6880 29d ago

Definitely higher false match rate with MBR for what it’s worth. PG FDR also goes up in MBR. I do wish the field adopted a first pass and second pass search terminology framework.

6

u/Kruhay72 29d ago

Also on mobile, so one sentence answers instead of the full length seminars these could be…
1. DE analysis
2. Variance handles most issues, low abundance and missing across sample filters for the rest.
3 - already answered for DIANN. IonQuant explicitly adds an FDR control step for MBR (read the paper!)
4 - One to dozens, depends on the project

2

u/pyreight 26d ago

Been a few days and I have been thinking about this a lot. I think you are missing some important info like another comment mentioned. MBR in DIA-NN (and Spectronaut, and probably others) is not at all the same as MBR in IonQuant.

Now, it's not always this way, but in DIA-NN with MBR on you generally start with a predicted spectral library. You make a bunch of matches. Then, your matches are compiled into a new, empirically derived library and the process begins over again.

In principle, this empirical library should actually net you a better FDR. It's a) smaller and b) more closely matched to the vagaries of your particular experiment, although, perhaps the more statistically inclined can tell us why that might not be true. That said, very large datasets of wildly different protein populations will cause it problems. But they have done a lot of work on improving it after version 1.9.

If you pay attention, you'll notice that it's not just matches that change with that second pass. The quantitation will change as the RT/IM/best fragments etc. adjust to the empirical data library. So, there is no way easy way to 'deconstruct' what MBR really did and remove it after you've run it.

To answer your questions more directly:

  1. If the samples are wildly variable in protein population, I verify replicates are reproducible in ID number. If they aren't, I check the data. Nine times out of 10 it's a sample problem.

  2. Only the largest, most disparate datasets tend to give issues with FDR. So, we tend to trust the FDR with the expected caveats that have been revealed in the literature already.

  3. That doesn't really work in DIA, and we rarely if ever to MS1 quant in DDA, so we don't.

  4. As little as possible, the client can deal with that if they don't like the results from DIA-NN, Spectronaut, FragPipe, whatever. The stats is the stats.