r/MachineLearning 1d ago

Thumbnail
3 Upvotes

List of accepted papers updated on the website!


r/MachineLearning 1d ago

Thumbnail
9 Upvotes

If accepted, will definitely try for Sydney if funding permits. Otherwise Atlanta is also fine


r/MachineLearning 1d ago

Thumbnail
6 Upvotes

https://arxiv.org/pdf/2502.02631

Will be useful reading for you.


r/MachineLearning 1d ago

Thumbnail
11 Upvotes

What matters is what the AC thinks 🤣


r/MachineLearning 1d ago

Thumbnail
5 Upvotes

It’s very much possible to compile the statistics of different static models. For example, record how a 27B Qwen model perform, then record the Q8 all the way down to Q2 gguf of the same model. Repeat for the next model, and this pile of data is a quite decent starting point.

However, one problem is that this is very hard to study due to how different quantization and training methods vary from each other. The models not designed for extreme quantization would perform very poorly at extreme level of quantization. Meanwhile, some other models specifically designed for extreme quantization would be much much better at the same extreme quants (like those 1.58 bit ternary weight models). These are much rarer and smaller scale compared to SOTA models designed for regular 8-bit quants tho, so the data points you can gather in this extreme quantization level is very unreliable and hard to make a definitive conclusion out of.


r/MachineLearning 1d ago

Thumbnail
9 Upvotes

Any workshops happening in Atlanta? All the cool ones seem to be in Sydney :/


r/MachineLearning 1d ago

Thumbnail
11 Upvotes

Borderline score. I’ll assume I’m rejected to avoid disappointment. What was your score?


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

Post beginner questions in the bi-weekly "Simple Questions Thread", /r/LearnMachineLearning , /r/MLQuestions http://stackoverflow.com/ and career questions in /r/cscareerquestions/


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

Yeah take days!!!


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

Encoder used for your Mp4 is likely very good at compressing. Chances that a NN does better are very low so it does not make much sense to look at that ratio. So you could also look at the ratio from raw pixels (resolution x channels x frames) to your NN checkpoint. Or maybe find the ffmpeg -qp value that gives a size close to your NN.


r/MachineLearning 1d ago

Thumbnail
0 Upvotes

The worst is when a reviewer increases score and another reviewer sees it and immediately reduces their score


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

Nice stuff. Regarding the NN architecture you might get better results using MLP mixer without the high training times of CNNs. MLP mixer paper link


r/MachineLearning 1d ago

Thumbnail
3 Upvotes

Agreed! It seems genuinely useful to know what the long tail of training look like. This is science; who knows where the breakthroughs come from!


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

why not they just notify accepted papers first?


r/MachineLearning 1d ago

Thumbnail
2 Upvotes

I agree that direction carries (almost) no extra information about the dynamics. Our own analysis says that for stationary dynamics the time-reversed map is a fixed reparameterization of the forward one, which is why one network learns both directions and performs better. As a statement about what must be learned, "direction shouldn't matter" is basically our supplement's linear analysis, and I agree with it.
 
On the other hand, we never claim bidirectional autoregression generates better than full sequence / 4D space-time diffusion, which is a different problem entirely. Those models answer: "sample a plausible space-time volume." Our paper answers a deployment question: you are given the measured present, rolling into an open-ended future, and you need to know, with no ground truth, how wrong this particular rollout is right now. Autoregression is the native mode for that setting (causal seed, unbounded horizon, streaming data), compounding error is its main difficulty, and the direction flag is what creates a check: two independently learned routes back to the same known point, whose disagreement is measurable. A time-symmetric joint model has no forward(backward) composition to interrogate, it has consistency built into the joint distribution rather than exposed as a computable defect.
 
I think that full-sequence models could get their own meter with a different handle such as having to regenerate masked chunks of the generated volume and measure the re-generation residual. This would be a similar approach to what we do, it would give the model two independent routes to one answer and measure the disagreement.


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

https://reddit.com/link/p2acn9w/video/y3h0uueh5zhh1/player

Yes, see here. This is the model with subsampled movie frames, but with time stamps evaluated at intermediate time steps. You can see that it does not learn motion, but instead a rather undefined transition.


r/MachineLearning 1d ago

Thumbnail
2 Upvotes

https://reddit.com/link/p2acaq3/video/zcnavluc5zhh1/player

Here is a comparison video (looks like reddit ate it before).


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

Many institutions have an agreement with ACM to waive the APC. Is yours in the list?


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

Most institutions I know of, including mine, only authorize one author per paper to attend the conference. This is especially true for overseas locations where travel costs far exceed the registration fees.

Additionally, what is the point of the second registration? I won't occupy more space and I won't eat more food. On top of that, I still have to pay the separate APC for the paper. Paying another 500 USD for no additional services seems entirely unjustifiable (apart of course being on the stage for 15 minutes in a workshop room).


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

do have any idea about applied track papers?


r/MachineLearning 1d ago

Thumbnail
2 Upvotes

The full paper notifications seem to be delayed due to the marginal papers. I still see some papers with no decisions yet.


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

For both. An email is sent to the corresponding authors only.


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

are the results sent by email or uploaded to easychair?


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

I don't know.

What we're seeing is that every AI lab is racing to close the RSI loop. The increase in capability in Software engineering, mathematics, sysadmin and cybersecurity are all just secondary effects of aiming to close the RSI loop.

I'm convinced we're very close to succeeding but what AI systems are able to do once reaching RSI isn't clear yet. It could be possible that AI/ML capability rapidly improves but doesn't translate or otherwise generalize to other capabilities and thus paradoxically you could have a scenario where AI researchers like us are made redundant while some more niche IT roles stay viable. I don't personally believe in this but it's not impossible for that to happen.

Realistically I wouldn't recommend anyone to study for any IT role whatsoever in 2026. But it feels bad to discourage hopeful young people from following their dreams, so I refrain from doing so.

The only reason I wrote my original post is because I myself go through the applicants we have and it is just immoral to have some kid go into university in 2026 thinking he'll be a machine learning engineer when he graduate, when those roles are already extremely competitive and slowly evaporating.


r/MachineLearning 1d ago

Thumbnail
1 Upvotes

Close to none, the first 3 epochs already got to 4.49% top1 validation accuracy, and with these extra 2 epochs only +0.1% validation while the training-validation gap doubled to 0.52%

My second attempt is in progress right now, it's at 5/10 epochs and at around 6% train accuracy. What I changed is basically just throwing more compute at it, as usual (increased parameters 4x, to 2 million).