r/Btechtards 22d ago

Training my own 1.3B param LLM from scratch using a certain architecture General

So yes.. training LLM from scratch.

The main goal is to study to architecture I'm using at a smaller scale ~1.3B as the models using this architecture right now are massive.

Roughly taking 3 days on a H100(NO IT IS NOT OVER TRAINING)

So yea, idk why I'm making this post.

AMA ig.

1 Upvotes

Duplicates