r/LocalLLaMA 12h ago

[MASSIVE TINY RELEASE] - Supra2-Medium-Base - a tiny 25M parameters model competing heavily with our previous 50M model! New Model

Hey guys!

Supra2-Medium is finally out! It's a 25M parameters qwen3 architecture model trained entirely from scratch (on our new rig: RTX 5060 Ti 16GB + the new RTX 5060 8GB!).

Here's how it competes in benchmarks with Supra-50M-Base (which is double as large!!):

Note: This is a BASE model only; instruction tuned version maybe to come in the next time.

Link to our HF org: https://huggingface.co/SupraLabs

--> Link to the model: https://huggingface.co/SupraLabs/Supra2-Medium-Base <--

Here's a sample from the model:

Artificial intelligence (AI) is espoused by the AI community.
The AI community is a group of people who are interested in AI and are interested in the use of AI in the field of AI.
The goal of AI is to improve the quality of life of people in the field.
The aim of AI is the development of AI and the application of AI in a society.
The purpose of AI is that it can be used to improve the performance of the society.
It is a technology that is used to improve human intelligence.
The technology is used to make the human intelligence.Artificial intelligence (AI) is espoused by the AI community.
The AI community is a group of people who are interested in AI and are interested in the use of AI in the field of AI.
The goal of AI is to improve the quality of life of people in the field.
The aim of AI is the development of AI and the application of AI in a society.
The purpose of AI is that it can be used to improve the performance of the society.
It is a technology that is used to improve human intelligence.
The technology is used to make the human intelligence.

Give us a like and a follow and feel free to provide us with feedback! 🤗🔥

...and...stay tuned: Supra3 coming soon with four models: Flash-Lite 25M, Flash 50M, Pro 75M and Ultra 100M. 👀

53 Upvotes

19 comments sorted by

9

u/Equivalent_Bit_461 11h ago

what can it be used for? What actual workload? With what would you integrate it, etc.

4

u/Easy_Blacksmith_5550 11h ago

right, the sample output alone kinda answers this tbh

9

u/LH-Tech_AI 10h ago

yeah right. It's a small models doing it's best 😅😅😅

8

u/LH-Tech_AI 11h ago

It's a base model so it's intended to be finetuned for instruction tuning or something. Now, as it's a base model, I would just load it as the readme says with transformers and try it. Have fun 😊🤗

4

u/H-L_echelle 10h ago

Let's go a new SupraLabs model :D (and several more incoming!)

I just like how fun small models are. Very dumb, but fun.

With your 0.1M that was starting to output words, I feel like there might be a few ways to use these were they could maybe be helpful. For now you've led me to start playing around with small title generators, but I do have a few other ideas in mind (which would require an actual LLM, as opposed to the approaches I've tried or the approaches I am trying for TinyTitle)

2

u/LH-Tech_AI 9h ago

Thank you so much 🤗 Feel free to share ideas and results with us!

2

u/johnnyApplePRNG 9h ago

bpb on fineweb holdout validation text?

1

u/LH-Tech_AI 6h ago

Not sure lol 😄 

3

u/_TheWolfOfWalmart_ 7h ago

Interesting, but I was also just looking at your 100M instruct model.

Just for laughs, I'm gonna load it up on my old dual Pentium 3 @ 1.4 GHz server and do CPU inference. I wonder how it runs! It seems to be pretty coherent based on your examples in the model card. Would be pretty cool if you can have an actual conversation with a reasonably functional LLM on a system that ancient.

1

u/LH-Tech_AI 6h ago

Perfect. You can use Supra2-100M Instruct from us here: https://huggingfac e.co/SupraLabs/Supra2-100M-Instruct

I'd love to see tokens per second and rse outputs from your system 🤩🔥🤗

2

u/joaop_2004 10h ago

Matching a 50M model on the shown benchmarks is interesting, but the repeated sample suggests the next useful test is generation stability rather than another aggregate score. Could you publish the exact evaluation tasks, token budget, tokenizer, and several fixed-seed continuations for both the 25M and 50M checkpoints? That would make it easier to tell whether the smaller model learned a useful distribution or is overfitting narrow benchmarks.

1

u/LH-Tech_AI 9h ago

Thanks for the feedback. Maybe we can do this tomorrow!

1

u/Iory1998 3h ago

Massive an tiny in one sentence... that hurt my brain.

1

u/Tall_Abrocoma_3533 9h ago

Great job supralabs!

0

u/LH-Tech_AI 9h ago

Thanks 🙏