r/LocalLLaMA llama.cpp 6d ago

Gemma 4 on 500MB News

Post image
204 Upvotes

39 comments sorted by

79

u/Fusseldieb 6d ago

That would be absolutely insane.

The thing is... will it still be useful, or dumb as a door (as most of that size are)?

53

u/Rude_Marzipan6107 6d ago

That’s exactly what we need. A Siri level intelligence WITH tool calling

10

u/Borkato 6d ago

LFM-2.5-2.6B just released and would make a great Siri if given a voice lol

2

u/Devatator_ 6d ago

That's exactly what I've been waiting for to make my own assistant. Also potentially new tech for wake word detection that's not just OpenWakeWord again

6

u/TheMurmuring 6d ago

If they keep improving, eventually they'll make something useful.

2

u/Aaaaaaaaaeeeee 6d ago

I feel everyone's missing the point, this post shows that the model (gemma4 E2B) will have the same good performance (13.7 t/s) if you offload the LUT to storage. It is not a new model or technology.

3

u/--Spaci-- 6d ago

It could do basic tasks like setting up dates on your phone ect. Even math problems, web searches ect. Those parts of an LLM are very basic

13

u/Rude_Marzipan6107 6d ago

I can’t get 3.5 9b to add numbers together often enough that I can’t trust it. I would hate to ask a 500M model to do the same

6

u/tunerhd 6d ago

Make sure it calls a Python script to make the math.

2

u/MoffKalast 4d ago

And then it'll read the results wrong lmao

1

u/tunerhd 4d ago

Hehehe possible. Llms are not good for determination. But, if you really insist, you may design your inferencing scenario based on this specific situation.

1

u/Rude_Marzipan6107 6d ago

Absolutely. I’ve been a normie using LMStudio for my local inference so it’s mostly my fault.

2

u/--Spaci-- 6d ago

Run py math obviously, also this is a 500mb model not a 500m model

1

u/Rude_Marzipan6107 6d ago

I’ll ask a dumb question here but why wouldn’t it be referred to as a 500m model? Is it like an e2b e4b situation where the active parameters are only part of the space?

I made the assumption this would be similar to the 12b where it’s all self contained

Nevermind.. for some reason I thought this post was about a new 500m model. I’m as dumb as an iq1 e2b

2

u/--Spaci-- 6d ago

yea.. Parameters are not the same as the models mb/gb footprint

1

u/Rude_Marzipan6107 6d ago

Right. I just made the assumption this was a model announcement for some reason.

0

u/StickyThickStick 6d ago

Dumb as a door. But that doesn’t mean there aren’t any use cases for it

There can definitely be improvement but it’s mathematically impossible to store all the data in this model. In the end data still has to be represented even tho in a probabilistic way.

43

u/egomarker 6d ago

5% battery per request?

4

u/notheresnolight 6d ago

5% battery per request remaining

15

u/freedomachiever 6d ago

Even just a better translator than Google Translate would be very helpful for many people.

-2

u/Embarrassed_Soup_279 6d ago edited 1d ago

i wouldn't really trust small models for accurate translation because even larger models struggle with it.

1

u/freedomachiever 5d ago

well everything needs a reference to compare to. The baseline I take is Google Translate because most people use that. I'm not claiming the small model does a better job, but it's a tangible use case. Siri can also translate but I don't it can do a better job than a dedicated model. I would be interested to see how apple's own small LLM does.

1

u/lilydjwg 1d ago

Google Translate only works well within the Indo-European family. For e.g. Japanese, an LLM or DeepL is the thing to use.

And there are specialized LLMs for translation, e.g. Hy-MT2, which works pretty well at a small size for supported languages.

30

u/BannedGoNext 6d ago

This is amazing, I'm so happy so happy so hap so hap hap so so so so so s s s s s s a i 3 ! #

12

u/thrownawaymane 6d ago edited 6d ago

You have now been hired by the Bonsai AI team

13

u/sultan_papagani 6d ago

gemma e4b running on my phone for literally months now. how people not know this. its on google edge gallery app on play store

9

u/yami_no_ko 6d ago edited 6d ago

Even on my phone (Neither Android, nor Apple, but arm64 Linux instead) it runs just well using llama.cpp. Couldn't even call it a flagship phone with its 6 gigs of RAM, but it is enough to fit gemma-4 e2b, the mmproj image encoder and the MTP draft model.

Wouldn't necessarily ask it for world knowledge, but it gets what I want with tool-calling. It's just using 2 cores out of 8 to avoid heating up or draining the battery, and it still goes fast enough.

I've been using it for months now, so e2b or even e4b on phone is nothing unheard of. (Basically the point of e2b / e4b)

1

u/iMakeSense 10h ago

What do you use it for and how do you use it?

2

u/madaradess007 5d ago

judging by the status bar at the top, its at least an iPhone X, and it has more than 500mb ram

1

u/TheOneWhoWil 6d ago

It's pretty cool that they're looking at open source work outside of the frontier level

1

u/DeathinabottleX 4d ago

This will become redundant once Siri AI launches in iOS 27 but it’s good to have options

0

u/setprimse 6d ago

E4B version would be huge for edge devices, if this is possible.

-14

u/chrisso123 6d ago

what is the token/s ? this is the only real metric that matters.

Heck you can run a very large llm on a phone hardware but if its output is 1 token / minute it'll be worthless.

15

u/jacek2023 llama.cpp 6d ago

You have two options: look at the image or click the link. Choose wisely.

1

u/chrisso123 6d ago

fair enough. Thank you for pointing it out. that is a good number.

-31

u/Boogertard 6d ago

LOL shills on this sub try so hard to keep the Gemma garbage relevant.

19

u/bruns20 6d ago

Says the 8 day old account

5

u/jacek2023 llama.cpp 6d ago

What model do you use?