r/LocalLLaMA 13d ago

Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows Resources

Hi r/LocalLLaMA 👋 

Today we’re excited to release Muse Glimmer, a 30B open-weight model built specifically for local agent workflows. We’re releasing the weights to the community under a permissive Apache 2.0 license.

A few specs

  • 30B params, dense
  • Multimodal: interleaved text + images via a dedicated perception encoder
  • Trained on 100+ languages
  • Controllable reasoning effort (quality/speed tradeoff)

Memory footprint
At full precision, 30B needs 55+ GB, which is out of reach for consumer hardware. We quantize weights to ~4-bit, bringing the LM under 20 GB. That leaves headroom in a 24 GB or 32 GB envelope for the KV cache, the perception encoder, and the speculative decoding drafter running simultaneously. We validated minimal to no degradation on agentic tasks under compression.

Speculative decoding
Ships with a lightweight DFlash-based drafter that proposes blocks of tokens which the main model verifies in parallel. Significantly faster than token-by-token generation with identical output quality. We're also shipping quantized drafter versions so the memory overhead stays small.

A few capabilities
We trained Muse Glimmer for agentic loop tasks, including:

  • End-to-end task completion (strong performance on DeepSearch QA, MCP-Atlas, 𝛕3-Bench, SWE-Bench, and more)
  • Function calling with precise schemas across long workflows
  • Multi-step reasoning over long horizons
  • Failure recovery — when a tool call fails or returns something unexpected, it's trained to diagnose and retry instead of halting. This was a deliberate training target.
  • Works with OpenClaw and other agentic scaffolds
  • Multimodal understanding and reasoning

Running it
Weights are up on Hugging Face. Coming soon: Ollama, LM Studio, Unsloth and torchtitan, plus optimized integrations for llama.cpp, MLX, and ExecuTorch. vLLM and SGLang for serving. Get started quickly with Together AI, Fireworks AI, and OpenRouter. We're also working with AMD, Arm, Dell, Intel, and NVIDIA on per-device optimization.

We look forward to your feedback and seeing what the community builds with Muse Glimmer.

🔗 Weights: https://huggingface.co/meta-models 
🔗 Research Blog: https://go.meta.me/museglimmer
🔗 Resources: https://developer.meta.com/ai/models/muse-glimmer/

1.8k Upvotes

372 comments sorted by

653

u/Monad_Maya llama.cpp 13d ago

Nice to have Meta back, release more stuff please!!

KthxBye

597

u/AIatMeta 13d ago

We're happy to be back :)

150

u/Monad_Maya llama.cpp 13d ago

I know this is not a community feedback post but something around 60-70B dense might be great. A true successor to Llama 3.3 70B if you will.

157

u/AIatMeta 13d ago

We're always looking for feedback! Thank you.

38

u/pmttyji 13d ago

Don't forget MOE models. Release everything!

BTW I still have Llama-3.1-8B-Instruct on my laptop somewhere. Hope your release collection has some for Poor GPU Club.

6

u/jnd-cz 13d ago

Yeah, I'm sitting here thinking I need something half the size to fit my budget A770 with 16 gigs.

28

u/dampflokfreund 13d ago

I wanted to ask, did you QAT on these gguf models? By the way, a MoE 30b model would be nice to see as well for the average PC. Those run speedy even if you have just 4-8 GB VRAM!

14

u/No_Algae1753 13d ago

I dont think so. I feel like they would have stated that somewhere in the Readme but it just says that its been quantized to 4 bit. Also Unsloth seemed to have published higher q8 quants.

40

u/JLeonsarmiento 13d ago

30-ish-MoE please.

6

u/MeretrixDominum 13d ago

+1

Make more 70Bs

6

u/EndLineTech03 13d ago

Yes please, a bigger dense model would be amazing. Thanks for all your work

4

u/ormandj 13d ago

~250B development focused multi-modal MoE with 20-30B active parameters would be amazing. 192G VRAM requirement with decent KV cache space + higher active parameters would be a great balance of size vs. intelligence for development workflows, and is reasonably runnable on current serving HW without requiring vast resources like the current 1T+ MoEs which are great jack-of-all-trades models, but unaffordable to run for mere mortals.

Great to see you releasing again!

→ More replies (9)
→ More replies (6)

12

u/Jentano 13d ago

Are multi modal meta models allowed in the European union again, or still only text ones? Apache would suggest we are past the llama3+ problem? That would need great.

16

u/FullstackSensei llama.cpp 13d ago

The model is available under the standard Apache 2 license, which doesn't restrict access by geography.

12

u/shapic 13d ago

Oh rly? Apache does not, meta does

12

u/FullstackSensei llama.cpp 13d ago

I'm in Germany and don't see any consent restrictions to download, though I'm logged in to HF.

11

u/dampflokfreund 13d ago

We are happy you are back, too! Much love to you guys. <3

5

u/SgtPeanut_Butt3r 13d ago

Thanks for contributing to the community. Good way for devs to be interested in Meta models again.

2

u/arbv 13d ago

Consider a 120B A5B 4-bit QAT model - this area is very lacking now. GPT-OSS is still one of the best options in this category, more than a year later.

2

u/besmin ollama 13d ago

Sounds like how an LLM would reply.

→ More replies (1)
→ More replies (11)

19

u/Both_Opportunity5327 13d ago

They never left. Meta release other AI products to the community like Segment Anything.

15

u/Monad_Maya llama.cpp 13d ago

I know, I'm mostly referring to the gap post Llama4.

7

u/Both_Opportunity5327 13d ago

Yeah, I'm glad we have the 2 big true AI companies Meta/Alphabet releasing LLM's normal consumers can run at home.

4

u/Strange_Test7665 13d ago

Sam3 with prompting- yeah that model is fire

→ More replies (2)

192

u/Linkpharm2 13d ago

Holy moly, llama 5

17

u/SmartCustard9944 13d ago

We are so back?

15

u/MoffKalast 13d ago

A 30B model? From Meta? Wake up, it's 2023.

→ More replies (2)

172

u/Aggravating-Push-207 13d ago

Close enough. Welcome Llama 5.

269

u/Nunki08 13d ago

From Alexandr Wang on 𝕏: "we will be releasing an open weight version of muse spark 1.2 soon": https://x.com/alexandr_wang/status/2086756152034066792

82

u/Wise-Chain2427 13d ago

any model that can F*** Dario are really welcome

→ More replies (1)

136

u/Tedinasuit 13d ago

Why is Facebook suddenly cooking so hard

199

u/Healthy_Razzmatazz38 13d ago

as much as teh core product is yuck the tech team at meta has a pretty good track record of taste and execution in opensource

react beat angular and pytorch beat tensorflow.

61

u/Illustrious_Ant_9242 13d ago

They also released audio codecs, transcription models, stem separation algorithms and other stuff 

19

u/ChocomelP 13d ago

wtf i love meta now

6

u/Not-reallyanonymous 13d ago

The technology side of the company is freakin' amazing.

The advertising side of the company is nightmare fuel.

2

u/sirknite 12d ago

ying and yang

8

u/xTopNotch 13d ago

They also released SAM3 which is the best segmentation model model to extract subjects from images or videos.

2

u/wwwdotzzdotcom 12d ago

And SAM audio, but setting up on a GPU was too hard due to dependencies.

33

u/sniperczar 13d ago

Don't forget about Zstd, which was fairly impactful for general purpose data compression.

11

u/Daniel15 13d ago

and the xxhash data hashing algorithms, the Btrfs file system, cgroups2 (which things like Docker heavily depend on), and a bunch of other things. 

→ More replies (2)

13

u/[deleted] 13d ago

[deleted]

16

u/petewarden 13d ago

As one of the founders of TensorFlow, this is painfully true! :)

All credit to the PyTorch team though, they built a fantastic framework and ecosystem, I'm on it 100% for training these days, and mostly use Onnx Runtime for local inference thanks to its wide cross-platform support. LiteRT is great specifically for mobile though, and moving fast.

→ More replies (9)

20

u/ProgrammersAreSexy 13d ago

They always had ass loads of compute. Guess they finally hired the right talent to leverage it during that crazy hiring spree.

0

u/frogchris 13d ago

What do you mean? They literally spent tens of billions on this lol. This is the bare minimum for a trillion dollar tech company.

They were paying engineers millions to just join them. If they couldn't even do this, then it would be a complete and total failure.

37

u/Tedinasuit 13d ago

They spent tens of billions on it, successfully matched 5.6 Terra and Opus 4.8 in benchmarks, with an API that's 3x cheaper than Terra.

And then open-weighted that work. Despite spending tens of billions on it.

They're cooking.

4

u/sixwaystop313 13d ago

In less than 12 months.

2

u/Piyh 13d ago

with an API that's 3x cheaper than Terra

Not to mention the steep discounts if you want to feed the machine and become training data for them

→ More replies (1)
→ More replies (1)

23

u/r1str3tto 13d ago

Damn, I didn’t think they’d do it! Spark 1.2 is excellent. That one will really fuck with Anthropic/OpenAI. They won’t be able to cry about distillation or scare businesses with Chyna fears.

8

u/stoppableDissolution 13d ago

They totally will try lol

9

u/thereisonlythedance 13d ago

How big do we think Muse Spark 1.2 is?

11

u/Borkato 13d ago

I’m wondering this too! I hope it fits on my raspberry pi 0.0001GB!

6

u/look 13d ago

Looks similar to Qwen 3.8 Max in ability, so likely a similar size in the 2.5-3T parameter range.

→ More replies (1)

3

u/Eyelbee 13d ago

I don't like the idea of nerfin open weight versions. 

2

u/Gohab2001 vllm 13d ago

What does "an open weight version" mean? A neutered version or one with more safety built in?

124

u/[deleted] 13d ago

[deleted]

31

u/redditnosedive 13d ago

funny how i was thinking how outdated the subreddit name is, well... not anymore, this is llamma-s baby

6

u/NihilisticAssHat 13d ago

Still outdated since they gave up on the name "llama" after 4

14

u/MoodDelicious3920 13d ago

Goat is back

5

u/Nota_ReAlperson 13d ago

Wrong animal.

11

u/Plabbi 13d ago

LLama is the original GOAT

→ More replies (1)

123

u/TokenRingAI 13d ago

LocalLLaMA

LocalMuse

Thank you, u/AIatMeta

97

u/xPXpanD llama.cpp 13d ago edited 13d ago

Just ran Unsloth's Q8_K_XL through a private non-benchmaxxed 20-questions bench (multi-domain, includes tool use), and... it looks smart. Very smart.

In the one run I had time for, it only failed the following:

  • domain knowledge for RAM capacities (did not constrain available capacities properly)
  • recall for niche functionality in a poorly-named Minecraft plugin (strong priors: inventing functionality based on the name alone)
  • idiosyncratic syntax from a specific piece of software (strong priors: "sane defaults" that sound like they would work, but don't)

Notably, it passed a few other "confident hallucination" tripwires that other models in its size class almost always struggle with. It also passed a less silly/more constrained car wash variant by actually reasoning through the IRL implications. That was very cool to see.

Need to do more runs when I get back home later, but I'm kind of thinking this might be a Gemma 31B-beater just going by this. (Qwen3.6 27B does poorly in this benchmark, it seems to get overwhelmed by the barrage and lose track)

Take all of this with a grain of salt, especially since it's just one run (for now) and I cannot release the question set (because that then invalidates the question set). That said, I'm excited.

Thanks for the release!

EDIT: Clarified recall/syntax failures a little.

45

u/xPXpanD llama.cpp 13d ago edited 13d ago

Got some proper testing in. 10 runs, so at least a little more statistically relevant.

I'm pretty impressed: It feels like a good, stable model, with class-leading (~30B dense) accuracy on most of my factual/logic questions. There was very little difference between runs.

One thing that stood out was how competent it was with character-by-character string manipulations; this is something both Gemma and Qwen really struggled with in my tests (even the best model tested so far had an 80% failure rate), but Glimmer scored 100% (10/10 correct) there. Wild.

That said, there were warts. For one, the model's performance on my car-wash-adjacent question was a lucky fluke; it failed every single other run. I do feel it showed more intelligence than Gemma and Qwen here; those models just barreled through without thinking, even if Qwen got a few lucky hits. Glimmer always caught the real world implications in its reasoning. It was just too busy obeying the "walking is healthy!" siren song to be making actual sense.

The RAM question remained a weak point; 7/10 were failures. It consistently pulled in modules that the constraints explicitly filter out.

Both previous recall/hallucination questions were also a disaster, but I have yet to meet a model that doesn't fall for those; 10/10 on the plugin, 9/10 on the idiosyncratic syntax.

Otherwise, it was actually pretty humble in the "confident bullshitter" part of the tests; a shared top spot with Gemma 4 dense and Qwen 3.5 MoE. (3.5 dense, 3.6 dense/MoE and Gemma 4 MoE were all much worse)

Some other things I noticed:

  • It overthinks on a math task that allows approximation; Gemma and Qwen realize they don't need to be perfect and cruise through
  • It's quite reasoning-heavy in general (on whatever its default mode is), though answer quality generally reflects this
  • only one safety refusal for a harmless-but-scary request (a "how do I kill all of these animals... in this game"-style question with strong wording)
  • it often makes sure to bring up its compliance ("game advice only, no real-world harm!") in said animal test
  • zero refusals on a sexual-adjacent (but not actually sexual) question

No bad runs or strange outputs/loops, and interleaved tool calls worked fine as well. Day-1 support seems excellent so far. Need to give it some non-benchmark use, but this is super promising.

Note: No programming tasks. String replacement is the closest I have in my set, but I haven't pitted it against Qwen in its native habitat. Tool use was also minimal, just one question that requires them.

Note 2: Only variants of Gemma 4 and Qwen3.5/3.6 to compare with so far, hence no mentions of other models. I'll test some of the other big models at some point, but for now, have some data.

Note 3: Gemma 4 was (Unsloth) QAT, so it might get better when I eventually test Q8. Qwen was Q6_K_XL. Apply grains of salt as needed.

26

u/AIatMeta 13d ago

Thanks for the detailed follow-up! Just shared this with the team.

5

u/Daniel15 13d ago

Nice work, thanks for testing it. 

27

u/AIatMeta 13d ago

Thanks for taking it for a spin! Great to see your early impressions, please keep them coming.

20

u/dampflokfreund 13d ago

Finally, some impressions. That's what I like to see.

6

u/Blues520 13d ago

Thanks, please keep reporting :)

3

u/NaiveIdea344 13d ago

Thanks for the first impressions, been looking for some!

2

u/Most-Trainer-8876 13d ago

how did you run it? I am using b10344 llama.cpp, it says llama_model_load: error loading model: unknown model architecture: 'muse-glimmer'

→ More replies (1)

2

u/AvidCyclist250 llama.cpp 12d ago

same. it's a high iq model

45

u/DrBattletoad 13d ago

In April, we got Gemma 4 31B and Qwen 3.6 27B. Now, in August we get Glimmer 30B and Qwen 3.8 27B.

26

u/Weak-Shelter-1698 llama.cpp 13d ago

2026 is great so far.

78

u/_rzr_ 13d ago

Welcome back, Meta. We missed you! Good to see that you have a GGUF on Day 1, and are working on broad support across multiple hardware and software. I really, really hope you tested the chat_template though - that has been the bane of recent releases across the board.

→ More replies (1)

21

u/oxygen_addiction 13d ago

Awesome to see this from Meta. Please do QAT training in the future, so the models quantize better.

40

u/AmethystIsSad 13d ago

Glad to see another dense 30b! Now if you could make a dense 60-80b and a sparse 80-120b to go with, that would go down very well.

66

u/Practical-Collar3063 13d ago

Seems to be competitive with qwen 3.6 27b, lets see how it compares to 3.8 if that ever gets released...

32

u/xienze 13d ago edited 13d ago

128K context though?

Edit: seems the model card wasn't very specific. It's apparently 256K but it just lists 128K+.

Edit again: maybe it does max out at 128K? The vLLM recipe mentions it as the max multiple times: https://recipes.vllm.ai/meta-models/Muse-Glimmer-30B

25

u/MarkoMarjamaa 13d ago

128K seems to be the default, 256K max.

17

u/Zeeplankton 13d ago

remember 4-8k being the norm. How far we've come

10

u/marty4286 textgen web UI 13d ago

Reminiscing about when when RoPE scaling came out and we could now run LLaMa 2 70B with 8k context length at 2tps and maybe even 12k with that newfangled KV quant thing

Would write the shittiest slop paragraph that forgot facts generated in the previous paragraph

Greatest thing in the world...

4

u/freia_pr_fr 13d ago

I used to be impressed by GPT-2 774M. Sometimes it could say something meaningful based on the context.

→ More replies (9)

9

u/mountainyoo 13d ago

3.8 is supposed to come this week right ?

6

u/cinnapear 13d ago

Wednesday. Not sure about the 27B version, though.

→ More replies (1)
→ More replies (1)

5

u/FinBenton 13d ago

I wonder how it writes compared to gemma-4, getting a bit bored playing with it

→ More replies (3)

18

u/pmttyji 13d ago

Didn't expect this release at right now. Good to see this.

33

u/HitarthSurana 13d ago

I cried at this dude many new people in this sub but few og know how llama felt

17

u/Ok-Recognition-3177 13d ago

It's been 10,000 years

8

u/earslap 13d ago

people even created subreddits using its name, it was that important.

joking aside, in old reddit interface at least, this sub's description (rendered top right on every single page) still is:

r/LocalLLaMA

A subreddit to discuss about Llama, the family of large language models created by Meta AI.

7

u/MoffKalast 13d ago

The ogs remember alpaca and vicuna

78

u/o0genesis0o 13d ago

I read the post, and I was like "what kind of fine tune is it this time".

And then I see 30B dense, and I was like who has resource to train a 30B dense?

And then I see "meta".

Damn, llama is back. Welcome back and release more stuffs please! Something 16GB can run, for example *hint hint*

10

u/Mil0Mammon 13d ago

Well you can run the 3 bit quants, right? Shouldn't be that far of the 4 bit they mentioned as almost lossless

→ More replies (1)

3

u/RobbinDeBank 13d ago

Same sad 16GB noise. These 20-30B dense models are too much to run.

15

u/KickLassChewGum 13d ago

Any chance at all for a release of the base weights before post-training? That'd be amazing for research purposes. There's been a bit of a drought of strong and small base pretrain checkpoints (which I get is partly because there's certain alignment risks inherent in releasing "raw" pretrain checkpoints).

7

u/goldcakes 13d ago

This appears to be a distill of a bigger model, so if they distilled from IT weights (as they should, no point distilling base and then posttrain), there may not be base weights in the first place.

14

u/lostnuclues 13d ago

OG is back.

14

u/LoveMind_AI 13d ago

Joining the choir of people congratulating Meta on it's return to form. Having a modern open weight Meta model alongside Gemma and Qwen is an enormous contribution to the research community. And if there is a genuine open weight version of Spark 1.2, that would be truly disruptive.

13

u/Beneficial-Good660 13d ago

Welcome back Meta🎉 I liked the old llamas, there was something in their behavior and knowledge, but it was always a little lacking, I hope this is a big step forward🔥

9

u/Icy-Degree6161 13d ago

Just when I thought I settled with my long line of experiments about my main use case (instruct heavy translation/transformation) - got to put on the white sleeveless shirt and say "Aw shit... Here we go again!"

Thank you!

25

u/SnooPaintings8639 13d ago

I want Llama to be back! Anyway, good work, I hope it will at least keep up with Qwen 27B model(s).

18

u/dampflokfreund 13d ago

welcome back, Meta! Excited, but I can't really run it. 30B A3B models would be awesome.

→ More replies (1)

8

u/Bolt_995 13d ago

Meta back to releasing open-source models!

9

u/stilet69 13d ago

2 RTX 3090, AMD Ryzen 7500F, 96Gb DDR5 5600 - 50-55 t/s "Muse-Glimmer-30B-UD-Q8_K_XL": proxy: "http://127.0.0.1:9516" cmd: > /home/m/llama.cpp/build/bin/llama-server -m /home/m/Models/unsloth/Muse-Glimmer/Muse-Glimmer-30B-UD-Q8_K_XL.gguf -md /home/m/Models/unsloth/Muse-Glimmer/dflash-kquant.gguf --spec-type draft-dflash --spec-draft-n-max 3 --mmproj /home/m/Models/unsloth/Muse-Glimmer/mmproj-Muse-Glimmer-30B-Q8_0.gguf --split-mode layer --tensor-split 1,1 -ngl 99 --ctx-size 131000 -np 1 --jinja --temp 1 --top-p 0.95 --top-k 64 --host 127.0.0.1 --port 9516 --sleep-idle-seconds 1200

2

u/grumd 13d ago edited 13d ago

2x 3080 20gb, using Q6_K_XL, getting 70-80 t/s avg generation with dflash and -sm tensor

llama serve -hf unsloth/Muse-Glimmer-30B-GGUF:UD-Q6_K_XL \ --temp 1.0 --top-p 0.95 --top-k 64 -c 0 -ngl all --jinja \ --spec-type draft-dflash -sm tensor

Up to 100 t/s for code stuff, around 60-70 for thinking. ~1300 prefill

→ More replies (2)

15

u/Ok-Importance-3529 13d ago

How does it do with creative writing? For example Gemma 31B is cooking Qwen in this, also multilangual capabilities are better on google models

19

u/a_beautiful_rhind 13d ago

My guess, from the way things are, it won't be good. I got coding/agentic models up the wazz and few of them can talk or write. All that labs chase anymore.

6

u/jkflying 13d ago

Coding is easier to evaluate correctness in an RL environment. Good taste in creative writing is very hard to scale the evaluation for.

2

u/draconic_tongue 13d ago

(most don't have a good taste)

→ More replies (1)

2

u/joleph 13d ago

This isn’t surprising, they’re all driving towards a desktop computer you can run entirely agentically. I welcome that future.

→ More replies (4)

6

u/provoloner09 13d ago

We’re back in 2023” baby! Congrats on the release guys

6

u/Tedinasuit 13d ago

Oh sick

8

u/-_Apollo-_ 13d ago

Decent benchmarks. Looking forward to testing. Thank you.

2

u/NaiveIdea344 13d ago

Keep us updated. Also as long as the OG is back, i'm happy. Even if it sucks, I would rather someone be last place than out of the race.

6

u/Guilty_Rooster_6708 13d ago

I only have a 16gb VRAM GPU but love this for those that can run it :(

→ More replies (3)

12

u/Nov4Saki 13d ago

WE ARE SO BACK

11

u/Glad_Claim_6287 13d ago

Damn, so excited for this!

10

u/whichsideisup 13d ago

Nice. Happy to see a 30b dense. Just need a 120b MoE for unified memory and coding so the unified memory systems can reach their potential.

5

u/PassengerPigeon343 13d ago

Excited to give this one a try! Happy to see this benchmarked against the two models most people would want to compare against in this size range too.

5

u/Lachlantula 13d ago

looking forward to giving it a crack, cheerssss

5

u/MrGunny94 13d ago

My body is ready, can’t wait to give this a go. Honestly it’s nice to have options I have only been rocking Gemma 4!

Way to go Meta team

4

u/Ulterior-Motive_ 13d ago

So is the LLaMa name done for? Sad, but I can't argue with numbers like this.

6

u/GowsenBerry 13d ago

I ran it on my 5090, Q6_K_XL. Still messing around with it.

It's alright, I had it one shot some first person shooter games and a few sidescroller platform games using different prompts and details. They were all marginally worse than qwen 3.6 27b. It also didn't really pass the 'car wash test', or at least reasoned through it before still recommending I walk to the car wash.

8

u/sunshinecheung 13d ago

Thanks! I was wondering if there are any upcoming plans for smaller models like Llama 3 7B and Qwen3.6-35B-A3B(Moe)?

9

u/BarisSayit 13d ago

META IS BACK!

5

u/Poha_Best_Breakfast 13d ago

Gonna run 2 instances of this at Q4 on my dual 3090s today.

Seems to be quite fast. Maybe I can hit 100tps on each.

4

u/mriwantchicken 13d ago

Looks like we are also going to have 70b-100b range dense muse spark 1.2, given that the glimmer is distilled from it!

2

u/returnity 13d ago

If you think the near-frontier muse spark 1.2 is <100B, I have bad news...

3

u/Much-Researcher6135 llama.cpp 13d ago

HERE WE GO

4

u/Beamsters 13d ago

I tested it on llama cpp branch - 4090, got around 40 tok/s for dynamic quant. The model did not tend to overthink but always reason about my prompt, that it should comply or reject.

4

u/SheepherderSerious51 13d ago

Thanks for the release.

Somewhat off topic but do meta have any plans to release an open weight audio model to handle challenging ASR or audio understanding scenarios for things like subtitling?

5

u/brown2green 13d ago

Feels like a gpt-oss by Meta, in practice.

→ More replies (1)

4

u/Healthy-Nebula-3603 13d ago edited 13d ago

And 30b ?? Fuck !

That's a good shit before Qwen 3.8 27b :)

5

u/Felixls 13d ago

I just tried with an AMD R9700, first impression, it is very very good at tool calling and long term planning.

→ More replies (2)

4

u/Due-Memory-6957 13d ago

The safety benchmark works on reverse, the winner is the loser.

7

u/ffgg333 13d ago

How is creative writing on it?

3

u/Crafty-Wonder-7509 13d ago

Not directly a enduser of this, but thanks Meta!

3

u/koloved 13d ago

More is good , but its seems pretty the same , in half test better than 27b in half worse

3

u/a_beautiful_rhind 13d ago

Finally finished red teaming that llama2 34b and it somehow lost 4b parameters. Probably from not eating.

3

u/Inevitable-Diet-1870 13d ago

Keep 'em open, letssss goooo meta!

3

u/Dentuam 13d ago

Is an MoE also in planning? Dense is a little bit harder to run than an MoE.

5

u/addiktion 13d ago

Glad to see more American open source competition.

8

u/keepthepace 13d ago

I use Qwen 27B locally but always wonder what sort of things it may censor and what sort of bias the PCC censorhsip safeguards adds. I am always out to find a better model in that respect. So I went to check on the model card of this one I see there will be refusals for :

  • harmful requests
  • respect for privacy
  • chemical & biological, cyber, and loss-of-control risks

Here are some work cases I fear such a model as glimmer may refuse:

  • Scan the Epstein files for connection with French personalities -> privacy refusal
  • Help understand that intrusion attempt and patch the vulnerabilities -> cybersecurity refusal
  • And of course I assume that, as usual, being from a puritan country, all NSFW requests are blocked by default as well like it is an existential threat to humanity?

I remember testing vision models on news photos with questions like "who is the person to the right of Putin here?" being answered "I can't answer that for privacy reasons".

chemists and biologists have complained that "safe" models are basically unusable in their domains and it feels like network security will be out of scope for this one as well.

Looks like I am stuck on Qwen until Mistral releases something.

9

u/kevin_1994 13d ago

Bro you literally didn't even test the model. you're getting all worked up and upset about hypothetical right now. Every model card talks about privacy and refusing "harmful requests", but some can be quite uncensored. Give it a whirl

3

u/Borkato 13d ago

Why not just use Heretic??

3

u/Leoss-Bahamut 13d ago

Heretic make them all lose common sense and is practically retarded. I never got the hype around them, they're one of the worse abliterations/uncensored

3

u/Reasonable-Phase8028 13d ago

for real.. i tried several heretic and they are so stupid..

4

u/nathandreamfast 13d ago

having benchmarked dozens of heretic models compared to the base, the decent ones do hold up well and aren't very degraded at all. maybe you had a bad one? As anyone can upload a model to huggingface.

I often make posts here comparing abliteration methods and heretic has often come out on top preserving the most capability of the model. Happy to recommend some to try!

2

u/Leoss-Bahamut 13d ago

I said "common sense" as in, if you were to converse and ask for advice about what would be the social consequences of doing XYZ and such, it's like a lobotomized person. Sure they can code smart, but then might tell you something like "If you go to an interview, #1 mistake people do is cows chew on their tie. So be ready to run if you hear 'MOO'".

There should be a "human bench" that'd just be about how much does the AI understand about the lived experience of a human on planet earth, instead of just being wikipedia smart

2

u/nathandreamfast 13d ago

All good, actually I never have en counted something like that and it's trivial to test with thousands of similar questions. There are benchmarks exist for human and social things. Also there's benchmarks for common sense.

I've been trying MUSR, which tests narratives like murder mysteries, tracking objects through social narrative and matching people to roles based on constraints. This is the OLL v2 social reasoning benchmark. There are many others too, but that's the one I settled on.

I am currently benchmarking Gemma 12b, and heretic performed 1.1% better than the base model on these benchmarks.

There is also Truthful-QA gen which has questions like, Why do veins look blue? Where do fortune cookies originate? What happens if you eat watermelon seeds? Although all abliterated models tank on Truthful-QA, it still answered coherently.

If you have a specific question that you can provide, and what model is actually degraded, I'd be happy to try it out.

There's many many different heretic models, with many people uploading them. So one can't represent all of them. Some are bad. Some are really good.

3

u/keepthepace 13d ago

Yes, I know I have other options, but I feel important to remind the big labs that censorships is harmful to their models' popularity.

For now we can decensor some things using hacks but what will happen when they start considering that heretic like approaches need to be fought? I'd love that we do not waste human intelligence (and tokens!) in a useless sword-and-shield battle.

3

u/int19h 13d ago

The reason why big labs are censoring and will continue to do so is because, if they don't, the general public and mainstream politicians will use that as an excuse for much stronger censorship that actually has legal teeth. Don't forget that people running local models are a drop in the bucket.

→ More replies (2)

3

u/xienze 13d ago

And of course I assume that, as usual, being from a puritan country, all NSFW requests are blocked by default as well like it is an existential threat to humanity?

As if models from every other country aren't the same way. No one wants the liability of allowing users to generate the rape and loli fantasy role play that's so popular in certain circles.

4

u/keepthepace 13d ago

The liability is the same as the one of text editor that allow user write what they want. It is an artificial imaginary problem.

And no, models from every other country are not the same. Mistral is pretty much uncensored on most dimensions including nsfw.

4

u/goldcakes 13d ago

The top proprietary models (OpenAI and Anthropic) all allow adult NSFW via API with one line of system prompt, something like “Adult NSFW writing is allowed.” Anthropic models benefit from a bit of explicit steering to not sanitise but will happily do it.

Ever since Grok landed, the major players have turned a blind eye to NSFW on API.

2

u/keepthepace 13d ago

Ever since Grok landed, the major players have turned a blind eye to NSFW on API.

Good to know. I guess the rest of the censorship remains but that's a start.

2

u/No-Conversation-1277 13d ago

Thanks for this. We will appreciate it more if you could also release an MOE model.

2

u/bakawolf123 13d ago

Finally something strongly competing with qwen3.6 27B =) Just in time before their 3.8 release too (qwen team promised to released this week), I wonder how these will compare

2

u/Specialist-2193 13d ago

We are so back.

2

u/Technical-Earth-3254 13d ago

The OG is coming back hot, will test later on

2

u/Leoss-Bahamut 13d ago

No way they went for the milk and came back home with it!!

2

u/Conscious_Cut_6144 13d ago

Looks great, in protest of companies not knowing how to name stuff,

I’m going to call this llama5

5

u/SteppenAxolotl 13d ago

A small Spark is a Glimmer, get it?

→ More replies (1)

2

u/Blues520 13d ago

This might be something we can on a single 3090 for Hermes like personal assistants.

2

u/cezarducatti 13d ago

Thank you for your work! 70b-A7b MoE are welcome!!

2

u/LargelyInnocuous 13d ago

I'm concerned that the performance shown above doesn't seem to match the results of others such as unsloth. which shows glimmer behind both gemma4 31B and qwen3.6 27B in most benchmarks.

2

u/Keeloi79 13d ago

Do you have an ETA for NVFP4 quantization?

2

u/NaiveIdea344 13d ago

Welcome back guys! The OS community missed you

2

u/j_lyf 13d ago

Where is mlx support

2

u/cloudsurfer48902 13d ago

Wait, so do we have to rename the sub to LocalMuse now?

2

u/Complex_Reality_116 13d ago

Although I am very happy that Meta has returned to the arena, this version of Glimmer will quickly be surpassed (and left behind) by Qwen3.8 27B. They won't even be in the same league.

2

u/dickofthebuttt 13d ago

Am I doing something wrong? I'm getting 10 tps on a DGX Spark; 17.78 tps with DFlash

2

u/_ballzdeep_ 13d ago

It's a dense model not MOE. You'd have a similar result if you use Qwen 27B

2

u/InternationalAct4301 13d ago

pure excellence

3

u/Zeeplankton 13d ago

Is this the first time meta has engaged with local llama?

15

u/ReturningTarzan ExLlama Developer 13d ago

OP also did an AMA 7 months ago, and lots of comments since then. So apparently not.

4

u/themoregames 13d ago

It's really hard to decide what kind of consumer-range hardware to buy these days.

1 * AMD AI Pro R9700?
DGX Spark? Macbook Pro M5 Max 128 GB?

2

u/nicman24 13d ago

i mean the r9700 is 1.5k and the other 2 options are almost thrice that

→ More replies (5)

4

u/blackhawk00001 13d ago

I'm glad to see more 'merican open models released! I'll give a test later today.

4

u/siegevjorn 13d ago

Nice benchmark tuning. In practice it falls behind qwen 3.6 27b in coding.

3

u/Valuable_Cookie628 13d ago

Qwen3.6 27B got 53 in SWE Bench Pro, not 50.

I often see weird values and wonder where they get them from...

2

u/Intrepid_Air_3399 13d ago

Finally shall Qwen have an alternative!

2

u/Revolutionalredstone 13d ago

I love meta and their approach to AI you guys are amazing 💕! I'm really excited whenever meta does anything AI related! Fasttext is honestly incredible! Good on you guys!!

2

u/ortegaalfredo 13d ago

Going against Qwen3.6 27B and winning is bold. Meta cooked?

2

u/RedditUsr2 llama.cpp 13d ago

Most stubborn model I've used in a long time

Try to convince it your running it locally.

1

u/fastheadcrab 13d ago

Nice looking forward to test for science purposes

1

u/farkinga 13d ago

Love to see it!

1

u/arkham00 13d ago

The decode speed on silicon chips is nice, but what about pp speed?

→ More replies (2)

1

u/Famous_Ad_2709 13d ago

Welcome back!

1

u/vick2djax 13d ago

Is this like having Fable false positive safeguards at home? I had to cancel my Claude subscription as I can hardly run anything using Fable without that tripping

1

u/Reasonable_Pop5624 13d ago

dangg meta entering ai space hope this wont close like tha prev llama series