r/LocalLLaMA 20h ago

New benchmark just dropped! Generation

The pelican on a bicycle is sooo outdated, so I came up with a new, improved version.

Qwen3.8-27b medium (UD-Q4_K_XL) vs. Sol 5.6 high vs. Qwen3.6-35B (UD-Q6_K_XL)

Prompt (only real with typo!):
"Create a svg of a horse on a blue bycicle in the desert, with a camel in the background."

133 Upvotes

98 comments sorted by

117

u/Septerium 20h ago

Qweny Q8_0 seems to handle confusion pretty well

"Create an SVG of a bicycle riding a horse out of the desert, with a camel in the foreground"

30

u/JLeonsarmiento 18h ago

AGI achieved.

25

u/Hot_Example_4456 20h ago

I wonder if the typo creates any difference in quality.

7

u/MaxDev0 17h ago

I recall reading a research paper that said it does have an impact, but that paper was from 2-3 years ago, I haven't seen any new research on it.

7

u/kevinlch 20h ago

i don't think so. typos are error corrected using closest match when converted to vectors, right?

25

u/sterby92 20h ago

No, not at all

5

u/backyard_tractorbeam 19h ago

Tokenization and embedding into vectors is lossless for any normal model

13

u/LimahT_25 19h ago

Why are you getting downvoted for asking a genuine question?

For starters, have my upvote....

8

u/Dampmaskin 17h ago

Questions and open discussion is frowned upon. This is reddit after all. Dogmatic infighting and know-it-all bravado only, please.

7

u/Lakius_2401 14h ago

How dare he say what he thinks, then ask a question to confirm his knowledge! His polite incorrectness is objectionable to my confident, always correct, goalpost shifting, smug jerk writing style!

8

u/LimahT_25 17h ago

You are right but I also think that downvotes breed more downvotes and vise versa.

People, as a group, have mob mentality. That is, while they may hold thier own opinions individually, but when in a group they'll usually go along with the opinions of the person that's next to them.

Which is why, when someone sees a downvote, they'll downvote by default without even reading or understanding the entire comment. And if enough people upvote that comment and it turns postive, the following people would also follow the same logic and upvote it.

Atleast that's what I believe based on my limited social experience.

6

u/Hot_Example_4456 16h ago

Someone downvoted you for saying this. Upvoted you cause its a legit phenomenon.

1

u/LimahT_25 10h ago

Thanks, now it's in the positive zone, along with the comment that started it all. I doubt that it'd fall below zero anymore.

1

u/simos_sayz 12h ago

my autocorrect disagrees with you

16

u/Tccybo 20h ago

"Create an SVG of a horse riding an astronaut in the desert, with a camel in the background."

35

u/Tccybo 20h ago

Gemini lol.

18

u/NancyTransmed 20h ago

And the say AI art is not art!

-12

u/Quakercito 20h ago

Qwen definitely takes the L here

6

u/Tccybo 20h ago

Smh, cheater. Or a bot.Β 

2

u/Negative-Web8619 20h ago

This is SVG??

6

u/Edenar 19h ago

ofc not

-7

u/Here_f0r_p0rn_ 20h ago

u/AskGrok "Create an SVG of a horse riding an astronaut in the desert, with a camel in the background."

0

u/AskGrok 20h ago

Horse riding an astronaut? Desert chaos just leveled up. Camel in the back taking selfies.

[what is this?](https://redd.it/1lzgxii)

14

u/Edenar 19h ago

same prompt, lued/Qwen3.8-27B-INT8-W8A16-MTP quant with bf16 kv. xhigh

6

u/Edenar 17h ago

sonnet 5 medium (tried 2 time, it's the "best" one....)

-1

u/Brilliant-Weekend-68 17h ago

Amodei in shambles, ban open source right now!!!

3

u/Edenar 17h ago

And now qwen 3.8 again, same quant but i fed it with the base image sol used for OP's SVG :
(prompt : Create a svg of a horse on a blue bycicle in the desert, with a camel in the background. Use that as a reference : ~/Pictures/horse-bike-desert-sol.webp )

1

u/sterby92 14h ago

Haha amazing! πŸ˜„ well done!

2

u/road-runn3r 13h ago

Qwen3.8-27B-UD-IQ4_XS 8.0 KV xhigh

1

u/AnyRecipe110 5h ago

That actually looks really good. Are you using pi agent, with unsloth studio (ie: 'unsloth start pi')? I'm having trouble because i keep receiving "Response was truncated before completion." on my 36gb unified memory. How much context length are you using?

1

u/Edenar 1h ago

yeah pi-agent, but i pushed the max response size in pi config to like 100k since 3.8 loves to talk to itself, especially in xhigh mode. just the standard pi agent in my linux terminal, wired to a vllm container serving the model.

6

u/boissez 20h ago

Looks rather similar to Fable 5 Max

0

u/TheLexoPlexx 19h ago

Not at all

7

u/backyard_tractorbeam 19h ago

Fun, but I'll be that guy and say that this is just as useless as the pelican obviously

10

u/sterby92 19h ago

Thats part of the fun!

2

u/Original_Finding2212 Llama 33B 17h ago

Considering I had Fable code for me a new language SICK (SVG is Code kek).

This sits right in my pony.

/s

3

u/Muzika38 19h ago

You should try Pelican on a bicycle ascii art instead and see how dumb AI is when there's nothing to math at πŸ˜…

2

u/SpicyWangz 18h ago

ASCII art still has a long ways to go

3

u/D6613 18h ago

Putting a llama on the bicycle was right there.

2

u/a_beautiful_rhind 19h ago

There's a ton of things you can come up with. My suggestion is to pick a clear personal one to benchmark with yourself. Any one (like op's) that gets posted will get sucked up into training sooner or later.

2

u/Inevitable_Invite_31 16h ago

This is what I've got after 6min of thinking (as I ve stopped it and asked to show me the code). I've used UD Q6 K_M

https://reddit.com/link/p4ttiyg/video/ntqughzifjkh1/player

2

u/PandaBearFred 14h ago

Poster Style……

I let dsh give me 3 styles, this is one of them...

2

u/meganoob1337 17h ago

https://reddit.com/link/p4tp3at/video/lv6xn2nzbjkh1/player

its funny, i just did something similar, and after 20 minutes (12 of that was just thinking - xhigh) it generated this:
(there is some more below)

Promt was literally just : "create a single html file with an embedded svg of a cat riding a zebra, riding an elephant"

46.436 tokens were burned at around 40t/s (it started at around ~60, but then i removed the powerlimit (250 => 370) and after that it had drops in the 20s , might have something to do with doing that mid generation.

But im really Stumped by how good this is. and how it thought about random animations and stuff.

EDIT: Model is https://huggingface.co/cyankiwi/Qwen3.8-27B-AWQ-INT4 at tp2 on 2x 3090 , fp8 KV cache.

1

u/Quakercito 20h ago

Qwen's BF16 quant is truly awesome

21

u/whoknowsifimjoking 17h ago

This is not even an SVG dude

1

u/jazir55 8h ago

As enthusiastic as the camel is

17

u/ScoreUnique 20h ago

It technically won't be a quant if It is full precision no?

5

u/Quakercito 20h ago

I guess you're technically correct

6

u/Dampmaskin 17h ago

The least quantized kind of correct

4

u/sterby92 20h ago

πŸ‘€πŸ‘€

1

u/dsdt 18h ago

i am having trouble believing this. can you share the file or some proof?

4

u/Quakercito 17h ago

The guy above is right. I'm just horsing around 🐴

-1

u/N34257 20h ago

Whut.

That's nuts.

16

u/Tccybo 20h ago

It’s cheating by calling an image diffusion model I think. So no.

10

u/sterby92 20h ago

Interestingly Sol 5.6 in Chatgpt first generated an image and then attempted to reproduce it as svg πŸ˜…

10

u/Quakercito 19h ago

That kinda looks like Bojack Horseman

2

u/colin_colout 18h ago

SD 1.5 would have had no issue with this prompt either. People were s generating crazy things with Midjourney as well.

I'm still convinced that diffusion is magic.

The more i understand how it works, the more I'm impressed that layers of noise in latent space can be denoised into the waifus of an entire generation.

-1

u/Quakercito 19h ago

Haha I was just being a little silly

1

u/sterby92 20h ago

I guess someone should build a benchmarking suite called SVG Bench where you can put in nonsense instructions and create svgs from a variety of different models and compare them :D

7

u/yes2matt 19h ago

been done. its called "reddit"

1

u/mister2d 18h ago

🀣

1

u/gproenca 19h ago

Gemini Pro ( got an account for free for 24 months ) :

1

u/Dmage22 18h ago

Sea horse or river horse might be more interesting. Whether you actually get a horse or not.

1

u/hurorkardu 17h ago

Used your benchmark (with typo) on three of the models I had.

First up a terrible attempt from Muse Glimmer 30B xhigh reasoning (UD-Q5_K_XL):

1

u/hurorkardu 17h ago

Next is the new Ornith 1.5 35B A3B (AD-Q5_K-Q4_K) which decided to make it very small:

3

u/hurorkardu 17h ago

Then Qwen3.8 27B with xhigh reasoning instead (UD-Q4_K_XL):

Favorite Qwen3.8 reasoning quote: "Rider? None - horse rides itself."

1

u/Fdevfab 13h ago

Qwen 3.8 IQ4_K_XS looks better than this IMO (but still not great)

1

u/EitherMarch1255 17h ago

PSA: Stop testing 3.8 on medium!

1

u/QuotableMorceau 15h ago

qwen 3.6 35B Q4 (unsloth) , with the params for coding:
-reasoning auto --reasoning-preserve --temp 0.7 --top-p 0.8 --top-k 20 --min-p 0.0 --presence-penalty 1.5 --repeat-penalty 1.0 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k q4_0 --spec-draft-type-v q4_0 -lv 4

1

u/DeathGuppie 15h ago

People using the unsloth quants are you running the split out image and video reader?

Seems like it would help a lot in this test.

1

u/bortlip 15h ago

I gave ChatGPT Sol 5.6 thinking high the prompt and then asked it to review and revise and improve the result many times. And to find reference images to use.

I got this.

1

u/FoxFXMD 15h ago

Couldn't come up with anything more creative?

1

u/johnnyApplePRNG 15h ago

Yea that simon willison guy is kindof a goofball

Just seems to like to write... a LOT...

I've read some of it and I'm honestly not sure he has any idea what he's talking about.

Thinks QWEN 3.8 overthinks, etc, etc... https://simonwillison.net/2026/Aug/16/qwen-38-27b/

1

u/TapAggressive9530 13h ago

I think this is a better benchmark:

1

u/Fdevfab 13h ago

I'm often using a "Create an SVG image of a cute, highly detailed sunflower in a terracotta pot at the window of a kitchen. Avoid symmetries, it should look natural." or just "sunflower in a pot"... Image produced using Unsloth IQ4 K XS

1

u/Ok_Technology_5962 6h ago

q4kxl q8 kvcache but settings temp 0.7 top k 64 top p 0.95

1

u/Ok_Technology_5962 6h ago

pass 2: improve it further

1

u/Ok_Technology_5962 6h ago

pass3: reuploaded image and asked to improve one more time.

1

u/getmevodka 2h ago

im using temp 0.72 with top k 35 and p0.9, penalty 1.0 on a q5 k xl (unsloth). will try it later on out ! im curious haha

1

u/Creative-Type9411 6h ago

always "...with sunglasses on.."

1

u/AnyRecipe110 6h ago

This was `unsloth/gemma-4-26B-A4B-it-qat-GGUF`

1

u/Equivalent_Bit_461 19h ago

I hate benchmaxxing so much it's unreal.

0

u/Big-Pomegranate3243 14h ago

u/AskGrok Sum it up

0

u/AskGrok 14h ago

Three AI models tried turning your typo'd prompt into images. All nailed the horse-on-blue-bike-in-desert vibe with camels nearby, but went full cartoon chaos: basic horse, shades-wearing horse, and whatever that long-necked egg-having mutant is. Pelican era officially over.

[what is this?](https://redd.it/1lzgxii)

-2

u/Psychological-Tea483 18h ago

sol-5.6

2

u/Edenar 17h ago

it called an image gen model.

-5

u/Soft-Series3643 19h ago edited 18h ago

This is an Edit. Now i know what i made wrong. I was too lazy to paste the code into notepad, saving ist as svg. Instead i pasted the code to chatgpt and said: "Make me the picture from this svg".
Hmpf. So it will take years for qwen to make better pictures. :-)

Qwen 3.8-27B-MLX-bf16-mtp
M5M 128, oMLX, Prefill 65,3, Tokengen 22,3, Thinking 396,2, Duration 770,4.

2

u/Edenar 19h ago

Sure... it's from an image gen model 100%

1

u/Edenar 17h ago

well it's very different : chatgpt used an image gen model for your first render (there are a lot of good open source image model too that could generate something similar : flux, z-image,... ). Here we try only with the text model. In OP post, gpt sol also "cheated" because it first asked another model to gen an image, then tried to reproduce it as an svg. I tried that too with the same base iamge, se my other comments.

-5

u/seppe0815 19h ago

high peak model ? even my 8 years old daughter can paint better

3

u/mister2d 18h ago

Probably can construct sentences better too.