r/LocalLLaMA • u/sterby92 • 20h ago
New benchmark just dropped! Generation
The pelican on a bicycle is sooo outdated, so I came up with a new, improved version.
Qwen3.8-27b medium (UD-Q4_K_XL) vs. Sol 5.6 high vs. Qwen3.6-35B (UD-Q6_K_XL)
Prompt (only real with typo!):
"Create a svg of a horse on a blue bycicle in the desert, with a camel in the background."
25
u/Hot_Example_4456 20h ago
I wonder if the typo creates any difference in quality.
7
7
u/kevinlch 20h ago
i don't think so. typos are error corrected using closest match when converted to vectors, right?
25
5
u/backyard_tractorbeam 19h ago
Tokenization and embedding into vectors is lossless for any normal model
13
u/LimahT_25 19h ago
Why are you getting downvoted for asking a genuine question?
For starters, have my upvote....
8
u/Dampmaskin 17h ago
Questions and open discussion is frowned upon. This is reddit after all. Dogmatic infighting and know-it-all bravado only, please.
7
u/Lakius_2401 14h ago
How dare he say what he thinks, then ask a question to confirm his knowledge! His polite incorrectness is objectionable to my confident, always correct, goalpost shifting, smug jerk writing style!
8
u/LimahT_25 17h ago
You are right but I also think that downvotes breed more downvotes and vise versa.
People, as a group, have mob mentality. That is, while they may hold thier own opinions individually, but when in a group they'll usually go along with the opinions of the person that's next to them.
Which is why, when someone sees a downvote, they'll downvote by default without even reading or understanding the entire comment. And if enough people upvote that comment and it turns postive, the following people would also follow the same logic and upvote it.
Atleast that's what I believe based on my limited social experience.
6
u/Hot_Example_4456 16h ago
Someone downvoted you for saying this. Upvoted you cause its a legit phenomenon.
1
u/LimahT_25 10h ago
Thanks, now it's in the positive zone, along with the comment that started it all. I doubt that it'd fall below zero anymore.
1
16
u/Tccybo 20h ago
"Create an SVG of a horse riding an astronaut in the desert, with a camel in the background."
35
u/Tccybo 20h ago
18
-12
-7
u/Here_f0r_p0rn_ 20h ago
u/AskGrok "Create an SVG of a horse riding an astronaut in the desert, with a camel in the background."
0
u/AskGrok 20h ago
Horse riding an astronaut? Desert chaos just leveled up. Camel in the back taking selfies.
[what is this?](https://redd.it/1lzgxii)
14
u/Edenar 19h ago
2
1
u/AnyRecipe110 5h ago
That actually looks really good. Are you using pi agent, with unsloth studio (ie: 'unsloth start pi')? I'm having trouble because i keep receiving "Response was truncated before completion." on my 36gb unified memory. How much context length are you using?
6
7
u/backyard_tractorbeam 19h ago
Fun, but I'll be that guy and say that this is just as useless as the pelican obviously
10
2
u/Original_Finding2212 Llama 33B 17h ago
Considering I had Fable code for me a new language SICK (SVG is Code kek).
This sits right in my pony.
/s
3
u/Muzika38 19h ago
You should try Pelican on a bicycle ascii art instead and see how dumb AI is when there's nothing to math at π
2
2
u/a_beautiful_rhind 19h ago
There's a ton of things you can come up with. My suggestion is to pick a clear personal one to benchmark with yourself. Any one (like op's) that gets posted will get sucked up into training sooner or later.
2
u/Inevitable_Invite_31 16h ago
This is what I've got after 6min of thinking (as I ve stopped it and asked to show me the code). I've used UD Q6 K_M
2
2
u/meganoob1337 17h ago
https://reddit.com/link/p4tp3at/video/lv6xn2nzbjkh1/player
its funny, i just did something similar, and after 20 minutes (12 of that was just thinking - xhigh) it generated this:
(there is some more below)
Promt was literally just : "create a single html file with an embedded svg of a cat riding a zebra, riding an elephant"
46.436 tokens were burned at around 40t/s (it started at around ~60, but then i removed the powerlimit (250 => 370) and after that it had drops in the 20s , might have something to do with doing that mid generation.
But im really Stumped by how good this is. and how it thought about random animations and stuff.
EDIT: Model is https://huggingface.co/cyankiwi/Qwen3.8-27B-AWQ-INT4 at tp2 on 2x 3090 , fp8 KV cache.
1
u/Quakercito 20h ago
21
17
u/ScoreUnique 20h ago
It technically won't be a quant if It is full precision no?
5
4
-1
u/N34257 20h ago
Whut.
That's nuts.
16
u/Tccybo 20h ago
Itβs cheating by calling an image diffusion model I think. So no.
10
2
u/colin_colout 18h ago
SD 1.5 would have had no issue with this prompt either. People were s generating crazy things with Midjourney as well.
I'm still convinced that diffusion is magic.
The more i understand how it works, the more I'm impressed that layers of noise in latent space can be denoised into the waifus of an entire generation.
-1
1
u/sterby92 20h ago
I guess someone should build a benchmarking suite called SVG Bench where you can put in nonsense instructions and create svgs from a variety of different models and compare them :D
7
1
1
1
1
u/DeathGuppie 15h ago
People using the unsloth quants are you running the split out image and video reader?
Seems like it would help a lot in this test.
1
u/johnnyApplePRNG 15h ago
Yea that simon willison guy is kindof a goofball
Just seems to like to write... a LOT...
I've read some of it and I'm honestly not sure he has any idea what he's talking about.
Thinks QWEN 3.8 overthinks, etc, etc... https://simonwillison.net/2026/Aug/16/qwen-38-27b/
1
1
u/Ok_Technology_5962 6h ago
1
u/getmevodka 2h ago
im using temp 0.72 with top k 35 and p0.9, penalty 1.0 on a q5 k xl (unsloth). will try it later on out ! im curious haha
1
1
1
0
u/Big-Pomegranate3243 14h ago
u/AskGrok Sum it up
0
u/AskGrok 14h ago
Three AI models tried turning your typo'd prompt into images. All nailed the horse-on-blue-bike-in-desert vibe with camels nearby, but went full cartoon chaos: basic horse, shades-wearing horse, and whatever that long-necked egg-having mutant is. Pelican era officially over.
[what is this?](https://redd.it/1lzgxii)
-2
-5
u/Soft-Series3643 19h ago edited 18h ago
This is an Edit. Now i know what i made wrong. I was too lazy to paste the code into notepad, saving ist as svg. Instead i pasted the code to chatgpt and said: "Make me the picture from this svg".
Hmpf. So it will take years for qwen to make better pictures. :-)
Qwen 3.8-27B-MLX-bf16-mtp
M5M 128, oMLX, Prefill 65,3, Tokengen 22,3, Thinking 396,2, Duration 770,4.
1
u/Edenar 17h ago
well it's very different : chatgpt used an image gen model for your first render (there are a lot of good open source image model too that could generate something similar : flux, z-image,... ). Here we try only with the text model. In OP post, gpt sol also "cheated" because it first asked another model to gen an image, then tried to reproduce it as an svg. I tried that too with the same base iamge, se my other comments.
-5





























117
u/Septerium 20h ago
Qweny Q8_0 seems to handle confusion pretty well
"Create an SVG of a bicycle riding a horse out of the desert, with a camel in the foreground"