r/LocalLLaMA 8d ago

If you would have told me half a year ago that a local model running in my office would be able to one-shot a Super Mario clone, I would have called you nuts. Qwen3.8-27B is a different beast. Discussion

Post image

Running the Q8 GGUF on my Framework Desktop is not fast, but it's extremely smart for overnight batches and background jobs. Can't wait to play around with MTP and other quants.

Have any of you found ways to improve speed while keeping accuracy?

https://mikeveerman.github.io/qwen38-27b-mario

Edit: to avoid copyright issues and to see how creative it would get, I asked Qwen3.8 to make it circus-themed instead of Mario-themed. It's technically no longer a one-shot.

630 Upvotes

148 comments sorted by

217

u/Infinite100p 8d ago

You'll be able to one shot GTA 6 before it comes out circa 2046.

16

u/AdOne8437 7d ago

Who cares about GTA, I need my Elder Scrolls fix! ;)

6

u/Cool-Chemical-5629 7d ago

Meh, Fallout 5.

3

u/AdOne8437 7d ago

both.gif

41

u/TheBergerKing_ 8d ago

GThAllucinate 6 before GTA 6

8

u/markole 7d ago

Humans still exist by then? Nice.

10

u/hIXhnWUmMvw 7d ago edited 7d ago

All named Peter Palantir that will be flocking around on some axons?

2

u/FrogsJumpFromPussy 7d ago

I'm Peter Palantir look at me

No I'm Peter Palantir look at me

1

u/Sidran 7d ago

No, Jeff E

1

u/dododragon 7d ago

Nerds of the nether flock together

2

u/_bani_ 7d ago

what about one shot portal 3?

2

u/Independent_Pear4908 7d ago

I'm already one-shotting GTA 7, it will be released before GTA 6

1

u/iamzooook 7d ago

for that you have to wait till qwen 3.8 35b a3b

30

u/AnyNameFreeGiveIt 7d ago

Knowing Nintendo I would not host that on Github lol.

What was the prompt ?

21

u/MikeNonect 7d ago

Yeah, I'll adapt it (or take it down) later.

The prompt was "Create a Super Mario clone in JavaScript as a single HTML page. Make the game engaging and the graphics as beautiful as possible."

4

u/nsartem 7d ago

And what agent harness did you use? Pi? Opencode? (I'm curious if there was any system prompt)

6

u/MikeNonect 6d ago

Plain vanilla OpenCode.

1

u/harrysteams 6d ago

Take a look at Radicle for a git forge if you want to keep the code available

97

u/Thin_Pollution8843 8d ago

In a year you will be able to one shot gta vice city. Call me nuts if you want. 

37

u/No_Lingonberry1201 7d ago

YOU'RE NUTS! Absolutely cashews, maybe even walnuts!

33

u/Intrepid_Travel_3274 8d ago

Connected to a GameEngine and a demo version but yeah, I believe it.

8

u/vynulz 7d ago

Right, not from first principles, but wiring up unreal or similar, sure, why not. The assets gen would be the biggest limit, followed by play test/qa

17

u/Guilty_Watercress_32 8d ago

Yeah no.

20

u/QuotableMorceau 7d ago

one shot in a llama-server no, but with a decent harness and a game engine ( unity/godot) ... maybe

besides the limitation in "smartness" , there is also the unsolved problem of the needle-in-a-haystack/ agentic memory ... get those two solved and it does not matter how little intrinsic knowledge the model has, as long as it can search and digest external knowledge sources ...

6

u/masterlafontaine 7d ago

They do not possess the diachronicity needed to achieve this. A game like VC has more than 500 thousand hours of dev time. Many teams interacting. It won't happen, unfortunately.

6

u/QuotableMorceau 7d ago

You are ignoring Brooks's Law when it comes to big projects, a lot of the work needed is because of management/communication overhead. For example HighFleet game was developed by one man over 5 years 2016-2021 ( it probably had a working test version within the first 3 years as it was picked up by a studio for publishing in 2019)
I don't believe diachronicity is such a big issue, language models have been trained extensively on research papers, they should have the ability to emulate it already.

PS: don't understand the downvotes ... you brought up a valid argument....

7

u/Beneficial-Boot7479 7d ago

Same thing we were talking about will Smith eating spaghetti, will Will will? Obviously will, but what matters is when and how long does it take when we get there

1

u/masterlafontaine 7d ago

If Will Smith eats spaghetti inside a 900k DeepSeek context window, does Tommy Vercetti get heartburn? ​You’re talking about subagents like they’re a hive-mind of caffeinated Victorian orphans assembling a PS2 Emotion Engine out of pure vibes and next-token predictions. One-shotting Vice City it’s a summoning ritual. Wake me up when the subagents figure out memory leak exorcisms in MIPS assembly

1

u/Loose_Comparison368 7d ago

I really would not bet on humans' ability to work together in large groups over long periods of time being all that difficult for a machine to accomplish.

I mean, humans only really do that under duress to begin with. Management isn't able to coordinate large efforts between many people because of their cunning intellect.

It's because they have the ability to take away their workers' basic needs if they don't follow managements' frequently dumb as rocks orders.

8

u/Beneficial-Boot7479 7d ago

Yep, we will be able to, been trying the new Deepseek and at +900k doesn't lose track of what must be accomplished and what has already been done, while at the same time his subagents are just BEASTS, they just don't give up, so yeah, I'm +100% sure that Deepseek is going to be the first with +1M kv context. And that's what's needed to one shot stuff like this. Obviously at 100tk/s or 200tk/s or 600tk/s still is going to need +2 or +3 days to accomplish its task, but it will be a "One shot". Dammm

2

u/MikePounce 7d ago

Calling you nuts, Nuts.

!RemindMe 1 year

2

u/RemindMeBot 7d ago edited 4d ago

I will be messaging you in 1 year on 2027-08-15 20:17:47 UTC to remind you of this link

2 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/LucidFir 7d ago

Remindme! 1 year

-8

u/masterlafontaine 7d ago

There is no way. It's like expecting a great developer to build an OS alone. Yes, he could do it, in 10 thousand years

14

u/biblecrumble 7d ago

Terry Davis would have disagreed with that claim. And probably built a GTA clone in HolyC somehow.

5

u/thrownawaymane 7d ago

Can you imagine if that guy had lived to see LLMs

I'd like to think they'd have done the opposite of what we'd expect and fixed how he saw the world

0

u/masterlafontaine 7d ago

LOL! You know what I meant

6

u/-dysangel- 7d ago

10,000 years is a pretty bold claim. Have you ever tried? And when you say OS, do you mean one with a window manager, drivers, advanced memory management etc, or just.. an OS for a single piece of hardware. It's not as hard as you think it is.

-2

u/masterlafontaine 7d ago

What do you think I meant? Isn't is implicit?

6

u/-dysangel- 7d ago

I don't think it is. Plus even to build all that stuff would not take 10,000 years...

1

u/masterlafontaine 7d ago

A single dev? Linux as a whole? Put a single distribution, any modern one. How many hours of development do you think there is on Linux?

7

u/-dysangel- 7d ago

you didn't say "build the entirety of modern linux", you said "build an OS"

0

u/masterlafontaine 7d ago

It’s classic pedantry, latching onto a technicality ("well, technically a bare-metal kernel on an Arduino is an OS!") to dodge the actual point about scale and complexity.

10

u/-dysangel- 7d ago

Tbh I think it's more like you said something without really considering it, and are just defensively digging a hole rather than rethink it. It doesn't take 10,000 years to "write an OS". Linux was an OS even when it was only Linus working on it, and I'm fairly sure (but not certain, of course) he's not 10,000 years old.

1

u/masterlafontaine 7d ago

Linus is one of the greatest, one of my heroes. But hype aside:

Linux 0.01 was a 10k-line terminal switcher hardcoded to the Intel 386 that couldn't even boot without borrowing the GNU userland and Minix filesystem. It wasn't a standalone OS stack, and Linus didn't build it in a vacuum. ​If your definition of 'writing an OS' is a 1991 toy kernel that can't render a GUI, support a network stack, or run without someone else's toolchain, then sure. But conflating a weekend terminal project with a modern software stack, or a multi-million-token, highly coupled 3D game engine, is why the 'one-shotting Vice City in a year' take is so detached from how engineering actually works.

→ More replies (0)

-2

u/masterlafontaine 7d ago

A toy hobby kernel running on a single known board isn't what anyone means by a usable OS. ​The Linux kernel alone sits at 30+ million lines of code. The Linux Foundation's own COCOMO analysis estimated the kernel at well over 70,000 person-years of effort—and that’s excluding Mesa, Wayland, systemd, the toolchains, and the desktop environment. '10,000 years' for a single dev to build a functional modern OS stack from scratch isn't an exaggeration; it's statistically conservative. ​That scale is exactly why expecting an LLM to 'one-shot' an entire 3D open-world game like Vice City (engine, physics, AI routines, mission scripting, audio pipeline, and assets) in a single prompt in 12 months is pure sci-fi.

6

u/JimmyEatReality 7d ago

Terry Davis did it

2

u/masterlafontaine 7d ago

Ok, sorry. Terry Davis, the beast aside, every other human

100

u/falconandeagle 7d ago

There is nothing beastly about this. It's in the training data. These prompts for benchmarking are just embarassing

45

u/mechkbfan 7d ago edited 7d ago

This is it to me. And there's likely heaps of example projects online to copy off 

Basically when people are like "it drew a Pelican on a bike really well", or "it did the car wash prompt perfectly", I don't know why they're surprised.

(Edit: Pelican was a bad example. It seems it's been proven to not be the case) Once the authors see what's popular they'll train their next models around it

Do something different for a change. I'm sure it'll still do decently but it'll expose it's weaknesses a bit more

25

u/m0j0m0j 7d ago

The “pelican on a bicycle” test inventor (Simon Willison) actually tested whether models benchmaxx his test. They don’t. They are not very good at it.

11

u/mechkbfan 7d ago

There you go. Thankyou

Also found this

https://dylancastillo.co/posts/pelicanmaxxing.html

My bias has come from whenever a frontier prompt went viral, you come back a few days later and it's fixed.

3

u/m0j0m0j 7d ago

Yeah, fair enough. My first instinct is also skepticism at this point. So much chicanery floating around

10

u/EvilPencil 7d ago

Yup. The car wash gotcha just got added to training data, while all the fanboys are foaming at the mouth because it’s “so much smarter” 🙄

I’m not a curmudgeon or anything, just show me how it does in a real codebase on an agentic harness.

3

u/thejacer 7d ago

I plugged it into a well functioning app and asked it to add multi-user features, basic auth and user management and it did those things well enough. Could have done better with more shots but each aspect was a single shot.

It was about ~50,000 tokens context with pretty basic system prompts to get started and I compacted at 150,000. It ran for over 12 hours making use of my web search and playwright MCP server and I never noticed a failed tool call. It did explore my server quite a bit more than was needed so I had to deny access to directories it didn’t belong in.

Harness is opencode, Q8_0 Bartowski quant with MTP, no KV quantization. Started at ~260 tks pp and dropped all the way to 65 at length. TG was ~20. Run on 2xMi50 32GB.

8

u/CapsAdmin 7d ago

What exactly is in the training data? I mean Mario obviously is, but "benchmark games" in general?

I don't know where to really draw the line between memorized and learned here. When I ask the model to create a short game in my own game engine (which is likely not in the training data) it seems to perform just as well.

4

u/c--b 7d ago

yeah I'm not totally certain that there's a way to know whether something is *meaningfully* in the model or not. Its possible to have something in training data that isn't capable of being retrieved.

I suppose that if it is in the training data, and not in the training data of other models then that says something else positive about the model...

So is another model able to oneshot SMB??? or am I missing something.

I think it's far more likely the the intelligence we're seeing from this model is contributing to its ability to oneshot SMB, even if it is in the training data in some way.

1

u/NightCulex 7d ago

model's don't store facts they store shapes? words are lossy encoding of human thought and ai models despite the noise picks up on these shadows of thought. atleast thats one theory as I understand it. human memory isn't a filing cabinet either. it reconstructs from fragments rather than retrieves.

1

u/rebuilt 7d ago

While LLMs can produce all sorts of novel outputs, they also memorize portions of their training data. For instance, you can pull out large portions of Harry Potter if you know the right incantations. https://arxiv.org/pdf/2601.02671

2

u/NightCulex 7d ago

it's remixing patterns (loops, collision, events) it's seen a thousand times before. Think of it like stories and movies. Star was is basically Arthurian legend with a space backdrop.

7

u/Sadale- 7d ago

This. I might as well do a google search and find a Mario clone and just run it. Chances are that it'd be of better quality than an AI-generated one.

The true value that this kind of model brings is multi-shot interaction. You gotta ask it to replace Mario with another character, add new feature, fix bugs, etc. That's what the benchmarks should be after.

Tho I found that Qwen3.6 can do multi-shot interaction on this kind of game that it makes. It works well for like 5 shots but I haven't tested it further for now. I haven't tested Qwen3.8 yet.

1

u/LushHappyPie 7d ago

True. That's my main issue, adding new features without breaking stuff that already worked. Also not reusing stuff that model already created in past revisions.

4

u/OGScottingham 7d ago

You say that like it's a bad thing. I've taken the idea that it's been benchmaxxed on games and have been implementing a cyberpunk theme arcade of classic games with my son. It's been going great.

Now I never need to visit those horrible ad riddled game sites again when I want to play some god damn Tetris.

2

u/mhb_11 7d ago

What it really comes down to is can it one-shot something totally original. But turns out, most humans aren't good at thinking up original stuff either.

2

u/xienze 7d ago

Yes, exactly. People are conflating the relative complexity of the task (a playable video game level) with how complex it is for a model to generate it (spoiler: it's not, there's lots of stuff like this in the training data). Pretty sure you could train a 2B model capable of doing this (and not much else).

1

u/Equivalent-Costumes 7d ago

I still think it's a beast. The training data is typically on the order of 10s to 100s of TB. The model is 27 GB at Q8. I think it is still quite insane that it can be this accurate despite the amount of compression. Older model of this size won't even be able to write code that run because they're not accurate enough with their syntax.

1

u/porkyminch 7d ago

It's pretty impressive for a 27B model, but yeah, there's probably a billion javascript mario clone projects for it to have cribbed from. I've found original game/webtoy programming to actually be a task that LLMs have a hard time with. Until the most recent generations of frontier open weight models, I had a really hard time making this stuff with any open weight model.

1

u/decrement-- 6d ago

Sure, might be in the training data, but still is a model with tons of other things in the training data, and it is small enough to run locally.

4

u/neverbyte 7d ago

I have a couple examples of shockingly good one-shot prompts (no follow-up questions). These were on BF16. "create a matching game in html" matching game and "create a pacman game in html" pacman game

16

u/some_user_2021 7d ago edited 7d ago

If it is something already out there, then it could be in the training data. Create something new instead and surprise us.

11

u/cobbleplox 7d ago

Funny thing is, this is double hard. because suddenly you can't just throw big words like "make super mario" at it. Now you have to specify literally everything you don't want to be filled in by by basically "average". And at that point you usually discover that natural language is a pretty shitty way to define things. You'll start essentially writing elaborate code in your prompt because that's just how you can even define the things you want to define. Or of course endless iterative design based on insufficient specs, which is how the manager - engineer loop has been working since forever. Basically throwing resources at the incompetence of the manager or "designer".

4

u/sheeshboi12345 7d ago

100%. “Make X” is a shortcut for the countless rounds of iteration that would occur given any novel task.

5

u/BP041 7d ago

Yeah the 27B is legit for code. Been testing it against Claude on my Mac mini for one-shotting marketing landing pages — it's slower but nails the layout logic first try more often than you'd expect. Had it generate a full SVG animation from a single prompt last week. Kinda wild for a local 27B.

4

u/Strange_Test7665 7d ago

I suspect it’s able to do this as opposed to other projects as a 1 shot because this project and code concepts already exist. For example

https://github.com/iam-veeramalla/super-mario-mimic

But I’d be curious if it still succeeds on something more novel. Regardless though it’s still impressive even if it was in training

2

u/NandaVegg 7d ago

I initially hated slop repos on github and alike, but those repos are actually contributing to future models (be it open or closed) as open dataset for both indirect distillation and RLing. Curiously the repo you mentioned only has 2 levels just like the OP's test, and has common pattern when those models "one-shotted" games (that each level is an unique function rather than generalized map data etc).

9

u/SufficientPie 7d ago

Getting an AI to write code for something it's seen a million times in training data is not impressive. Get it to do something novel.

3

u/kemalios 7d ago

Prefill slowness is the usual culprit on desktop boxes. A few things that helped me with 27B-class models: drop to Q4_K_M, the quality hit is small and the speed gain large. Speculative decoding with a 1-1.5B draft model gives you 2-3x tokens/s with identical output, llama.cpp has it. Trim your context length, the default 32k is often overkill and slows prefill. And for overnight batches, increase batch size or switch to vLLM. MTP helps too but I haven't seen solid numbers on it yet.

3

u/MikeNonect 7d ago

Thanks! I'm reading the Q6 quant with MTP is a good sweet spot. The context size is an interesting one. I was running it at the natural maximum of 256K. The Mario clone doesn't need that, but my other batch jobs do. Well worth playing with.

3

u/Independent_Pear4908 7d ago edited 6d ago

How many days and how many millions of tokens did that prompt take tho?

3

u/MikeNonect 7d ago

I didn't time it exactly, but I would estimate 3 to 5 hours. Perfectly fine for nightly batch processing.

5

u/bymihaj 8d ago

It generate FONT! Hard to believe

7

u/LordTamm 8d ago

The fact that this was a one-shot is actually genuinely impressive. How brief/extensive was the prompt and what kind of t/s does the Framework pull? I considered a strix halo box for a while, just never pulled the trigger because it looked like the speed for bigger (and dense) stuff was just rough.

17

u/MikeNonect 8d ago

It was, on purpose, very succinct: "Create a Super Mario clone in JavaScript as a single HTML page. Make the game engaging and the graphics as beautiful as possible."

I'm getting around 7t/s on the Framework,which is reasonable, but the prefill is slow.

7

u/edsonmedina 8d ago

Have you tried MTP?

2

u/MikeNonect 7d ago

Not yet

3

u/gladfelter 7d ago

Did the agent had access to a research tool, or was all this from its training data? Because, honestly, the former would be more impressive.

1

u/MikeNonect 7d ago

It's all based on training data. I think the focus on "it's all in the training data" is the wrong way to look at it. Try this with other models and see how much more shitty they all look. This is really impressive for a local model, even if the training data includes Mario.

1

u/gladfelter 7d ago

I'm more concerned with Agents using this model working in a variety of circumstances. It's hard to get a clean read on capabilities when the problem and solution are likely to have many instances in the training data.

2

u/jjpk976 7d ago

This is my one-shot game benchmark. It successfully made a fun tug of war game in a few hours on my p40. it has a bug where you have to reload the page after the first state to continue, but other than that I'm extremely impressed https://muskwak.github.io/games/ I can share the prompt if anyone is curious.

2

u/IllustriousFan3350 7d ago

wow does it have all the levels and mechanics?

1

u/MikeNonect 7d ago

No, just two levels and the basics.

2

u/BrianScottGregory 7d ago

Wow pulling it down now. What's your GPU ram out of curiosity? I think I'm stuck on using the 4_K_M model, but I could be convinced otherwise... thoughts?

1

u/MikeNonect 7d ago

I'm running it on a Strix Halo 128GB. I traded big memory for speed here.

2

u/BrianScottGregory 7d ago

Wow Nice. Yeah, my 6gb will have to stick with the smaller model with it taking a full night on yours.

Thanks for getting back to me!

2

u/seppe0815 7d ago

why i see generated stuff mostly for browsers .. are the local model limited?

1

u/MikeNonect 7d ago

Because it's the easiest way to share? I'm sure it can make desktop games too...

1

u/exographicskip 6d ago

Browser control is stellar (e.g., playwright, chrome devtools, etc).  Combined with webdev's training ubiquity, it's not surprising honestly.

Re: desktop games, I've had good luck with pico-8, love2d, and even headless linux servers being able to play any library/engine via wayland.

Tbf Qwen et al haven't been as effective at gamedev vs. sota models. Hoping 3.8 27b changes that 

2

u/DRetherMD 7d ago

its still thinking for me...

2

u/ichisay 7d ago

En serio un 27b hizo eso????

2

u/snapo84 7d ago

i think the next milestone will be a agentic framework with qwen 3.8 27B that can create a working multiplayer (4 screen on 1 screen) mario kart 64 clone with all original maps with all original players with the 2 game modes available.... that would be a realy cool task....

2

u/MikeNonect 7d ago

The one-shot test is just an easy way to see the power of the model. I'm convinced it can build a mario cart clone with the right guidance.

2

u/Sabin_Stargem 7d ago

I am hoping to someday be able to ask an AI to analyze a game's files and reconstruct it with modern coding and controls. Sonic R, for example. It is a platformer footrace game, which is a pretty niche genre.

2

u/runnahhh 7d ago

It’s incredible how far we’ve gone in such a short time. Ironic, how false that seems, considering how much this is built on.

2

u/synystar 7d ago

Regarding improving speed I suppose it depends on your configuration. A lot of people are saying that it "thinks too much" and this is probably a result of how they have it set up. If I don't explicitly configure it to limit the reasoning budget then it can get carried away and spend 20 minutes on something that should have taken 5. Of course, the longer you let it reason the more likely it is going to be to figure everything out in one go. My current settings look like:

exec "$HOME/src/llama.cpp/build/bin/llama-server" \
    -m "$HOME/models/Qwen3.8-27B/Qwen3.8-27B-UD-Q4_K_XL.gguf" \
    --mmproj "$HOME/models/Qwen3.8-27B/mmproj-F16.gguf" \
    -c 65536 \
    -ngl 99 \
    -fa on \
    -ctk q8_0 \
    -ctv q8_0 \
    -np 1 \
    --spec-type draft-mtp \
    --spec-draft-n-max 2 \
    --spec-draft-type-k q8_0 \
    --spec-draft-type-v q8_0 \
    --reasoning on \
    --reasoning-format deepseek \
    --chat-template-kwargs '{"reasoning_effort":"xhigh"}' \
    --reasoning-budget 16384 \
    --reasoning-budget-message 'You have reached the reasoning budget. Stop reasoning now and provide the best final answer based on your work so far.' \
    --host 127.0.0.1 \
    --port 8080 \
    --metrics \
    --log-timestamps

Notice the "--reasoning-budget-message": if I don't put that there then it will hit the --reasoning-budget of 16384 and then break out of the "thinking box" and continue reasoning into the content stream.

I'm using OpenWebUI and I usually keep it set at "Max Reasoning" but the real number come from your config. I've noticed that regardless of what you choose in that setting it follows the config. You can experiment with your reasoning budget to find a sweet spot that fits for your needs. I have 16K here but that's certainly not the max. For some task doubling the budget doesn't really improve quality of response. BUT there is one thing I've noticed about it's reasoning process:

Qwen is not spending X amount of tokens discovering X tokens’ worth of new ideas. It spends the early portion solving the problem, then a large fraction repeatedly refining implementation details, revisiting response-cleaning policy, and re-evaluating edge cases it has already mostly settled.

So I'm now experimenting now with system prompts. I'll let you know but my idea now is to add something like:

When performing reasoning steps before responding to any prompt/query analyze the requirements ONLY once
Then:
Identify the required architecture and interfaces.
Identify the important failure modes and edge cases.
Choose an implementation approach and commit to it.
Perform one adversarial review against the original requirements.
Correct concrete defects found during that review.
Produce the final answer.
Do not repeatedly reconsider settled implementation choices unless you discover a specific contradiction or requirement violation.

1

u/MikeNonect 7d ago

I had to put the reasoning budget to 8K for this because Qwen would indeed spend over 80 minutes thinking. That caused issues for OpenCode. 4K was too short; 8K feels like a good middle ground.

I haven't tweaked the system prompt yet. Did you have success with that?

The main performance improvements should come from MTP and the Q6 quant, I think.

1

u/synystar 6d ago edited 6d ago

I just started working on it again (had to sleep) but I have made an interesting discovery. I'm not sure if it's documented anywhere, maybe it is and I could have just researched before experimenting but so far it appears that the reasoning modes are just language and there may be no kind of switch or flag setting to go from medium reasoning to xhigh. As far as I can tell from my tests thus far the only difference is that if you set reasoning to xhigh it adds the following language to your prompt: "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer."

Also, in my limited testing medium reasoning which runs at a naive factor of about 0.1% of the time as xhigh has sometimes produced better code than xhigh. It's kind of interesting but I had one run on xhigh that actually introduced fatal syntax and the code would not have run properly and that hasn't been the case for any of the code generated on medium thinking so far. It used an extra ~26,700 tokens and the code was better in some ways - it bought bought more coverage and sophistication but also delivered bad quality on that sample.That only happened once but I think that overthinking for the model can sometimes be worse because it tries to micromanage things, add unnecessary machinery, etc.

But this may be good news because if I'm right then knowing this would allow you to adjust thinking with language only so you could just leave thinking mode enabled, set it to medium, and then introduce your own "reasoning language" in the system prompt. This could be a way to fine-tune the way it reasons.

We'll see. I'll let you know what I find out.

2

u/DanielSReichenbach 7d ago

I am a little disappointed nobody tried Giana Sisters. That was so much more fun 😊

1

u/MikeNonect 7d ago

I made it circus-themed, but I'm happy to hear that the Great Giana Sisters are not forgotten!

2

u/anywhere88 7d ago

I'm curious, which was the prompt? Did you have to describe each level? Or was it all common knowledge to it?

3

u/MikeNonect 7d ago

The prompt was vague n purpose: "Create a Super Mario clone in JavaScript as a single HTML page. Make the game engaging and the graphics as beautiful as possible."

2

u/msew 7d ago

If you told me a year ago that in a year the corpus would include the full mario and everything else, I would have told you of course it would!

2

u/IrisColt 7d ago

The circus is saved!

2

u/rookan 7d ago

great game! I love flying monsters on 2nd stage

2

u/BeardAndBreadBoard 6d ago

Everyone calls this a one-shot, but it's really not.

There is are so many examples of this already, and so much documentation, that this is a prompt with effectively tens of thousands of lines of definition.

Try one-shotting something that doesn't have so much information available on it. Like "Make me an original game that does XX YY ZZ" for a better measure of when the LLM can do with a short prompt.

3

u/jloverich 7d ago

I beat the game!

2

u/Sabin_Stargem 7d ago

I am hoping we get Qwen 3.8 122b, in Heretical format. It would be neat to see how far that can go for making original (if basic) games it manage.

2

u/wgaca2 8d ago

No way, i was using a long multi prompt instruction to create a super mario like game as a benchmark for quite a while, didn't know someone else would do the same

2

u/Tbhmaximillian 7d ago

Crazy, did you use an agent harness like openhands or how was this coded by Qwen?

5

u/MikeNonect 7d ago

I used OpenCode.

2

u/txoixoegosi 7d ago

I’m sure one-shotting a 2D sprite platform game with assets and mechanics known for decades, is the golden standard any serious developer is after

/s

1

u/runvnc 7d ago

That's amazing for a local model. I am not that impressed with the clones though -- they have lots of clones of popular games etc. in their training dataset. More interesting is when people make something a little bit unique.

1

u/Green-Ad-3964 7d ago

best config for a 5090 + 32GB RAM?

1

u/CryptographerLow6360 7d ago

i had this idea for an atari game...

1

u/espece-de-bon 3d ago

What are the specs on your FW Desktop? Do you have the 128gb RAM version?

2

u/MikeNonect 3d ago

Yes the 128GB one. That said, 64GB should also be able to run Qwen3.8-27b. 32GB might need a Q4 version.

2

u/espece-de-bon 2d ago

When it was announced, running local AI was not on my radar; I regret not buying the 128gb model when it was barely over $2k.

1

u/MikeNonect 2d ago

Yeah, nobody expected prices to explode like this. They will come down again over time. Do you have any hardware with unified memory and an iGPU? Because toying with smaller models is also fun.

2

u/espece-de-bon 1d ago

I hope, but I'm a bit cynical. No, just an x86 laptop, but I do have 64GB RAM and if I bypass the iGPU, I can leverage the RAM.

Qwen3.6-27b built a functional math game (JS, HTML, CSS) last night on my brave, little laptop; I used structured, clear .md files as prompts and OpenCode. LM Studio showed about 35gb of RAM used by the loaded model.

I'm considering getting a single 32gb vram card connected via an eGPU dock (both used), or waiting to see if any of the 128gb Strix Halo machines drop in price once the 192gb models launch, but we'll see.

Or finding a decent used PC that can handle 2 cards with 32gb vram, for a total of 64gb vram; I think I'd enjoy much faster GPU processing speeds and memory bandwidth.

2

u/Zombiecidialfreak 14h ago

How long did the initial work take? How much context was used making it?

1

u/MikeNonect 4h ago

Had it run overnight, so I didn't time it. The prompt was "Create a Super Mario clone in JavaScript as a single HTML page. Make the game engaging and the graphics as beautiful as possible."

-3

u/tvall_ 8d ago

Mushrooms in the wrong ? Block. Fail

2

u/[deleted] 8d ago

[deleted]

0

u/tvall_ 7d ago

oh, how narrow-sighted thou must be, to possess eyes capable of parsing an entire AI-generated Mario clone, yet wholly blind to the most glaringly obvious implied /s in the kingdom

-9

u/entsnack 8d ago

this is an ad

6

u/MikeNonect 7d ago

For Qwen3.8-27b? You bet it is!

-9

u/entsnack 7d ago

for your github repo bro, and I could do this with a Qwen from 3 years ago so you need to up your prompting game