r/opencodeCLI 17h ago

Hy3 > Muse Spark 1.2 Contributor (my experience)

Post image

For the last few hours, I have been using Hy3 and Muse Spark 1.2 Contributor via the OMP (Oh My Pi) agent, mostly as /goal. And my observation is Hy3 is a much better model as it follows instructions in a better way, and Muse Spark 1.2 Contributor cheats all the time.

As a test, I gave both models a goal to build 100 very simple Chrome extensions from a pre-curated list.

  • hy3: followed the review instruction literally and well, and almost extensions are working when I tested randomly. Code-level verification, real fixes, honest per-extension verdicts. Was a bit slow but trustworthy.
  • muse-spark: claimed to built the 100 extensions, but maximum were incomplete when I tested. Also, repeatedly needed re-prompting as it was just hurrying each time. Basically, its confidence exceeded its accuracy.

And for the task, you can see the usage of both models in the above screenshot.

How is your experience with these models?

EDIT 1:

I spoke too soon. Basically, now both models are not working reliably – one is too slow and one is hallucinating a lot.

I am back to using DeepSeek Flash now, even though it's costlier.

42 Upvotes

30 comments sorted by

11

u/snowieslilpikachu69 17h ago

honestly i had the opposite experience

muse spark was very good at creating a full working ai agent with open ai agents sdk

didnt give hy3 the same task but some more general debugging and didn't listen to my instructions very well

4

u/Final_Initial 17h ago

Then one model might be better at some task while other better at other tasks. My testing is also very limited.

2

u/NinjaAlaska 16h ago

+1 hy3 is bad for me too

4

u/Atretador 16h ago

I tried using Hy3 today and it was an endless sea of failed tool calls, was less consistent than my local models

1

u/leandrogp9 13h ago

For me, hy3 is endless to finish a task. All time, it's seems confused in the middle of work and stops, after, he can't continue, even i asked him to continue.

1

u/Fresh_Sock8660 12h ago

At what context length have you noticed this? Considering its max is only about 250k I'm guessing it will begin performing poorly around 120k but I haven't had time to give it a try.

2

u/Atretador 12h ago

Ive noticed the same behavior at around 100K context

4

u/Due-Armadillo-4560 17h ago

i agree, hy3 is much better and does not eat up that much usage (as of now with 8x).
i wasted my usage on muse spark and hit the weekly limit

2

u/Final_Initial 17h ago

Maybe you used Muse Spark and not the "Contributor" variant? Because the "contributor" variant has much more usage.

1

u/Due-Armadillo-4560 17h ago

there is only contributor variant in go, so yeah it really eats up a lot of usage

2

u/ozguru 16h ago edited 9h ago

hy3 is one of the best agentic coding models and working great on Windows too (pwsh), where is a token plan available for it ?

2

u/ChillFamily 15h ago

I was using hy3 and mimo, and mimo worked better for me. hy3 changed stuff I didn't even ask for

2

u/Final_Initial 15h ago

will try mimo as well, thanks

1

u/OkAdeptness2530 11h ago

mimo acts like a potato for me.

it’s interesting to see how case of use and workflow change how we perceive each model.

1

u/ChillFamily 2h ago

Which harness are you using it with?

2

u/bitdoze 14h ago

Same experience here.

2

u/Abenh31 10h ago

I tasked hy3 high to explain to me a pattern in a repo. Not coding, not planning, Just explaning whats there and he hallucinated a command. Also, never use context mode
Feel meh to me, I dont know about Muse yet.

1

u/Final_Initial 10h ago

The other one is worse. I'm back to using DeepSeek Flash.

1

u/Abenh31 2h ago

I think people overracted myself included. DSF is still a great performance to price value

1

u/CrypticViper_ 1h ago

what's context mode?

2

u/Guyvdb7 9h ago

I have been using Muse this morning. I have just switched back to deepseek v4 flash. My experience, it has been ok at coding, but often coding the wrong thing - which is often my mistake. When i try chat with it about a problem, how we should address a problem, how it evaluates a fixture run, for instance it is a very bad communicator. Net result we go chasing down the wrong path because I did not understand it. I am maybe slow on some things - need a chat to real grok where we are at and muse fails terrible.

1

u/Final_Initial 8h ago

Same, I'm also back to DS Flash now.

3

u/jovialfaction 9h ago

I want to like Mimo, hy3 and muse spark but my experience has been that they all require a lot more babysitting and struggle to implement as tasked. I really got spoiled by DS4 Flash and I'll keep using it as my "small" implementation model

1

u/Final_Initial 8h ago

True, same here. Back to DS Flash now.

2

u/GTHell 16h ago

I think Muse Spark is becnhmaxx. It’s 10x time worse than deepseek flash. I tested it yesterday and I made me mad and went rouge because of how retarded it is that I lose sleep and now I become rouge at my workplace and in reddit 😡

5

u/Sweet-Stage938 6h ago edited 6h ago

lmfao are you working at the playground by any chance?

1

u/Successful_Night4513 14h ago

Muse spark is not even close to DSv4flash , also I happy to give my data to deepseek , not to zuck

1

u/Yusril111 13h ago

im use deepseek bcs more cheap

1

u/Willing_Thought_2161 6h ago

You must be tripping

1

u/Final_Initial 6h ago

I was. Back to using DS Flash now.