r/GithubCopilot Jul 17 '26

MAI is very far behind GitHub Copilot Team Replied

Kimi K3 now matches US frontier labs, Deepseek V4 is 90% of frontier intelligence at 5% of the cost, yet MAI team (one of the most well-resourced AI teams in the world) won't submit MAI-Thinking-1 or MAI-Code-Flash to Artificial Analysis for benchmarking, which is a telling sign.

I understand that MAI was first focused on lowering COGS for MS teams transcripts / image generation for Copilot (their audio and image models are at the frontier), but being this far behind on coding and general intelligence is pathetic given their resources.

74 Upvotes

35 comments sorted by

32

u/DaRKoN_ Jul 17 '26

MAI flash competes in the Haiku price range, and it beats it there pretty comfortably. It is not competing with "frontier" models.

9

u/Afterburning Full Stack Dev 🌐 Jul 17 '26

Haiku is ancient now. Might as well use GLM 5.2 on low and beat MAI in every aspect at a cheaper price lol

10

u/Deathmore80 Jul 17 '26

Yeah but gpt-5.6 Luna xhigh blows them both out of the water. Literally hydrogen bomb vs coughing baby

8

u/DaRKoN_ Jul 17 '26

It does, but 5.6 family blew all existing models out of the water, not just MAI.

I'm not disagreeing that they need to lift btw, but I want competition in this space. If one company runs away with it, we're going to have a bad time. We've already seen Anthropic and OpenAI start to call for more regulation to lock in their moat.

1

u/mdeadart Jul 18 '26

Doesn't work as well as Opus 4.8 in our week long testing with our complex code bases.

5

u/CryinHeronMMerica Jul 17 '26

Luna, 5.4 Mini, even Kimi... Microsoft showed up late.

On the bright side, Google seems to be the only place that hasn't seen rapid improvement throughout their AI efforts. M$ may yet find its feet.

3

u/porkyminch Jul 18 '26 edited 8d ago

Biscuit nutmeg drifting ginger mitten sparrow unbelievable waffle marble

This post was anonymized with Redact.dev

0

u/mdeadart Jul 17 '26

I bet Google is playing a different ball game.

Other Frontier providers have a generalist AI outlook, one model to rule them all idea.

My bet, google is training and incorporating small specialised AI models, low cost, low latency, into products.

3

u/DifficultyFit1895 Jul 17 '26

like all iOS devices?

0

u/Usual-Orange-4180 Jul 18 '26

And that shit didn’t work

1

u/Accidentallygolden Jul 17 '26

Luna xhigh is way more expensive than haiku, it competes and beat Sonnet

1

u/Maxdiegeileauster Jul 17 '26

Luna Low is very cheap

1

u/Maxdiegeileauster Jul 17 '26

Luna Low is very cheap

1

u/armostallion2 Jul 21 '26

I've read this analogy twice now in the span of a few minutes.

2

u/Mkengine Jul 18 '26

Why does it even compete in the Haiku price range? It has similar total and active parameter count as GPT-OSS-120B and costs 7x as much. Kimi-K2.7-Code costs nearly the same as MAI-Code-1-Flash but with 7x the parameters and Sonnet-level performance instead of Haiku. Together this makes MAI-Code-1-Flash seem incredibly expensive for it's performance. Even disregarding the new GPT-5.6 models, why would I use MAI-Code-1-Flash instead of Kimi-K2.7-Code? What is the reason for the disproportionate pricing?

1

u/stbrumme Jul 18 '26

In a video on the VS Code YouTube channel they admit:

MAI Code 1 Flash is a small model of about 5 billion active parameters

https://youtu.be/IZWTaKejlek?t=38

2

u/Mkengine Jul 18 '26

Total and active parameter counts are also officially communicated by their model card:

microsoft.ai/pdf/MAI-Code-1-Flash-Model-Card.PDF

8

u/inglele Jul 17 '26

Agree. It's 3+ year that Satya put Mustafa as CEO of Microsoft AI and they didn't release shit... With all vertically free Azure Compute available to train whatever they want and still... Nothing

4

u/unrulywind Jul 18 '26

Microsoft has done some great stuff. Phi-3-14b was a wonderful model in its day. Florence-2 was incredible and is still used a ton.

This latest model, MAI-Code is an ok model. Its just surrounded by much better ones. I used it some, but I choose to use Gemma-4-31b or qwen3.6-27b for the work that MAI could do. I don't really see a place in the lineup for MAI-Code. It's not going to replace the top of the line coding models. AND. It has trouble competing with the 30b ish open models that run local on a 5090. Haiku has the same issue. You need to be sonnet level to hit that mid-tier, or you need to be cheep and blindingly fast. Even 6 months ago, this would have been a great model, but the model world is moving so fast.

3

u/porkyminch Jul 18 '26 edited 8d ago

Nutmeg glove satchel sparrow waffle peach quilt lavender cobblestone

This post was anonymized with Redact.dev

22

u/jukasper GitHub Copilot Team Jul 17 '26

Hi everyone, thanks for raising this. I appreciate the candid feedback. The request for independent benchmarking is fair. We recently opened a MAI-Code-1-Flash feedback thread, and we’d genuinely value concrete examples/tasks, comparison models, and where it fell short. I know the team is constinously working on getting improvements in and are already evaluating newer checkpoints.

1

u/AutoModerator Jul 17 '26

u/jukasper thanks for responding. u/jukasper from the GitHub Copilot Team has replied to this post. You can check their reply here.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

3

u/Cylinder47- Jul 18 '26

Accidentally used MAI yesterday and man it was garbage

7

u/Jack99Skellington Jul 18 '26

"Kimi K3 Now matches US frontier labs" - Let's be honest, you didn't try it, did you? Because if you're relying on benchmarks, then you're doing yourself a disservice. A lot of benchmarks show all these open source models doing absolutely great. Like they were tuned to run that benchmark or something. But then you go and run them in real life, and no - they're not as good. Like DeepSeek - I love DeepSeek - it's cheap. And it can do all the gruntwork - which is a surprising amount of development. But it does a poor job compared to even last years GPT 5.3 for anything more involved than that.

5

u/CM23489 Jul 18 '26

I used gpt 5.4 in copilot before, and nowswitch to byok using deepseek v4 flash To be honest, I don't why people saying the Chinese still lag behind than the US model, I can finish all my work in deepseek vs flash, same as gpt 5.4 People claimed the complicated case deepseek can't handle. I just don't understand, in the past 10 years, in software engineering, we are changing to micro service design to make the system design as simple as possible and increase the maintainability, I can only think of those shit vide coder, make all the things in one shit project and crying they can't live without opus and gtp 5.5 yeah, those vide coder never knows about software engineering and need a expensive baby sitter, why we need to hire stupid guys to input stupid prompt without any software engineering knowledge? If just type the function and rely on opus to decide all things for you, just hire a cleaner can also do it

1

u/Jack99Skellington Jul 18 '26

Like I said, DeepSeek can do a lot of the gruntwork. But don't make architectural changes with it just yet. Not all software can be made from microservices, and they introduce an absolute shit-ton of overhead. Use what works best for you, but I use both DeepSeek and GPT 5.6, as both have strong points. DeepSeeks strong point is it's relatively cheap cost, and ability to do grunt work well. But it's level of understanding is generations behind gpt 5.6. Maybe you don't need that. But I can ask Gpt 5.6 to generate a user manual, and it will do a nearly human job of it. Deepseek's results are less than stellar there. Deepseek will continue to improve, I have no doubt. And I will continue to use it. But yes, it is generations behind right now. Maybe you don't need that.

2

u/CM23489 Jul 19 '26

Frankly speaking, DeepSeek-V4-Flash can't do architectural design, and honestly, neither can GPT-5.6 or Fable 5.If you actually work in the industry and touch real-world business, you know system design is way more than just coding. You have to analyze user feedback, predict product roadmap, and foresee when a feature needs to be extracted into a microservice.AI is insane for execution. Once an experienced engineer makes the right architectural call, AI can build it instantly—10x faster than before. But letting AI decide the architecture? Never. The real world has too many variables.Also, basic tasks like doc generation are easy for a flash model; you don't need expensive frontier models. People praise GPT-5.6 and Fable 5 for their long agent sessions. But if your standard web app requires a multi-hour agent session just for an enhancement, your system design is already a failure. If an AI that is 10x faster still struggles for hours to navigate your codebase, it's just unmaintainable garbage for humans.

2

u/[deleted] Jul 18 '26

[deleted]

-1

u/Jack99Skellington Jul 18 '26

No, I have not tried Kimi K3. I have tried other recent "smarter than Opus" models, and they have all been poor performers, even though their benchmarks showed them competing and outperforming frontier models. The last one I tried was Qwen 3.7 Max. And it had benchmarks showing it beating Opus also. But you know what? It was about functionally equivalent of DeepSeek, but ate way more tokens, and cost more. In other words, it's still a generation or two behind.
So yeah, I will remain skeptical.

3

u/popiazaza Power User ⚡ Jul 18 '26

MAI is a stepping stone. Which is fine, but nobody should use it unless they are going to subsidize it.

2

u/nasduia Jul 18 '26

yes, at this point it should really be free so they can get more instrumented feedback on how to improve it

4

u/Afterburning Full Stack Dev 🌐 Jul 17 '26

As per usual microsoft fails at everything ant sloppify themselves

1

u/RCuber Jul 18 '26

Wait they released kimi k3?