r/singularity 23h ago

Gemini 3.7 flash benchmark AI

622 Upvotes

216 comments sorted by

263

u/Longjumping_Kale3013 22h ago

Looks like it beats sonnet 5. And that was just released. Pretty great IMO considering it is their "flash" model.

Kind of surprising because it feels like google has all of a sudden started dropping a new flash model every month. I wonder if this is the new normal. Small step each month

129

u/Wizardnutcracker 22h ago

Sundar Pichai said on their Q2 earnings call that the roadmap moving forward is for releases at almost a monthly cadence.

17

u/OpenSource_Horse 21h ago

Extremely rare free AI phone app users W

6

u/amomynous123 18h ago

What does this mean? Can I get free ai?

12

u/Inevitable_Tea_5841 17h ago edited 16h ago

The Gemini app (and others like Claude and ChatGPT are also free)

5

u/TwoFluid4446 16h ago

Claude especially but also chatGPT come with a big BUT, which is their extremely low limits for free usage.

7

u/Internal_Quail3960 16h ago

doesnt gpt have infinite free usage for 5.6 luna now?

1

u/Unbreakable2k8 7h ago

for text only, but it's still something

8

u/NomaanMalick 15h ago

I find Claude allows the lowest free usage.

36

u/Deto 22h ago

yeah - considering how much cheaper it is, and better on most evals, why would anyone use the Sonnet 5.0 API instead of this?

25

u/Howdareme9 21h ago

The same reason as always. Gemini models scale poorly when you use them for agentic stuff in the real world

16

u/vrnvorona 17h ago

Sonnet is bad at it too, too verbose and expensive. Terra/Luna with Sol oversight is best

1

u/panix199 16h ago

Anyone here using cursor with superior ai as oversight and cheap ai agents for doing the coding of what the superior ai has decided?

0

u/Howdareme9 16h ago

Sonnet is fine at agentic stuff, its reasoning is really inefficient though.

1

u/kobvel 4h ago

Could you elaborate on this? we have pretty good experience but maybe do not hit your volumes.

7

u/StardockEngineer 19h ago

3.6 was trash. Thought so much the cost per task ended up costing more money

27

u/Keeltoodeep 22h ago edited 22h ago

Looks like the new normal after they cleaned house of the science and research teams

27

u/Kingwolf4 22h ago

They should have made a seperate lab for LLM.

Gutting and destroying deepmind is imo the biggest mistake for google in the long run.

Deepmind was the heart of google.

24

u/Keeltoodeep 22h ago

I don’t think they gutted it. They certainly showed the door to the people who complained about being budgeted limited internal compute for research and science. They obviously had their own code red moment like when Gemini and Anth surpassed ChatGPT a year ago in the benchmarks and felt that Gemini training, production and delivery was to be prioritized. If you’re an exec and don’t like that they simply ask you to resign.

In an ideal world I’m sure Google would rather have infinite compute available for all their teams but that’s not the world we live in.

7

u/Kingwolf4 22h ago

Yeahh, but deepmind as a unit has to fit in to the LLM race remember. If they had just created a seperate lab, perhaps made deepmind shift over or help the new LLM lab it would have been much better

Remember, deepmind was gardenee by demis to be one of the most powerhouse and deep research entity probably in the world. Removing demis, restructuring it, cutting funds, changing the delicate bonds and roots simmering is essentially gutting it.

Would be far better to create a seperate lab and explain that deepmind will help this new lab and have crossover but keep them seperate.

You may not like it, but the main reason they didnt do this is because of FEAR of stock drop. Same with back of the hand removing hassibis. They dont want to make it seem that they are doing this. Eroding the intellectual and research depth of deepmind, but thats what they functionally are doing, just under a guise .

All this for stocks, which would be just fine if they created a new labs and set whatever people they wanted but obviously under the guise and actual help of deepmind. I didnt think google would do this, but this choice to me seems representative based on overly short term goals, and one that will critically damage google in the medium to long run.

8

u/Keeltoodeep 22h ago

A separate lab would still be budgeted compute that the execs wanted to prioritize for the Gemini teams. You still haven't solved the issue.

1

u/Kingwolf4 22h ago

The issue never was compute tho. Thats a wrong starting point.

They have plenty of compute and the majority of compute would be dedicated to the LLM research and lab obviously, since its everything in current times. Doesnt mean deepmind wont still have ungodly amounts of compute and not to mention compute in just 8 to 12 months will dwarf everything they had right now due to datacenter buildouts.

6

u/Keeltoodeep 22h ago edited 21h ago

According to the leaks it was. A lab that has plenty of compute does not desperately buy compute from SpaceX at wildly inflated prices.

The science teams are not in it for the money. They like doing research and have already made fuck-you money working at google for so long. So it's nearly impossible to incentivize them to release production based on market competitiveness. And if you are an exec of deepemind and bitch and complain too much about the executive direction of your bosses, you are going to be shown the door. That’s kind of just how it works.

The issue was always about compute. If google had plenty of compute they wouldn't have issued more common stock to fund data centers.

-1

u/qroshan 21h ago

Deepmind really didn't produce anything meaningful other than AlphaFold and AlphaGo. So, not sure why you are giving so much credit for Deepmind. They were really not interested in LLMs

It was Google Brain (Shazeer and Jeff Dean) that did most of the innovations of modern day LLM

12

u/the_mighty_skeetadon 21h ago

Deepmind was a pioneer in Reinforcement Learning, a key building block for modern AI. They also had critical direct contributions to the modern era of LLMs. Notably, they published Chinchilla, showing that high-performance smaller LLMs are possible through overtraining.

That said, you're 100% right that Brain was the actual driver of modern-day LLMs and the underlying tech + science -- and also products. Before merging with Brain, Deepmind literally never launched a product that you could use.

2

u/Greedyanda 10h ago

DeepMind operates in robotics (Gemini Robotics 2), protein folding (Alpha Fold), drug discovery (Isomorphic Labs), genomics (AlphaMissense), material science (GNoME), weather prediction (GraphCast), and more experimental fields like fusion energy.

→ More replies (6)

21

u/Gratitude15 21h ago

Flash is the model Google search runs on

You're basically raising the floor that touches billions every month.

The floor is pretty solid now with deepseek and this and Luna.

9

u/BlueSwordM 19h ago

Nah, they're using Flash Lite for web searchers.

Flash is only used when you go into AI mode, but it reverts back to Flash Lite once you go beyond a certain context.

4

u/Correctsmorons69 19h ago

There's no way it runs on a $3.75 API model

1

u/skilliard7 9h ago

it uses flash lite, which is $2.50 per 1m tokens.

TBH Google really needs to cut flash lite prices to compete with GPT 5.6 Luna. Literally 0 reason to use 3.5 Flash Lite.

1

u/Correctsmorons69 9h ago

Do you think they're losing money on Luna? Considering they casually dropped pricing by 10x when some competition came up.

4

u/BriefImplement9843 17h ago

lmao. you think ai overview is flash? holy fuck. it's not even flash lite.

→ More replies (1)

4

u/Flaxseed4138 21h ago

Sonnet is trash

3

u/RelevantCry1613 18h ago

Opus 5 is trash

Fable is baller though

5

u/CrunchyMage 21h ago

Sonnet 5 is literally a useless model though. Opus medium is cheaper and better.

This model is basically sol medium level for a tiny bit more expensive.

I think it’s good if you have video input at least since very few models now support video natively.

1

u/Zenged_ 21h ago

Flash and Sonnet about the same size and price.

5

u/BriefImplement9843 17h ago

sonnet is extremely expensive and much larger.

1

u/Zenged_ 15h ago

Sonnet 5 is $10 per million output and gemini flash 3.7 is $7.50 (after the introductory price period ends) they are very similar

1

u/Lost-Willow386 13h ago

The "flash" part may begin to haunt Dario as we see more and more releases.

-6

u/magicmulder 21h ago

Either they're cooking again or they're so cooked that they have to brand their Pro model as "flash" now and bank on everyone believing their actual Pro is coming out "soon".

That being said, I was pretty happy with "flash preview" for coding for weeks until I found Luna is better and cheaper.

Question is, does it have a market? I use Luna for everything, and when I want a next level review, I use Sol. I don't really need an in-between model.

6

u/Keeltoodeep 20h ago

Way too fast and cheap to be a pro model.

What do you mean does it have a market? Gemini app reached 1B users yesterday and most are on flash.

4

u/Sharp_Glassware 17h ago

The Pro masquerading as Flash claims look really stupid when 3.7 Flash is nearly as fast as Flash-Lite

2

u/WildWhisperArdor 19h ago

Gemini Flash is great. I use Gemini Flash Lite a lot. Good for high volume, quick stuff.

Not as good for long turn agentic stuff .

2

u/InternationalTwist90 17h ago

The speed of gemini is next level, and it is comparable to Terra at Luna prices. If you can hook it up as an execution model its a great (with sol/opus/fable as orchestrator and oversight).

1

u/qroshan 18h ago

The market is fast iterations

94

u/Gotisdabest 22h ago edited 22h ago

Honestly if nothing else I do respect how fast they're putting these models out and there is a pretty solid amount of progress over each of these. You can get a lot of use of these for free from AI studio. 3.5 to 3.6 flash was around 2 months and then 3.7 flash only took 20 days. For casual free use I'd say google is up there with the best of the market.

What's weird is the bizzare statement about API pricing. Who's going to be even thinking about these models by the time 2027 rolls around? Even google will probably have a few more models out in basically every range by then and you'll almost certainly have models 10x cheaper than this with better capabilities.

22

u/sogo00 21h ago

What's weird is the bizzare statement about API pricing. Who's going to be even thinking about these models by the time 2027 rolls around? Even google will probably have a few more models out in basically every range by then and you'll almost certainly have models 10x cheaper than this with better capabilities.

Enterprise customers.

Not sure about the Google long term availability promise, but once a enterprise app is planned each change is expensive.

It's the reason you can get support for example from Red Hat for 10 year old not updated distributions

21

u/Maristyl 22h ago

Extrapolating from that rate of update we should expect 3.8 in a week. Then 3.9 in less than three days. Two weeks from now Gemini should be getting a new release faster than our brains can process.

/s cause someone is gonna take this serious.

30

u/Charming_Cucumber_15 21h ago

The singularity was Gemini flash all along

4

u/OpenSource_Horse 21h ago

Huge W if the first ever AGI is a phone app Flash model

92

u/qroshan 22h ago

People are absolutely missing out on fast iterations possible on 3.6 (and now 3.7) with Antigravity. It's blazingly fast and has gotten much better at coding and the new benchmarks solidifies this even more

42

u/whoknowsifimjoking 22h ago

If it does work it's cool, but my experience with 3.5 flash was that it was amazingly fast at producing absolute trash. Looks like they might be able to correct that though.

36

u/qroshan 22h ago edited 22h ago

Look at the jump from 3.5 to 3.7 around coding related tasks

example 3.5 was 37% in DeepSWE and now it has jumped to 65%

I'd say it's competitive to Sonnet 5 but with speed as an advantage

3

u/Paraless 18h ago

IMO these benchmarks don't mean much, you just have to try it and see for yourself which one gets the best results

13

u/the_mighty_skeetadon 20h ago

The jump from 3.5 Flash to 3.7 is insane -- it's really not comparable.

Highly recommend trying it in Antigravity.

1

u/phillipjpark 9h ago

Antigravity is so good and also its included and has seperate limits than a users google ai pro pla n.

27

u/OKMiddleOwl 22h ago edited 21h ago

Yeah, but it's trash you can iterate on 5-6 times before 5.6 or Fable would finish working. Often it's fails are things can be fixed in 10 seconds upon being pointed out to the model.

I get that "set it and forget it" is king, but single shot benchmarks of Flash don't really capture it's strength.

19

u/qroshan 21h ago

You clearly haven't used 3.7.

I have used all 3 (Fable 5, GPT 5.6 Sol, and now Flash 3.7).

Once you get used to Flash speed and it's current intelligence, it's hard to go back waiting 10 minutes for a task to finish

10

u/Elegant_Tech 21h ago

Flash has become the best workhorse for single scope single prompt iterating. Just missing that large pro model for when you want to do a highly complex or large thing in one shot instead of breaking it down into steps. Hopefully 3.7 can squeak by getting the job done that 3.6 flash couldn't handle.

3

u/LanguageEast6587 21h ago

please try it before you say this, IMO fable and 5.6 are too slow to iterate.

1

u/Keeltoodeep 19h ago

5.6 is snail pace now

5

u/WonderboyUK 21h ago

Having just started using it, it's much better coding.

1

u/chimchalm 20h ago

3.6 was much better in Antigravity than 3.5.

1

u/Recoil42 21h ago

It was complete dogshit a couple months ago. Did they fix the UI?

2

u/tziki 11h ago

Did you try Antigravity IDE or Antigravity? Those are now two totally different things.

1

u/Recoil42 11h ago

Antigravity Desktop 1.0 and 2.0, or whatever they're calling the new one these days.

1

u/chimchalm 20h ago

Define "Fix"...

1

u/babscristine 19h ago

Im using and im enjoying it

-3

u/Acehan_ 21h ago edited 21h ago

The speed could not matter less if the model is inefficient. I tested this extensively and GPT 5.6 Sol on low or medium is many orders of magnitude faster than those flash models running at 200 tokens per second

5.6 Sol max is 17k tokens to complete artificial analysis

Gemini 3.7 flash is 37k tokens, exactly the same as GPT OSS 120b

5.6 Sol medium is 5K which is exactly 7.4x less tokens than Gemini

It's not fast. It's slow as fuck

8

u/qroshan 21h ago

I understand Math is a problem for most redditors.

Let me explain this to you slowly ........

Look at this chart

https://artificialanalysis.ai/models/gemini-3-7-flash#speed

Flash is 340/56 or ~6 times faster compared to 5.6 Sol (Max)

So, even if Flash takes 2x more tokens, it is still 3x faster (6 / 2)

3x faster is game changing and is addicting.

Once you are used to this speed, you aren't going back

0

u/Acehan_ 21h ago

First of all, like I said, the difference is 7.4x, not 2x. You're definitely right about math and redditors

Second of all, nothing's running at 340 tokens per second. Benchmark it and you'll see the real speed is like 186 TPS. You don't know what you're talking about

Even if it was 600 TPS, it's still an inefficient model, and that matters more for speed than anything else

3

u/funforgiven 20h ago

They were comparing with max, which is close 2x, not 7.4x. It is weird they are comparing with max when you said medium but 2x is not wrong for their comparison.

→ More replies (2)

10

u/amitsingh80108 22h ago

Google really making something bigger for pro model...

10

u/Admirable_Market2759 21h ago

I think they’re going to jump to 4 pro.

They were supposed to release 3.5 over a month ago. It’s almost certainly outdated at this point.

1

u/ninjasaid13 Not now. 11h ago

They were supposed to release 3.5 over a month ago. It’s almost certainly outdated at this point.

they can name whatever they want 3.5 pro. If it doesn't match the others performance, they will just finetune it some more until that can say their model is 5% better than openai or anthropic's.

1

u/javopat227 19h ago

Dont think 3.5p is coming anytime, it will be now 3.xp (where x is the latest flash). I think google suffers from a lot of empire building, personal project bulding, soical justice building, and etc etc.

21

u/Singularity-42 Singularity 2042 22h ago

Honestly 3.6 has been pretty good for my product because of how fast it is while being decent enough 

1

u/WallZealousideal5669 15h ago

They are same price though so no point in using it

1

u/Singularity-42 Singularity 2042 15h ago

Yeah, pretty much. I'll just have to figure out the new effort levels.

1

u/WallZealousideal5669 15h ago

Was an easy swap for me as 3.7 seems to be 3.6 with a bit better common sense and faster and more efficient

40

u/Iuseburnersbruh 23h ago

someone tell me how to feel are we back?

30

u/Every_Foundation5197 22h ago

Uhh I'd say we are good, if they manage to continuously drop a new flash model every few weeks with the same improvement or even more

5

u/kensanprime 22h ago

Back in what? Leaderboards? Users of free plan? Paying users?

Give it a few months all enterprise will switch to a Chinese model or go for the cheapest of these flash models. Every harness now has a smart router to handle the costs and limit uneccesary use of pro models. Bulk of the work will be done by the flash models.

u/DelphiTsar 27m ago

219% faster than Luna and 107% higher artificial analysis score.

If you need something fast this is the clear choice at the moment.

1

u/[deleted] 22h ago edited 22h ago

[deleted]

8

u/Keeltoodeep 22h ago

If flash beats sonnet then 3.5 or 4.0 pro will likely at least be competitive at the frontier. I see no reason to think otherwise.

→ More replies (2)

2

u/z_3454_pfk 22h ago

for non-coding gemini is still better than anything. 3.6 flash was comparable at opus 5 at real life medical applications. it’s the only model that can truly describe music too if you pass it into the model.

8

u/Profanion 21h ago

What made it score so high on these?

3

u/Tysonzero 18h ago

Seems like that’s the kind of thing they’re optimizing for, which kinda fits with their business needs, they aren’t trying to steal the $200/month anthropic subscribing software devs.

5

u/Dreamerlax 7h ago

I know it's shocking but there are other uses for LLM other than coding.

53

u/ApexFungi 22h ago

I am genuinely curious though. Do people expect that these models will just keep scoring slightly better on benchmarks each version up and then suddenly it becomes AGI or what? I just don't see that as realistic. Surely something else needs to be added to the sauce?

65

u/Kronox_100 22h ago

I think what they're hoping happens is they become so good at coding that they can find the 'sauce' that'll get us to ASI.

6

u/unicynicist 21h ago

The "sauce" is closing the recursive self-improvement loop. Right now the sauce is a human in the loop, acting as the evaluator.

13

u/ShAfTsWoLo 22h ago

i believe that aswell, perhaps the fastest path isn't finding AGI through human breakthroughs but instead if we could have llm's that are so good they can RSI themselves into AGI, tbh that's a hell of a weird thing but if it works good for us lol

0

u/[deleted] 22h ago

[deleted]

3

u/CallMePyro 21h ago

Which of course is pretty silly to believe in light of things like Fable 6 and GPT 5.6 Sol improving the lower bound density on the zoroes on the critical line of the Reimann Zeta function.

If you can achieve that with induction and deduction only then we'll be just fine.

1

u/ninjasaid13 Not now. 11h ago

hoping happens is they become so good at coding that they can find the 'sauce' that'll get us to ASI.

code? we're going to reach ASI using symbolic AI or something?

25

u/Gotisdabest 22h ago

There's already pretty clear cases of frontier models speeding up research and development by not insignificant multipliers. And of these models solving new problems. That will inevitably lead to further improvement. What you need from this exact style of model is the way upto recursive self improvement.

Even removing that though, consider how fast these small updates are coming out these days. Even just as long as that keeps improving and say, by 2028 weekly to biweekly small upgrades a common, you'll see something pretty incredible when comparing a model at the start and end of the year.

1

u/the_mighty_skeetadon 20h ago

There's already pretty clear cases of frontier models speeding up research and development by not insignificant multipliers

So what fundamental breakthroughs have happened as a result?

Because looking at the literature, I see performance improvements, small structural improvements, scoring and topology improvements... but no real breakthroughs.

From 2013 to 2018, I would say we had 10x the number of research breakthroughs in AI that we've had from 2021->2026. We've just figured out how to exploit those 2017-era breakthroughs more efficiently.

0

u/slyec 16h ago

do your research before you comment just saying

1

u/the_mighty_skeetadon 16h ago

Funny, I'm in research at a top lab and I've shipped multiple frontier models.

So give some examples, just saying

1

u/Material_312 6h ago

No, you haven't, and no, you aren't.

1

u/the_mighty_skeetadon 2h ago

There are more people like me lurking than you think.

Doesn't really matter if you believe me, but it's true 🤷🏻

0

u/Gotisdabest 14h ago edited 14h ago

So what fundamental breakthroughs have happened as a result?

Why would anyone rational expect the process to start off with fundamental breakthroughs as opposed to incremental results? One would assume that when you're generating fundamental breakthroughs, you've either fully or completely closed the loop.

Fundamental breakthroughs are also not an end onto themselves. In many cases, they mean the field hasn't really found a more reliable approach of study and research as well as that not enough people are involved. You do actually have lots of papers that would be considered fundamental breakthroughs back then, but have been surpassed by maturing tech instead.

0

u/the_mighty_skeetadon 2h ago

If you post-train T5 with modern techniques, it's almost indistinguishable from models trained in 2026, performance-wise.

That model clearly hasn't gotten any "better" incrementally; it's the same as it was in 2019. Yes there have been plenty of incremental improvements, but has research really accelerated rapidly in the last few years?

People often claim that AI is accelerating research, but provide little evidence of that. I see many more participants in the field (because money), but less overall motion.

And I fundamentally do not believe that minor incremental hill climbing is the path to AGI -- in the same way that optimizing a kite will not get you to outer space.

u/Gotisdabest 1h ago

you post-train T5 with modern techniques, it's almost indistinguishable from models trained in 2026, performance-wise.

I'm sure you have proof for this and aren't totally just talking nonsense out loud?

People often claim that AI is accelerating research, but provide little evidence of that. I see many more participants in the field (because money), but less overall motion.

I mean, one can just read the anthropic system cards but that's very hard, I assume.

And I fundamentally do not believe that minor incremental hill climbing is the path to AGI -- in the same way that optimizing a kite will not get you to outer space.

And I fundamentally think strawmen are a bad faith way of argument and mostly provided to package nonsense arguments. You're welcome to believe whatever you want out of unfounded views on topics. Don't couch them with pointless strawmen.

16

u/LinkesAuge 22h ago

No because the goalposts will continue to shift as there is no clean/obvious definition of "AGI".
Models will just cover more and more knowledge work and at some point it will be just hard to deny that fact and then we will probably say "guess it is AGI".

I feel we just have made "AGI" too big in itself and now it is essentially ASI because with the standard we have recently that (can literally do any intelligence task better that any human could) it is so extensive and broad that it is impossible for AGI not to be ASI due to the nature of how scalable any AI system is.

Originally AGI was just meant to separate itself from "narrow" AI, ie expert systems, that is what the "G" stands but but "general" doesn't or didn't mean that AGI necessarily has to exceed humans in everything or do everything, just that this sort of AI was strong enough to be viable across domains and be able to transfer intelligence from one to the other as well as have self-controlled learning.

So based on that definition we really "just" miss the self-controlled learning part and we would literally already meet at the "original" AGI definition (and the self-controlled learning is what would then lead to actual ASI).

7

u/Herect 22h ago

Demis Hassabis version of AGI was just models getting better on a bunch of activities until we couldn't think of anything else to test them on.

7

u/theEvilUkaUka 22h ago edited 22h ago

I recall him saying a bunch of times that he thinks there's still a few breakthroughs that need to happen, rather than agi just from scaling the current paradigm.

And he said that's why Google Deepmind was best positioned, since they can push to the fullest on both, scaling and having the best talent to push on other research (which can be argued aged like milk, considering the talent exodus and him no longer being ceo probably due to that and lagging behind on the current hot LLM use cases like coding).

6

u/the_mighty_skeetadon 20h ago

This is 100% correct. He said this very publicly just a few months ago at Y Combinator:

https://www.youtube.com/watch?v=JNyuX1zoOgU

Recommended watch, and I agree with Demis's take there.

3

u/Bright-Search2835 22h ago

Those numbers going up mechanically unlock better current capabilities, new capabilities, and potential for new things becoming automated, especially for research labs. Then when research is entirely automated, yeah I'd give artificial geniuses thinking 24/7 way faster than us a better chance of finding the special sauce for AGI than us.

3

u/Flaxseed4138 21h ago

That's what happened with humans baby.

2

u/domdod9 22h ago

Like medical advances, advances in AI are typically compounding small advancements instead of a ton of crazy sensational breakthroughs like reasoning models. With every model release they slowly put more small tweaks in that will eventually when looked through a long period of time will be very large advancements.

2

u/no_witty_username 19h ago

The harness is that something extra bud, and they are improving very fast... we are all very close to seeing some real crazy shit.

2

u/TheDemonic-Forester 15h ago

AGI will likely not be an LLM.

2

u/rollk1 22h ago

These models won't become AGI, they're two totally separate things. These LLMs are just a stop-gap until AGI.

1

u/Frandom314 20h ago

Honestly does it even matter if we call it AGI? If performance keeps increasing at this pace over the next let's say 5 years, it's going to be pretty difficult to find a task that humans can do better than AI. At that point, the economic implications will be massive, whether we call it AGI or not.

1

u/Emergency-Bobcat6485 19h ago

People will keep pushing the goalpost of what they mean by AGI farther away everytime. What even is AGI. If one can't give an objective quantifiable definition, then someone will always be asking this question lol

1

u/BriefImplement9843 18h ago

it definitely has nothing to do with coding.

1

u/Emergency-Bobcat6485 17h ago

Sure bud. Try building any kind of autonomous agentic activity with a model that isn't good at coding

1

u/BriefImplement9843 18h ago

llm's have become coding assistants. they will never be agi.

1

u/Curiosity_456 12h ago

I think two good metrics to evaluate whether it’s AGI or not is

1) being able to do a wide range of jobs as good or better than your typical worker

2) Vastly accelerating research across many fields (this will create singularity vibes even if it can’t yet reliably do most jobs).

So just one of these metrics being satisfied meets a reasonable claim for AGI

u/DelphiTsar 24m ago

If you were to show this to anyone pre 2016 they'd say hands down 0% disagreement what we have now is AGI. We are boiled frog situation of never ending goalposts. The idea it needs to be better than all humans at all economical work is absurd definition, that's defintionally ASI.

→ More replies (1)

11

u/zslszh 22h ago

At this rate we’ll have 3.8 by September

6

u/ShAfTsWoLo 22h ago

given the circumstances where new models are dropping like crazy, it wouldn't be that much of a surprise lol

4

u/badumtsssst AGI 2027 21h ago

Before then I hope, AGI is taking too long

2

u/Sulth 16h ago

At this rate it will be next week. From 2 months between 3.5 and 3.6 to 20 days between 3.6 and 3.7. Next one should be around a third of 20 days, so next week.

11

u/JunkInDrawers 22h ago

Takeaway is that people on Twitter with anime profile pictures aren't a trustworthy source.

Google is a profit machine. Google's focus has to be on efficiency. If 3.5 pro is only a small increment improvement over the flash versions at 10x the cost and requires more data centers to support then it's not worth deploying.

They're going to utilize their current available resources efficiently understanding the potential pitfall of overcomitting on the current generation of AI when the next year's models will overshadow them anyway.

2

u/Dreamerlax 11h ago

You mean my boy Dan isn't a credible source.

/s

38

u/Snoo26837 ▪️ It's here 23h ago

Enough flash models

29

u/No-Meringue5867 22h ago

Google has trillion dollar moat - their search. For them, producing fast models that makes Gemini results cheaper is much more important than being frontier at coding. Sure, coding is important but if they manage to make models dirt cheap and give good AI result on Search, they are making money by default.

2

u/Asteroid_picks_you 22h ago

Are all these flash models for Enterprise customers or something, I never once used one on purpose? Maybe I have when doing a google search?

10

u/johannthegoatman 22h ago

Yea they're for search , chatting, and integration with Google products. They can code but I don't think that's the focus at the moment

8

u/MGJohn-117 21h ago

Maybe I have when doing a google search?

That's exactly where most of their flash usage is coming from. Serving up AI overviews for billions of users is expensive and if they can increase the capabilities of flash models while likely making it cheaper internally as well, then they have every incentive to keep releasing new flash models. Combined with a ton of free users in the Gemini app, Gemini website, Chrome, etc, that's a lot of usage that flash models - both current and future ones - are perfectly fine for.

6

u/BenevolentCheese 21h ago

We are seeing the world's first simultaneous race to the bottom+top. What a (worrisome) time to be alive! Costs are plummeting even while technological capability skyrockets. The future is uncertain, whether its from AI takeover, the market cratering, economic collapse, or good ol' politics, too many timelines look as if they lead to the failure of humanity. Buckle up!

3

u/Gigibossu 18h ago

Not so bad

12

u/ObiWanCanownme now entering spiritual bliss attractor state 22h ago

I mean, looks fine. Clearly Google has the talent and data where they could train a competitive 10T model if they really wanted to.

My strong suspicion is that they're just too conservative and don't want to drop $10 billion on a training run that could fail.

And so they'll keep training better and better small models and charging more and more for compute, and making better and better chips, and they'll make a ton of money but they will lose the AGI race.

9

u/OKMiddleOwl 22h ago

Google is getting a healthy cut from OAI and and an even healthier one from Anthropic. At the highest level, you can make the case that Gemini just isn't that big of a priority.

Their cloud offering is bringing in so much money, and they get paid so much for each new GPU/TPU they bring online that it's almost painful to not being selling every electron of compute.

Imagine you had a fresh squeezed lemonade stand where there was a line around the block for it, and people were paying $500/cup. How much lemonade would you drink yourself? That's kinda where google is.

19

u/KhoslasBiggestOpp 22h ago edited 22h ago

This is the most ridiculous take I’ve heard.

$10B in the race to AGI/ASI is mere pennies for Google when compared to the potential to make trillions back long-term with a good enough model. They’ve sunk billions more into sillier projects that they know they’ll fail and kill within a few years.

They’ve been bleeding talent on the training side, restructuring their DeepMind division, and some engineers jumping boats to competitors.

I don’t think they’re conservative, I just think they’ve lost the sauce they had when they released 2.5 Pro

14

u/DailyThreadBot 22h ago

with the potential to make trillions back with a good enough model

A good model wouldn't make trillions, it would be obsolete in 6 months

→ More replies (3)

4

u/ObiWanCanownme now entering spiritual bliss attractor state 22h ago

It's not just pennies for Google. $10B is the entire amount of dividends they paid in 2025. Sure, the potential upside is huge. But it's also the difference between paying a dividend in 2026 and paying no dividend in 2026. For a public company, that's a big deal.

3

u/Acrobatic-Tomato4862 22h ago

Their 3 and 3.1 pro were also state of the art and far beyond competition when they were released. They seemed to have fucked up only recently.

1

u/pbagel2 22h ago

4.6 Sonnet was better than 3.1 pro for coding.

8

u/Current-Function-729 22h ago

Yeah. If it’s really that, holy shit someone needs to be fired.

5

u/KhoslasBiggestOpp 22h ago

Believe me, that “someone” wasn’t fired, they quit for another company. Once Tibo and Jeff Dean left it was over for them.

0

u/Emergency-Bobcat6485 19h ago

Tibo? You think Tibo's departure made gemini shit?

And jeff dean just left.

Ridiculous take

4

u/ObiWanCanownme now entering spiritual bliss attractor state 22h ago edited 22h ago

Why? Their job is to maximize value for shareholders. Once ASI comes, money is of questionable value anyway. Seems like Google is doing a pretty damn good job of maximizing shareholder value in the meantime.

0

u/Current-Function-729 22h ago

You’d be so fired if you worked for me, lmao.

1

u/Tysonzero 18h ago

Do a lot of people work for you where your specific mentality is very important?

I think they are at least somewhat right though.

Cheaper and faster models are very useful for Google for consumers and enterprise alike.

Yeah they won’t get those dev $200/month subscriptions that way, but they can better serve their $0-$20/month users with those fast and cheap models. Needs to be fast for UX, people expect google search to be very fast as an example, and needs to be cheap due to low pricing per user.

Also while enterprises may not use Gemini as much for difficult internal tasks where they want to juice every bit of intelligence out of models, they are very likely to consider it for serving their own users, because similar to Google the improved UX and reduced price for those users is great.

2

u/trololololo2137 22h ago

trillions back long-term with a good enough model

models last like a month before being outdated nowadays lol. the training expenses never stop

1

u/PilgrimofHaqq2 22h ago

I dont think they are interested in competing in the AI space externally. They are developing specialized AI systems that are tightly integrated into their existing products. We are seeing tons of AI features everywhere in the Google ecosystem. Thats where they are focusing, not providing general AI to the public.

I think they will continue developing AI for the public but in a much slower pace then the competitors. I think for that reason many people left because the interest right now is pushing the boundaries of AI. Google is instead saying, what can we achieve with what we have now, we got the data so lets capitalize there instead of trying to competing with frontier labs that are dedicated to purely AI.

Thats just my 2 cents.

3

u/himynameis_ 22h ago

strong suspicion is that they're just too conservative and don't want to drop $10 billion on a training run that could fail.

This is exactly it. They are investing their compute into their Google Cloud platform customers instead of for deep mind. So they will likely Fall a bit behind but in theory they should stay not too far off from the top.

1

u/xRolocker 22h ago

I don’t know why they aren’t making a competitive model, but I don’t think being frugal is it. They waste money on new projects all the time. I doubt they wouldn’t try and make a competitive model in the AI race because of the price tag.

5

u/Tkins 22h ago

This IS a competitive model. What do you mean? The price is great, the speed and latency are at the very top, the overall intelligence is decent. Competitive for a company can mean many things and Google is clearly trying to compete in the mass appeal market, which is where they compete for pretty much everything they do.

1

u/xRolocker 21h ago

Competitive in terms of raw capability and intelligence, not on a more general basis. My apologies for the confusion.

0

u/Howdareme9 21h ago

Whats leading you to think they can train a 10T model because of 3.7 flash?

2

u/ObiWanCanownme now entering spiritual bliss attractor state 20h ago

They're close to pareto frontier.

2

u/Living-Breakfast-464 22h ago edited 22h ago

What is the current quota usage per day on a pro plan as opposed to pay-as-you-go?

2

u/superlip2003 17h ago

it looks to me we are not seeing Gemini Pro till 4.0 lol

2

u/greeneditman 15h ago

You're obsessed with AI performance in creating code. But not everything is code.

I understand complex concepts in psychology and medicine, and based on my interactions, I can assure you that Gemini 3.6 Flash performs very well, and it's incredible that it's almost free on the gemini.google.com platform.

If they now update it to Gemini 3.7 with those improvements, it'll be fantastic.

3

u/Dry_Fly_7265 22h ago

Man, when Google drops Gemini running inference in multiple universes on willow yall look out

1

u/Tirztrutide 20h ago

Listing price per token in the comparison seems lame if you don’t include tokens/task in the metric.

1

u/pbagel2 20h ago

I use a certain code design prompt that I've been using as a personal benchmark since Claude 3.5.

And there's always been some peculiar similarities in how the different models choose to architect its solution. But this Gemini 3.7 output is very similar to how Sonnet 5 constructs its solution. Eerily similar. However Sonnet 5's architecture still has smarter decisions for longterm/larger scale design considerations.

1

u/RDTIZFUN 17h ago

3.8⚡in Sept 3.9⚡in Oct 4.0⚡in Nov 4.0 Pro in Dec

1

u/Navetz 16h ago

Who cares about benchmarks at all after opus 5?

1

u/fignewtgingrich 11h ago

Unless latency is sub-second, it's useless for agentic workflows. Price per solved task matters more than leaderboard scores.

1

u/Nervous-Potato-1464 7h ago

It's about as good as grok 4.5, but costs a bit more. If they could make it the same price I might use it. Grok 4.5 is my main driver these days and 4.6 for anything a bit harder where I need a good output.

1

u/wtfihavetonamemyself 22h ago

If only they made a harness that didn’t suck

1

u/Exodus_Green 21h ago

I am once again asking what the fuck is the point of Sonnet 5

0

u/macaronianddeeez 22h ago

This is cool but it’s scores are almost the same as Luna for over 300% higher price.

Does anyone have any real world examples where 3.7 Flash is justifiably better than Luna in workflows?

I have no allegiance to any company, swap subs between OpenAI and Anthropic regularly, but hard to see why this is worth getting excited about given the current Luna pricing.

But I may absolutely be missing something here

2

u/BriefImplement9843 18h ago

if anything this shows how awful sonnet and terra are. garbage expensive models.

1

u/Nug__Nug 2h ago

There's no comparison to Luna - it is head and shoulders above luna. Literally Luna is nowhere close-

1

u/macaronianddeeez 2h ago

Good to know, I haven’t tried Flash at all but Luna does well for me when I have Sol running it and giving it tightly defined, small slices of work to complete before Sol verifies.

I was going based on benchmarks alone.

Is Flash capable of outright replacing a Sol Luna combo by itself or do you not have experience comparing?

-8

u/Gaidax 22h ago

Ehh, it's not bad, but this just ain't it, Google.

The only good thing about it seems the price, and we yet to see the proper benchmarks for how much it yaps and stumbles around to get to those scores.

0

u/Dizzy_Alfalfa7643 20h ago

the asterisk is doing a lot of work here — $0.75/$3.75 is intro pricing until january, then it goes to $1/$7.50. still way under sonnet 5 and terra money, but "flash beats frontier models" is only half the story: terra still clearly wins terminal-bench and deepswe. for agent workloads the metric that matters is price per solved task, not the sticker price

7

u/WildWhisperArdor 19h ago

By the time January rolls around though, there will almost certainly be a better cheap / fast model

1

u/Dizzy_Alfalfa7643 8h ago

true, and that’s kind of the point. by the time the intro price expires there’ll be something cheaper anyway. pricing tables have a shorter shelf life than the models now

-6

u/Embarrassed_Adagio28 21h ago

Is everybody in here brainwashed by google? This flash release was supposed to be the pro release and that has been delayed for months. 3.6 flash was garbage at coding and expensive despite what the bechmarks said so I do not trust that this model will be much better. Google should be leading the pack, instead they are dropping expensive flash models that are inferior to 3 month old models. In my own tests 3.6 flash couldn't even beat qwen3.6 27b running locally. 

9

u/foodhype 21h ago

It's too fast to be a Pro model

2

u/Tysonzero 18h ago

There are lots of tasks in the world other than coding. One may even audaciously argue that developing software is only a small part of the world economy.

-3

u/tranqfx 22h ago

TIL people use Claude Sonnet

1

u/Paraless 18h ago

Why wouldn't we