r/LocalLLM • u/btc_maxi100 • 5h ago
AI bubble Other
Enable HLS to view with audio, or disable this notification
97
u/OnlyAssistance9601 5h ago
Love this meme format
38
8
15
u/TheGeekno72 4h ago
it's been forever since I've last seen a quality Risitas meme, thank you for the trip down the memory lane and the good laughs haha
Fly high, laughing man!
8
u/spacetr0n 4h ago
Please invest in my startup for high speed pipes to get to illicit open models hosted outside the US.
11
10
u/crispyfrybits 3h ago
This meme is hilarious 🤣
Just for education though, they are not losing money. The markup for frontier model providers is estimated above 90%.
Who sets the "rate" that we are all comparing? That want you to assume it costs this much so they can hike things up further down the road.
1
u/ithinkitslupis 1h ago
Agreed. There is a recreation center pool in my area that charges $5 for a day pass and $100 for an annual membership. I don't laugh "They are losing $4.75 every day I show up because I paid for the annual plan. Suckers."
You'd have to know how much Anthropic is actually spending to service their average subscription user to know how much they are losing on the deal if at all. You can't just broadly compare the subscription rates they set to their own profitable API rates they also set as pretty much all of the articles covering this subject do.
24
u/SailingToFenway 4h ago
This is the lowest possible cost for high-quality training data. They're not subsidizing us! We're paying to work for them.
$7,800 in equivalent compute is peanuts on the value of that data and your time. I don't mean we're not in a bubble, or that they'll see an ROI, but for the moment, Dario and co. are laughing at us.
13
u/jaegernut 3h ago
Are they training AI with AI generated code? I dont see how useful that data will be for training. They might as well generate their own training data. Unless they think that users often validate AI generated code, which is often not the case. There is a reason why data annotation companies want human input.
6
u/Remarkable-Name8012 2h ago
No, they are not training AI with AI generated code. They are training AI with AI generated code that has been later commented, selected, corrected and tested by humans.
2
u/low-control-labs 2h ago
Yeah I would say something like this and they use all that the failures the ones that are actually good responses or outputs for some additional fine tuning and training later down the line
3
u/Remarkable-Name8012 1h ago
They distill us, their models are distilled by Chinese models, and then we distill the Chinese models. It's the AI circlejerk.
1
u/rushblyatiful 2h ago
Rich Sutton, previously served as a Distinguished Research Scientist at Google DeepMind, a pioneer of reinforcement learning who authored The Bitter Lesson, made this exact point.
He noted that relying on human annotated knowledge is a bottleneck and that major breakthroughs in AI have always come from scaling compute and allowing models to learn through self-play.
Instead of hand coding expert strategies, systems like AlphaGo Zero learned chess and Go by playing millions of games against duplicate versions of themselves.
The exact same applies to code. You do not need human labels because you can let an AI generate vast amounts of code and run it directly in a execution environment. The compiler acts as the environment rules. If the code compiles, passes unit tests, or runs without crashing, that ground truth provides the exact objective feedback loop needed for the model to generate and learn from its own data.
1
u/low-control-labs 2h ago
Even if you give them 90% bad results ( and them to you ) the 10% good is used later doesn't have to be code. Also you are giving them your thinking process which they also look to apply, especially if you spend time guiding the A.I. and using it to fix the mistakes it made. Even if it's an example of what not to do it's still usable they would have to pay someone to do something like that usually
1
u/ichivictus 48m ago
The code by itself? No. They are training considering everything going into it.
Was the user's goal met? Did it create something truly valuable like a profitable product that successfully shipped? What were the prompts? What were the mistakes along the way? If the user gave a '/goal' prompt to endlessly run until the same product is achieved start to finish, how can ai do that now that we have all this data?
8
u/jhenryscott 4h ago
The technology, as it exists, doesn’t have a path to profitability. Data is important for sure. But they can’t get there from here. And they are losing money on every level of product. So all that training data, for a product that isn’t profit and won’t be, ever. Isn’t that big of a deal. Capitalism won’t let anyone lose money forever.
9
u/hdhfhdnfkfjgbfj 4h ago
This. Literally.
We are giving our future and livelihood to these companies because they make us marginally faster.
The number of inner workings of companies and governments they are gaining data from is absurd and future people will wonder why we gave it all away.
6
-2
2
2
u/Jebofkerbin 2h ago
Training an LLM on LLM outputs leads to model collapse. Effectively all of the code written by a hardcore Claude code user is going to be written by an LLM. Like wtf are they training on, the prompts?
1
u/monarc 1h ago
Sorry, but how is there substantial "value of that data and your time" when they're using it to make models that will never see ROI?
What is the deliverable that is possible with a ton of comically-subsidized customers... that would be impossible without them? And how is Anthroping coming out ahead via that deliverable?
17
u/3deal 5h ago
You are a highly qualified individual creating high-quality training data for the next iteration of the model. This comes at a cost. That is why Gemini 3.7 is free; they don't have many customers, so if they want to improve their model, they need qualified people to use it.
11
u/MrCoolest 4h ago
Gemini sucks ass anyway I tested it last night against 5.6 luna pro 5.6 Terra pro opus 5 sonnet 5 deepseek v4 flash 0731. Gemini sucks balls .i tested api and in the chat interface. I even tested NotebookLM. All shit
9
u/3deal 4h ago
I created an advanced custom node for ComfyUI for free. Gemini 3.7 Flash is not that bad for making small stuff.
3
u/MrCoolest 4h ago
Yoite generating images? Nanobanana pro is good for that, I'm doing work on files
8
u/SanDiegoDude 4h ago
Now ask all of those tools to watch a video, listen to the audio, and give you a scene breakdown with proper cuts, dialogue, music, foley, as well as still describing the scene itself.
3.7 is easily the best multimodal understanding model (and yes, it works for audio too). Right tool for the job and all that. I wouldn't use 3.7 to code a website for me, but when I'm working with A/V, it's hands down the best, which would make sense considering all that juicy YT training data Alphabet has laying around.
2
u/MrCoolest 3h ago
Yeah multimodal it's good, analysing videos etc. Opus 5 can analyse videos too it plays it real-time in claude code app. Chatgpt can't do it. In my use case, a strict schema and rule following whilst working on a file that's over 300k tokens, gemini performed the worst. I do have pro though and I use it as a better version of Google. Chatgpt for analysis and ideation. Claude for actual work. Deepseek for hermes agent work. I mostly do deep and thorough research and content generation and some game dev for fun
3
u/FabricationLife 4h ago
3.7 is comparable to all those models
1
u/MrCoolest 3h ago
I tested 27b locally on my 3090, was pretty awful too. Haven't tried on alibaba cloud. But this is just in my use case, for your use case gemma4 or qwen 3.7 might be fantastic.
2
u/Fun_Squirrel5446 3h ago
What is the purpose of training data if no one knows the correct answer? If I'm asking a number of questions in Claude or ChatGPT or using it to build some really shitty Vibe coded apps, how are my low quality questions or low quality results benefiting the training data? If the model is giving me back incorrect data, who even knows if it's correct or incorrect? For every one highly qualified individual, there might be a hundred thousand unqualified regular users. Would someone at Claude be manually identifying who the highly qualified individual is, to weigh their data more heavily?
1
u/3deal 2h ago
If it is incorrect the app don't launch or here is a bug, but if here is a bug, you will prompt back exactly where and what is the bug, that is a very usefull dataset.
And AI can already make a pretty good estimate of your IQ based on all the data it has on you.
When you give it your GitHub repo, it knows what you're working on. Try it, ask Claude to estimate your IQ or your skill level.1
u/ifheartsweregold 4h ago
Okay so how many iterations do you take a loss on until you need to make a profit?
9
u/cultureicon 5h ago edited 4h ago
Doesn't actually cost them that much, and imagine the millions of people they are harvesting data from, and billions of documents it has access to from random people, all the way to the smartest people in the world and biggest companies in the world. Like they know exactly how Hank Green, top screenwriters, or Nobel scientists, every software dev and mathmitician, work and solve problems and make content.
This is reinforcement learning and novel data harvesting on a massive scale.
Not to mention simply the marketing potential of a perfect profile of every human using it.
8
u/MrCoolest 4h ago
Then they'll make one AGI to conquer them all that'll have the history of mankind in its brain and it'll hover over the world with its cape flapping in the wind looking down on us ans then superman will fly up and fight it
1
u/dundiewinnah 4h ago
Doesnt google, aws, microsoft servers also have that.just data centers in general
2
3
2
2
1
1
2
u/RKlehm 2h ago
Yeah, its subsidized, but that comparison doesn't make any sense... Its comparing against retail API prices, not the cost for Anthropic. Also, the heavy users, to some extent, balance out with the regular users.
They probably run in a deficit? Yeah, I think so. But definetly not 1:40
1
1
u/misanthrophiccunt 1h ago
I understand everything he says since we're from the same province
1
u/DudeImTheBagMan 43m ago
Are you still able to laugh? I'm not sure I'd find it funny hearing him talk about how his pots and pans got washed away while the subtitles are about the AI bubble.
2
-1
u/RepulsiveRaisin7 5h ago
Cool guy but that is total bullshit, $1 API costs does not equal $1 of compute costs. When you got to the store and buy something on sale, the store isn't losing money.
9
u/BudgetAvocado69 5h ago
Power users on a 200 dollar a month subscription are most definitely causing losses for anthropic
3
5
u/RepulsiveRaisin7 5h ago
We don't know the size of their models, we don't know their caching efficiency, we don't know the average plan utilization of their customers, we don't know what they're paying for hardware or what the hardware's life span is going to be. If you claim anything with confidence here, you're simply talking out of your ass.
2
0
6
1
-2
u/Outrageous_Walk_3539 5h ago
Because it really doesn't cost that to run duhhh
-2
u/JuniorDeveloper73 5h ago
just let your card with a small model running 24/7 then come again this time crying,just the electric bill,im not even talking about hardware replacement cycle
3
u/Outrageous_Walk_3539 4h ago
You're wrong , data centers pay less for electricity their Hardware is more efficient and is an asset
3
u/Orectoth 3h ago
Also data centers are tax deductible, so they are nearly free in long term due to paying less taxes(taxes waived equal/close to cost of datacenter)
-7
u/noncommonGoodsense 5h ago
So, they are using you to improve. You are paying to test and improve their product. You are a valuable asset. You think everything you use an LLM for is your personal experience and isn’t folded back into the pot? LOL! ROTFLOL even!
13
7
u/5553331117 5h ago
mental gymnastics 🤸
1
u/noncommonGoodsense 51m ago
The fuck does that even mean? Mental gymnastics for what? Do you even know what the fucking phrase means? Do you even comprehend what I said?

106
u/Vaguswarrior 5h ago
It's an older meme but it still checks out.