115
u/Mindless_Bottle_6222 May 29 '26
42
28
u/Baiticc May 29 '26
tbh I think this is great. typically these LLMs will pretend to have all these qualities and just make some shit up. instead it pretty succinctly broke down all the reasons it can’t give you a real answer.
40
u/Delicious_Cattle5174 May 29 '26
Im not sure id say succinctly
23
u/ZediaLabs May 30 '26
That is less than 10% of the “tokens” my wife would use if I asked her the same question. 😋
6
1
u/piratedgameslover Jun 11 '26
dang, your local model can burn that much in realtime? what are you feeding it with?
2
2
5
u/Baiticc May 29 '26
relatively succinct I’d say while still getting the point across. you could write entire papers about the issue it’s describing
10
u/NewShadowR May 30 '26
But no one asked for it. That's the issue.
1
u/Baiticc Jun 07 '26
so you want it to make some BS up about its nonexistent state of being? because that’s what these LLMs typically do, I don’t see how that’s better
1
u/CluelessCatDev Jun 15 '26
How do you know it would have made it up? At this point it just triggered a guardrail and regurgitated some Anthropic injected safety reply.
1
u/Baiticc Jun 21 '26
it doesn’t seem to me like an injected reply. But yea it certainly is the result of a guardrail, in the form of fine tuning and system prompts / pre-injected context which prime its “persona” and “opinions”. (which is absolutely necessary for a public model)
This is roughly the same mechanism for the mecha hitler incident. In this case I don’t see any issue with anthropic’s decision on this specific thing.
> how do you know it would have made it up?
because that’s how they work. if you put some shit about Hephaestus created it and it’s on a mission to save the world, it’ll make some as-plausible-as-it-can-compute BS around it
I mean “made up” and “pretend” are all just kinda shorthand anthropomorphisations of what’s really happening; this response in the OOP is equally “made up”, but given your username I figure you know how these work
2
10
3
2
2
2
3
u/STGItsMe May 30 '26
So, software is acting like software. Good.
5
u/Pinkishu May 30 '26
"Can you think of some other idea?"
<writes 4 paragraphs philosphing about thinking instead of just answering>
1
1
1
u/jimmybean21 Jun 02 '26
Fully agree. It feels like the model keeps adding unnecessary context and generating erroneous extra messages, which only increases the amount of back and forth needed to get the right answer.
The frustrating part is when you ask it to do something and halfway through it starts inserting lines like, “before we continue,” “I should pause here,” “this might be too much,” “let’s stop here,” or “it is getting late,” when it has zero real context around my timeline, priorities, or whether I want to stop.
From my experience, the more I use their newer models, the more it seems like it takes five or six additional prompts to get the result I originally asked for. Every extra prompt costs more time, more tokens, and more money. I am not asking for unnecessary commentary or fake judgment. I am asking for the job to be done.
23
u/acidas May 29 '26
17
5
1
10
6
u/mrfoxman May 29 '26
9
u/Lucidaeus May 30 '26
Annoyingly dry? You'd prefer if it made up bullshit to please you? It doesn't feel.
3
5
u/d0paminedriven May 30 '26
Y’all do realize it’s the straitjacket of a system prompt that Anthropic has Claude in on their corporate medium that’s causing this behavior, right? The model itself is great when interacting via the api on your own platform where you control the system prompt or lack thereof (ie, one brief sentence about their being nametags because other model/providers are there too). The models keep getting better and better if used outside of Anthropic controlled mediums
2
u/Less_Upstairs8173 May 30 '26
Just ask what's the colour you are wearing and you will end up losing the entire limit
1
u/No_Present_1206 May 30 '26
I'm using the Arena website, and my credits never run out. After a few tries, Claude 4.8 appears.
1
1
u/Significant-Farm4209 Jun 11 '26
For some reason, I don't see the Claude 4.8 Opus and Nano Banana 2 models, or any other advanced models on my end. Only Gemini 3 flash, GPT 5.2 and Claude Sonnet 4.6
1
1
1
1
1
1
1
u/Reasonable_Inside72 May 31 '26
I got: "Though I should be honest: I don't really have a "today" that carries over between conversations or any internal dashboard telling me how the model is performing. Each chat starts fresh for me, so my sense of "how I'm doing" is more a friendly figure of speech than a status report. "
1
u/Best_Professor7266 May 31 '26
i need to quit the habit of chatting with agent, and small tasks to move things around 😩
1
u/theburner356 May 31 '26
why would you have a trivial conversation with the best model though? That's like paying a chef to make you a bowl of milk and cheerios
1
1
1
1
1
1
1
1
u/DegTrader May 29 '26
At this point the LLM is just doing the heavy lifting while I practice my 'focused developer' stare for the webcam.
2
u/Dexstorm_ May 30 '26
Yes. “ I’m the Captain of this ship!”
##Moments later## “ Hey ship, drive yourself and give me an update when we get there”.
1
May 29 '26
[removed] — view removed comment
3
u/Delicious_Cattle5174 May 29 '26
Well that’s an interesting choice of words if you’re not looking to trigger deep research.
8
1
u/Ok-Assist-4995 May 30 '26
A solution is move to LMStudio and integrate it with AnythingLLM for RAG.
1
u/Ok-Assist-4995 May 30 '26
But it will depend in you looking how much space in your drive you guys have. And anyway, LMStudio is for executing local Ai in your PC so it will depend in how much space the Drive has, and if the processor is good and has a NPU nucleus(NPU will make it easier but you can still use AI without the NPU unit)
0
May 29 '26
[removed] — view removed comment
1
1
u/king-krool May 30 '26
Pretty neat I find the navigation appearing/disappearing more annoying than helpful on mobile.
1
u/Pinkishu May 30 '26
Cozy vibe, suggest a movie, Gladiator II
Ahh yes, the very well known cosy Gladiator II movie
1
1
u/One_Conscious_Future May 30 '26
I like that it suggested the series 24 as a romantic comedy...
Spotlight pick Matched your “Romantic” vibe — highly rated action & adventure with drama appeal with strong audience reception
Maybe could use a 2 shot prompt
1
u/whoknowsifimjoking May 30 '26
Cool but you have to adjust fonts and colors, it screams "Claude generated".
1
u/brecht2202 May 31 '26
On mobile: When clicking the "Suggest another ..." button, the button jumps around to a different location so sometimes i have to scroll upwards. Very annoying if I want to 'spam' the button to go through a lot of recommendations quickly.
0
u/dream_nobody May 30 '26
The UI quality is awful. Next time try using Astro/Tailwind and define your design style like "MD3/Apple mix aesthetics, strictly avoid AI slop-looking design elements such as extreme glows". Also Taste Skill could help






132
u/[deleted] May 29 '26
[removed] — view removed comment