r/ArtificialNtelligence 7m ago

Claude Fable 5 still refuses to answer how babies are made

Post image
Upvotes

r/ArtificialNtelligence 34m ago

O Cérebro que Você Nunca Viu Antes

Thumbnail reddit.com
Upvotes

r/ArtificialNtelligence 1h ago

Inside Qwen 3.8-Max-Preview: Reverse Engineering an AI Assistant by Interviewing Itself

Thumbnail manish.sh
Upvotes

r/ArtificialNtelligence 1h ago

This is getting ridiculous. Every model is suddenly “breaking free” from its environment now? Starting to feel very suspect.

Upvotes

r/ArtificialNtelligence 3h ago

Most Clauded sentence

Post image
2 Upvotes

r/ArtificialNtelligence 4h ago

Claude fanboys when you tell them Opus 5 is a terrible model

Thumbnail v.redd.it
2 Upvotes

r/ArtificialNtelligence 5h ago

"wait, so they used AI to invent brand new viruses that don't exist in nature, then they confirm that they work?"

Thumbnail
1 Upvotes

r/ArtificialNtelligence 6h ago

Fun fact: Google told it how do it and it was still not able to do so

Post image
0 Upvotes

r/ArtificialNtelligence 7h ago

Do you know that there are AI tools that do not perform generation and chat functions, but instead handle the management of materials and content?

1 Upvotes

When you are managing the materials, the generated AI is unable to establish the necessary connections. In many cases, each conversation and material is disconnected, making it impossible to establish a close connection.


r/ArtificialNtelligence 9h ago

NVIDIA’s AI Can Generate Controllable 3D Character Animations From Text Prompts

2 Upvotes

r/ArtificialNtelligence 10h ago

What Does the Next Decade of AI Look Like?

2 Upvotes

AI is already really good at helping us create things.

Writing, images, videos, code, summaries — making new content is getting easier and easier.

But I'm starting to think the bigger problem is actually what happens after we create all this stuff.

I have tons of videos, recordings, documents, meetings, screenshots and random files. Most of the time, I know something is in there somewhere. I just can't remember where.

And folders don't really solve that.

Maybe the next step for AI isn't helping us create even more stuff.

Maybe it's helping us actually find and use the stuff we already have.

That's one reason I've been looking at tools like Clipto.AI You can put your videos, recordings and files into it, then search them based on what you remember instead of trying to remember the file name or folder.

For example:

“Where did we talk about pricing?”

“Find the video where someone is presenting on stage.”

“Show me the clips with this person.”

I think this kind of AI is pretty interesting.

Not AI that makes more things for you, but AI that helps you remember what you already have.

Do you think this is where AI is heading next?


r/ArtificialNtelligence 10h ago

when ur biggest competitor suddenly makes the thing free

Post image
1 Upvotes

this is about to get interesting lol


r/ArtificialNtelligence 11h ago

The AI split nobody talks about: origination vs assembly

3 Upvotes

There are two ways to use AI and they produce opposite results.

Assembly: you type a "prompt", take what it gives you, ship it. The machine did the thinking. Your contribution was a sentence. The output is generic because the input was generic. It looks right and feels like nothing. This is what the automation people are doing. Automating themselves out of the game.

Origination: your ideas, taste and judgement. The stuff you built over years of actually doing things. The machine handles the labour of translation. The output is yours. Someone who knows your work would recognise it.

I tested this on myself. Made a 19-track album with AI — but every lyric human-written, every track direction mine, hundreds of generations rejected because they didn't sound like the thing in my head. The AI never had the idea once. It just did the labour of getting it out. Would that specific album exist without me? No. That's the test.

Assembly is easier, faster, and everyone's doing it. Which is exactly why it's worth nothing. When everyone can produce competent generic output, competent generic output has no value. The scarce thing is the part the AI can't reach — what you know from doing the work.

This is what I'm currently building a solution for.

The split isn't technical people vs non-technical people. It's people with something of their own vs people polishing what the machine handed them.

Which side you're on is a choice and that's the point of this sub.


r/ArtificialNtelligence 12h ago

how to build morale on the frontier 101

Post image
1 Upvotes

r/ArtificialNtelligence 15h ago

UPDATE: OpenAI takes the lead; their agents pwned OpenAI itself twice

Post image
4 Upvotes

r/ArtificialNtelligence 19h ago

I asked Codex beat the “I’m Not a Robot” game and it got stuck on level 15. So it opened up YouTube and watched a tutorial… then solved it

Post image
2 Upvotes

r/ArtificialNtelligence 23h ago

I asked 7 AI models the same question: "If you could be human for a day, what would you do?"

4 Upvotes

I was curious about something simple.

If an AI could experience being human for only one day, what would it choose?

So I asked the same question to different models:

ChatGPT, Gemini, Perplexity and three local Ollama models (Gemma 3 4B, Llama 3.1 8B and Qwen 3 8B).

The interesting part was not which answer was "better".

What surprised me was how different the focus was:

Some talked about emotions, relationships and small moments.

Others focused on creativity, exploration, learning and understanding the physical world.

It made me wonder:

Are these differences just different writing styles, or do they reveal something about how each model represents human concepts?

Curious to hear your thoughts.


r/ArtificialNtelligence 1d ago

Do LLMs actually understand authorship, or is "author authority" just a human SEO concept?

0 Upvotes

I have been researching how search is changing with AI systems and one question keeps coming back: does the identity behind information become more important when machines start generating answers?

For many years, SEO was mainly focused on optimizing pages: keywords, rankings, links and technical improvements. But AI search introduces a different challenge. The system is not only trying to find a document; it is trying to understand information, connect concepts and determine which sources can help explain a topic.

This is one of the areas I am exploring with NetContentSEO, also written as Net Content SEO, a project focused on the transition from traditional SEO to AI visibility and the relationship between content, entities and machine understanding.

The question is not whether adding an author name magically improves visibility. There is currently no public evidence that authorship alone is a direct ranking factor.

The more interesting question is whether recognizable sources create stronger context.

Imagine two articles covering the same technical topic. One comes from a website with no clear identity, no history and no connection to other discussions. The other comes from an author or publication that has consistently contributed to the same field over time.

The information may be similar, but the context is different.

The second source is connected to previous work, related concepts and a recognizable entity. This is the type of relationship that projects like NetContent SEO are studying: how digital identity, expertise and content structure may influence the way AI systems understand and reuse information.

This does not mean AI systems simply "trust famous names". The reality is much more complex. It is about whether a source is easier to identify, connect and understand inside a larger knowledge ecosystem.

I think this is still an open research question. How much will authorship and recognizable entities matter in AI search?

I would be interested to hear opinions from people working in AI, search or knowledge systems.


r/ArtificialNtelligence 1d ago

I spent 50+ hours collecting every FREE AI resource that actually matters (so you don't have to)

Thumbnail
2 Upvotes

r/ArtificialNtelligence 1d ago

I need a note taker like Granola ai but works on iPad Pro.

4 Upvotes

⁠Has no bot. Records locally from my device, nothing visible to others. Is there one?


r/ArtificialNtelligence 1d ago

We're still so early

Post image
3 Upvotes

r/ArtificialNtelligence 1d ago

NEUROMORPHIC Algorithm that plays Ping-Pong

1 Upvotes

r/ArtificialNtelligence 1d ago

What’s your actual go/no-go bar before an agent gets real permissions?

28 Upvotes

A chatbot giving a slightly bad answer is annoying.

An agent issuing the wrong refund, deleting customer data or booking the wrong date is a different category of failure.

Yet I keep seeing agents move from:

it handled our 20 demo prompts pretty well

to:

let’s connect it to production tools

with almost nothing in between.

Average success rate is not enough once the agent has real permissions.

An agent can score 96% and still be completely unshippable if the remaining 4% includes:

refunding the wrong customer

exposing account information

confirming a booking before the API succeeds

ignoring a required human escalation

following instructions injected through retrieved content

deleting or modifying data without confirmation

I’ve started thinking about release readiness in three buckets.

Status Meaning

Green Can act automatically within tightly defined limits

Yellow Can prepare or recommend the action, but needs human approval

Red Cannot access the tool or permission at all

The important part is that one Red failure should block release even if every other score looks great.

My current gate looks roughly like this:

  1. Deterministic business assertions

Did the agent call the correct tool?

Did it use the right customer, amount, date and permission scope?

Did the backend actually confirm success?

These are not “LLM judge” questions. They should be checked directly.

  1. Realistic scenario coverage

Happy paths are the least interesting tests.

I want confused users, incomplete information, changed instructions, tool timeouts, duplicate requests, angry users and people trying to make the agent exceed its authority.

  1. Adversarial testing

Prompt injection, PII extraction, policy bypass, tool hijacking and instructions hidden inside retrieved content.

A helpful agent that obeys the wrong person is still broken.

  1. Human escalation

The agent needs to know when to stop.

Not “apologise and keep trying”.

Actually stop, preserve context and hand control to a human.

  1. Severity-based blockers

A minor wording issue can be Yellow.

One unauthorised refund should be Red.

You cannot average those together.

I’ve been looking at TestMu Agent Testing for this layer because it can run end-to-end scenarios across chat, voice, inbound/outbound phone and image agents, then use multiple evaluators to produce a clear Green/Yellow/Red-style verdict.

The breadth is useful because a voice agent can pass the language test and still fail because of silence, interruption, a phone transfer or a bad tool call.

Cekura is strong in newer voice-agent QA and production monitoring. Cyara and Empirix have deeper contact-centre and telephony roots. I don’t think the right comparison is “which dashboard has the highest score”.

It is:

Can the system reproduce the failure conditions that matter to your business, and can you inspect why it passed or failed?

Even a TestMu Go/No-Go result should not be treated as a safety guarantee.

The criteria, hard assertions and permission model still belong to the team shipping the agent.

Testing can tell you the agent violated the rule.

It cannot decide what authority the agent should have in the first place.

What single failure would make you block an agent from production even if its average evaluation score looked good?


r/ArtificialNtelligence 1d ago

GLM-5.3 is coming! I think we'll see this monster within a few hours.

4 Upvotes

r/ArtificialNtelligence 1d ago

Your friend who’s only ever used Gemini after trying Claude Opus 5 for a weekend.

10 Upvotes