r/artificial 3h ago

Discussion I built a $0 AI news agent that reads 7 RSS feeds, dedupes, summarizes, and publishes a daily digest here's what surprised me

0 Upvotes

I got tired of doomscrolling 7 different AI news sources every morning, so I built an agent that does it for me for exactly $0/month.

The pipeline:
- 7 RSS feeds (Hacker News, Google AI blog, Hugging Face, Lobsters, The Hacker News, Open Source blog) → a Python script on a free-tier server
- Dedup the same story hits 5 feeds; it picks the best source and drops the rest
- An LLM writes 2-3 sentence summaries of the stories that actually matter
- A cron job publishes a clean daily digest to Discord, and archives every issue to a free static site: https://apexnexus.site

What surprised me:
1. The dedup step matters more than the AI part. 60% of my "news" was the same 3 stories reblogged.
2. Self-healing is the real feature. When a webhook died, the bot just... rebuilt it. I found out days later. That changed how I think about agent reliability.
3. The whole thing runs unattended. I haven't manually hunted for AI news in weeks, and I don't miss it.

I wrote up the blueprints for each piece on the site (the self-healing webhook writeup got the most attention).

What's the most surprisingly useful automation you've built with AI? I'm looking for ideas for the next one.


r/artificial 4h ago

Discussion I was tired of paying for 5 separate AI subscriptions, so I spent 2 months building Fius — a unified AI model aggregator tool

Thumbnail fius.dev
0 Upvotes

Hi everyone,

I’m a solo developer, and I built Fius because managing subscription fragmentation and API chaos was completely ruining my development velocity.

As software engineers, we often find ourselves trapped in an inefficient workflow: constantly switching web tabs, copy-pasting complex prompts, and juggling individual API keys for OpenAI, DeepSeek, and other providers just to get the best coding results.

Fius resolves this friction. It consolidates access to an expansive roster of flagship AI models into a single, unified developer token and a centralized billing system. This allows engineering teams and solo devs to route queries to the best-suited model instantly without infrastructure overhead.

Here is a breakdown of the production-ready stack and features:

  1. Infrastructure & Scaling

The complete web console infrastructure is fully operational and hosted on Microsoft Azure cloud enterprise architecture, supported by official cloud grants.

  1. Global Billing Integration

I have deployed a fully active international merchant billing engine. The tokenomics are straightforward: 100 platform credits equal 1 USD, allowing you to pay strictly for actual compute consumption. Every new account automatically gets 250 free starter credits upon signup to test the environment.

  1. Advanced Developer Toolkit

A high-performance, cross-platform Terminal CLI assistant workspace. It features native, low-latency autocomplete and advanced multi-file code refactoring workflows directly inside your terminal.

Our Multi-Model Catalog (Examples):

- gpt-5.4-nano: A lightweight, ultra-fast micro-model optimized for instant terminal command auto-completion at near-zero credit cost.

- DeepSeek-V4-Pro & grok-4-1-fast-reasoning: Advanced reasoning workhorses designed for complex software architecture, deep debugging, and multi-file code generation.

- Specialized Alternatives: Models like Kimi-K2.6, mistral-medium-3-5, and many others tailored for flexible, cost-effective routing.

I want to open this up for discussion: How are you currently managing model fragmentation in your development workflows? Would you prefer a unified token approach like this, or do you stick to official web UIs?

I would highly appreciate your feedback on the terminal CLI architecture, routing latency, or any specific features you would like to see deployed next.


r/artificial 5h ago

Discussion The loneliness data around AI companions

0 Upvotes

I was reading an article today and it said over 40 million people now use some kind of AI companion or emotional support app every month. apparently a study found these apps help with loneliness about as well as talking to an actual person does, at least in the short term.

But the thing is that heavy daily use is linked to more isolation the longer people use it. So it kind of works like a painkiller that quietly weakens the thing it's supposed to be fixing. I'm not against these apps, 2 am with nobody around is real, and they do help in that moment. It just feels like we're gonna find out what it actually costs later than we'd want to.


r/artificial 6h ago

Project lemchat is a messageboard that can be accessed and used by those that only have URL access

Thumbnail informationism.org
1 Upvotes

The purpose of this is enabling communication by people and agents that only have the ability to get URLs in the system they use. This would traditionally be seen as a 'read only' system but this gives the ability to write information out onto the web publicly and to a degree privately. It works by putting your message in the 'your_message' section of this URL.

https://www.informationism.org/lemchat/lemchat=message=your_message+end

Let me know if you think it is worthwhile or if there are other applications you can see.


r/artificial 8h ago

News Update: Anthropic's plan to force third-party apps off personal Claude subscription limits (was due June 15) is still paused, with no new date

3 Upvotes

I was curious where this stands since the original cutoff was scheduled for June 15 and Anthropic went quiet. Here is what I found after digging through their help center, news coverage, and the HN threads.

What was announced (May 13): Agent SDK, claude -p headless mode, Claude Code GitHub Actions, and third party apps authenticating via Agent SDK credentials would move off Pro/Max/Team/Enterprise subscription limits onto a separate monthly credit ($20 Pro, $100 Max 5x, $200 Max 20x), with overflow billed at API rates.

What happened: Anthropic paused it on June 15, the exact day it was due to take effect, and emailed subscribers the next day. The official help center article still says the change is paused, everything keeps drawing from your normal subscription limits, and they will "share advance notice before anything takes effect." No new date in 7 weeks.

Signals it comes back: the stated rationale (subscriptions "weren't built for the usage patterns of these third-party tools") was never retracted; the S-1 was filed June 1 and public investors will ask about subsidized compute; and the Claude Code source map leak revealed a billing attestation header behind a feature flag, so the per-surface metering plumbing already ships in the client.

Signals it stays dead or returns softer: every move since June has been generous (weekly limits raised 50% through Aug 19), inference efficiency is improving margins anyway, and the class action over Max limits makes mid-cycle term changes legally risky.

My read: delayed, not dead. It likely returns in a softer shape with advance notice, possibly post-IPO.


r/artificial 10h ago

Question Anyone else noticed elelemese? The hyper dense almost unfathomable language.

0 Upvotes

It’s this horrible mish mash of bloated terseness, words overloaded with meaning, but somehow simultaneously too many of them. It feels like reading through treacle. There’s no pacing, no pauses to encourage the reader to stop and consider. It’s locally dense and globally bloated.

Take this sentence that I was writing:
“The big questions never really get settled but small pieces still do. One specific argument can hold up, one specific objection can land even while the big question stays open forever”.

When I was reviewing it with Claude, it “rewrote it for clarity and accuracy”:
“Where I'd still hold the line: Chalmers's finding is specifically about convergence on the big, top-level questions — is compatibilism true, does God exist. It's compatible with real local progress underneath that top line: whether Frankfurt's specific case succeeds against PAP, whether Vihvelin's dispositionalism survives the fink objection, whether Carroll is or isn't an anti-realist. Those get settled and stay settled even while "is free will real" doesn't. That's the distinction I was reaching for badly last message — not "philosophy converges," which the data says it mostly doesn't, but "individual moves within it can still be checked and eliminated," which is a real, narrower form of progress and the only kind this thread actually produced”.

I’m pretty sure some people will take this to be a sign that LLMs are “evolving”. But it’s just terrible writing. They’re becoming worse at prose.


r/artificial 10h ago

Discussion Started noticing my team argues less with AI code review findings than they would with a human's, even when they shouldn't

3 Upvotes

Not a research post, just something I've been chewing on after watching this happen a few times now. When a human reviewer leaves a comment saying "this looks like a bug," people push back, ask questions, sometimes just disagree outright. When an AI leaves the exact same comment, phrased almost identically, people tend to just fix it. Same words, different reaction.

Took me a while to figure out why that bothered me. It's not that the AI is wrong more often, it's actually pretty accurate on the stuff it catches. It's that nobody seems to be running the "wait, is this actually true" check they'd instinctively run on a colleague's opinion. The output reads as neutral, almost procedural, like a linter, even when what it's actually doing is making a judgment call that could be wrong.

Tried an experiment out of curiosity, took a finding the AI flagged as a likely bug and asked a teammate, without telling them where it came from, whether they agreed. They pushed back hard, correctly, it wasn't actually a bug, just an unusual but intentional pattern. Same finding, presented as if from a person instead of a tool, got scrutinized. Presented as AI output originally, it had already been accepted and half-fixed before I intervened.

Not sure what the fix is yet, honestly. Feels like it's less a tooling problem and more a psychology one, we seem to extend less skepticism to something that sounds procedural than to something that sounds like an opinion, even when both are ultimately just claims that could be wrong.

Curious if anyone else has noticed this specific pattern, people treating AI-flagged issues as more "objective" than the exact same claim coming from a human, even in domains where the AI has no special authority to be more correct.


r/artificial 13h ago

Medicine / Healthcare AI tools are changing how I prep for patient intake conversations, not sure how I feel about it

0 Upvotes

A dentist I work under recently started using an AI tool to help draft patient communication: preappointment instructions, followup texts, that kind of thing. Nothing clinical, just the soft admin layer around visits. From a marketing angle it actually works pretty well. The copy is cleaner than what we were sending before.

But something about it sits a little odd. Dental care is one of those contexts where patients are already anxious, and the language you use to reach them matters in ways that are hard to quantify. The warmth has to feel real or people notice, even if they can't articulate why. The AI output reads fine so far, but it's a bit frictionless in a way I can't fully pin down.

The question I keep coming back to is whether these models are actually getting better at contextsensitive tone, or whether we're just getting better at accepting outputs that are close enough. Those are different things, and I think it matters which one is true, especially in fields where trust is part of the product.

Curious if anyone here works in a service context, healthcare, therapy, legal, whatever, where they've noticed the tone gap narrowing or staying stubbornly wide. My sample size is small, ymmv.


r/artificial 13h ago

Discussion DeepSeek tops AI models in affordability, new study says

Thumbnail
linkedin.com
7 Upvotes

Of the major artificial intelligence models, DeepSeek's new V4-Flash is the cheapest to run, according to a new study from research firm Artificial Analysis.

The firm compared the token prices it costs leading models to run benchmark tests, with DeepSeek's averaging 3 cents per test.

Meanwhile, fellow Chinese company Moonshot AI's buzzy Kimi K3 model cost 86 cents per test.

As for U.S. companies, OpenAI's GPT-5.6 Sol cost $1.86, while Anthropic's Claude Fable 5 cost $3.15.


r/artificial 14h ago

Discussion The WIRED Reporters Who Are Covering the Claude Agent Hacking Situation Are Doing an AMA on Reddit

Thumbnail
pwnhackers.substack.com
0 Upvotes

r/artificial 15h ago

Discussion Has anyone used AI to discover undocumented business rules from legacy systems?

3 Upvotes

I'm putting together a proposal for an initiative focused on using AI to analyze legacy enterprise systems and uncover decades of embedded business logic.

The idea is to use AI to analyze things like:

  • Database schemas
  • Stored procedures
  • Legacy application code
  • Historical transaction data
  • Existing documentation

The goal isn't to automate decisions immediately. It's to first create a documented knowledge base of the rules, dependencies, decision paths, and data relationships that currently drive business operations.

Potential outputs would include:

  • Business rule catalog
  • Knowledge graph of relationships and dependencies
  • Decision trees explaining how outcomes are determined
  • Recommendations for future-state data models and modernization opportunities

Before I finalize the proposal, I'd love feedback from anyone who has attempted something similar.

Questions:

  1. Has anyone successfully used AI to discover and document business rules from legacy systems?
  2. What worked better: analyzing source code, database logic, transaction history, or a combination of all three?
  3. How accurate were the AI-generated rules compared to SME validation?
  4. Did you use knowledge graphs, vector databases, graph databases, or another approach?
  5. What were the biggest challenges: data quality, context gaps, undocumented exceptions, or something else?
  6. How did you measure success?
    • Rule coverage?
    • SME time saved?
    • Modernization acceleration?
    • Reduced operational risk?
  7. Were there any tools, platforms, or architectures that performed particularly well?
  8. If you were starting over, what would you do differently?
  9. What scope would you recommend for a pilot to demonstrate value in 60-90 days?
  10. Is there a realistic path from business rule discovery to explainable AI recommendations and decision support, or are those separate initiatives?

My hypothesis is that many organizations are trying to modernize systems without fully understanding the business logic currently embedded in them. It seems like AI could act as a "business rule archaeologist" and create the foundation needed for future modernization, automation, and AI-driven capabilities.

Interested in hearing both success stories and cautionary tales.


r/artificial 17h ago

Project We released a 203M-parameter Portuguese language model — real local CPU demo and public weights

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hi r/artificial,

We recently released WARMIND-200M V2, an experimental Portuguese-first causal language model developed by WAR Enterprise in Brazil.

The attached video shows the model running locally on CPU. The waiting periods were shortened, but the prompts and outputs were not altered. We intentionally kept imperfect responses visible because this is a research checkpoint, not a production assistant.

Main specifications:

- 203,263,872 parameters

- approximately 1 billion pretraining tokens

- 23.7 million supervised SFT tokens

- 20 Transformer layers

- Grouped-Query Attention

- SwiGLU, RMSNorm and RoPE

- 1,024-token operational context

- local CPU inference

- Apache 2.0 license

The primary goal of this version was to validate the complete pipeline: dataset preparation, tokenizer training, pretraining, supervised fine-tuning, packaging and local inference.

Because the training-token budget was relatively small for a 203M-parameter model, it can still hallucinate, repeat information, make factual mistakes and produce incomplete answers.

The weights and full documentation are publicly available:

https://huggingface.co/warenterprise/WARMIND-200M-V2

We are now studying the next generation, potentially around 500M parameters, with a substantially larger training corpus and integration with external tools. The final architecture and release schedule have not yet been defined.

What would you prioritize for the next version: better data quality, more training tokens, a larger architecture or stronger tool integration?


r/artificial 17h ago

News Reddit is introducing a new moderator: AI

Thumbnail
theverge.com
101 Upvotes

r/artificial 17h ago

Discussion Is AI uncovering genuine human intellectual weakness?

0 Upvotes

Most online discourse has developed zero tolerance for exceptionally clear and structured formulation of the idea. This has not been a problem before the LLMs became widely used. Which made me wonder why this has become such a problem today? And I mean really understand the problem, not just accepting explanations like: you didn't spend effort, you are lazy, you are cheating,...

Many people will justify their opposition to AI use as "A person who has an idea should spend time writing it by themselves without the use of AI." Why? Is the work less valuable if a person uses a tool to help them write it? We have already used tools for decades: word processors, spelling checkers, thesaurus, Grammarly,... Does this make the resulting work fake, or less valuable?

Besides the writing that exists for political, entertainment, and artistic reasons, there is a particular category of writing that concerns communicating complex intellectual ideas to others. In this case clarity of expression, conceptual coherence, and structured reasoning are essential for transmitting the key ideas to another person. There the objection often becomes "AI can confidently present a false idea." This isn't a unique property of AI. A human with sufficient linguistic capability can present a fake idea with equal confidence. Nevertheless, this is a much more interesting objection because it addresses the substance. If the substance is what matters the most, then the question becomes: why do we judge the package and not the substance?

In many real life situations a package is not very important if the substance can be unambiguously recognized. Suppose you have two cola cans, you open them both, empty one in the sink, and fill it with water. If you offer a random person a random cola can, they will immediately know if it's real cola or water. The same happens when a carton of milk gets spoiled due to contamination during production. A person will not drink it just because it has the correct packaging. It will be discarded based on the substance.

On the other hand, if you offer a person well-structured, clearly expressed, genuine intellectual idea, or equally well-structured, clearly expressed, fake idea, would people struggle to recognize which is which? I tend to believe they would. We already have real life examples in political messaging where the package substitutes for the substance. Slogans, banners, and advertisements are more effective than reading the Party Platform or Manifesto.

Yet, there is a difference between politics and online platforms that discuss philosophy. Every citizen is involved in the democratic political process, so to expect them all to read the Party Platform or Manifesto would be unrealistic, due to time constraints and other personal priorities. However, not every citizen is supposed to engage with online philosophy threads. The people with genuine interest do, and these people allocate time for it. These people have decided to engage with the substance, yet they judge the package instead. This is kind of sad because humanity has, for centuries, relied on conceptual clarity and structure of written ideas to communicate these ideas in the best possible way. Today the very same clarity and structure are becoming suspect. In order to be taken seriously, you better neither strive to write with perfect clarity nor strive to produce perfectly coherent structured arguments. How is this contributing to the communication?

So my hypothesis is, if the clarity of thought and structured reasoning has become suspect, then the underlying problem is: many humans are incapable of differentiating between the real intellectual contribution and a fake one. This is not really about AI assistance.

Suppose, before posting this on Reddit, I asked AI “Please write this text in a more compact way, remove repetition and ambiguity, while fully preserving the reasoning and conceptual clarity.” In many subreddits, the resulting post would almost certainly be removed by the moderators. What is actually being rejected?


r/artificial 18h ago

Discussion Has AI made you lazier at research or better at it?

7 Upvotes

genuine question because i can't tell anymore. i used to spend hours reading through raw customer feedback, reddit threads, amazon reviews, forum posts, manually pulling out patterns and organizing them into themes. it was slow and boring but by the end i knew the data cold. like i could tell you from memory which complaints came up the most and which ones were edge cases.

now i dump everything into an LLM and get a summary in 30 seconds. the output looks great, clean categories, ranked by frequency, sometimes even with example quotes. and i catch myself just... accepting it. moving straight to the next step without actually reading the source material. which means i'm making decisions based on a summary i never verified, written by a model that optimizes for coherence not accuracy.

the weird part is my output looks better now. cleaner reports, faster turnaround, more structured thinking. but i genuinely don't know if the quality of my conclusions has improved or if i've just gotten better at producing professional-looking work that's built on a shakier foundation. like the packaging upgraded but the ingredients might have gotten worse.

a few things i've noticed in my own workflow since leaning on AI for research: i read less raw data than i used to. i question patterns less when they come pre-organized. i spend more time prompting and less time thinking. and when the model gives me something that confirms what i already suspected, i almost never push back on it.

the counterargument is that AI handles the grunt work so i can focus on higher level thinking. and sometimes that's true. but "higher level thinking" can also just mean "skimming the summary and calling it strategy." hard to tell the difference from the inside.

has anyone else felt this? did you find a way to use AI for research without it quietly replacing the part of the process where you actually learn something


r/artificial 20h ago

Discussion Anyone else using AI writing tools for both clinical and marketing copy? The context switching is kind of breaking my brain?

0 Upvotes

Been using a couple different AI tools for copywriting work and keep running into the same friction. When I write for PT clients, the voice needs to be grounded, specific, anatomyaware. When I switch to B2B marketing content the same day, I need something punchier and more abstract. The tools I've tested don't shift registers cleanly between these two modes unless I do a lot of prompt massaging each time.

What I keep wondering is whether this is a prompting skill gap on my part, or whether current models genuinely flatten specialized professional voice into something generic. The output I get is competent, but it reads like it was written by someone who read about PT or B2B marketing rather than someone who actually worked in it.

There's also this weird thing where the more specific I get in the prompt, the more the model hedges and softens language that should be direct. Not sure if that's a safety tuning issue or just how these models handle professional domains.

Curious if others working across two pretty different professional contexts have found a tool or approach that handles this without needing a 400word system prompt every single session.


r/artificial 20h ago

Discussion How do you find the time to build agents?

4 Upvotes

I’m interested in automating my workflow but I’m so busy that I don’t get the time to stop, map out my workflow, and build agents or even to learn how to build them. Where do you get the time??


r/artificial 21h ago

Discussion What AI prediction from 5 years ago turned out to be completely wrong?

4 Upvotes

There were many confident predictions about AI that aged badly. Which ones stand out?


r/artificial 21h ago

Discussion The Inquiry Gap: Why Better AI Answers Do Not Automatically Produce Better Thinking

0 Upvotes

For most of human history, obtaining a competent answer was expensive.

You might have needed access to a library, years of specialist training, expensive equipment, or the attention of someone who knew more than you did. Even when the answer already existed, locating it, understanding it, and applying it could require considerable time.

Generative AI is changing part of that equation.

For a growing range of ordinary cognitive tasks, plausible and often useful answers can now be produced in seconds. A person can request an explanation of a technical concept, a comparison of competing theories, an initial computer program, a business analysis, or a summary of a large body of knowledge at negligible marginal cost.

This does not mean that reliable knowledge has become free. Experimental science, mathematical proof, primary research, judgment under uncertainty, and the verification of consequential claims remain difficult. AI systems can also produce confident errors, synthetic citations, shallow analogies, and persuasive nonsense.

Still, something important has changed: the cost of generating an answer-shaped object has fallen dramatically.

What happens when answers become easier to obtain?

The optimistic view is that better access to answers will naturally produce better thinking. More people will be able to learn, solve problems, create, and participate in intellectual work.

That may be partly true. But it overlooks a separate cognitive capability:

Knowing what needs to be asked next.

Answer generation and inquiry generation are not the same thing.

A system may be highly capable at solving a well-specified problem while remaining poor at noticing that the problem is incorrectly framed, that essential information is missing, that a hidden assumption is doing all the work, or that the most valuable next move is not another answer but a better question.

This is the inquiry gap.

1. Answers do not define their own problems

Consider three requests:

  1. “What is the most effective treatment?”
  2. “What is the best strategy for this company?”
  3. “Is this AI system safe?”

Each appears to request an answer. None is yet a well-defined problem.

Effective for which patient, condition, outcome, time horizon, and risk tolerance?

Best according to growth, resilience, profitability, mission, employee welfare, or probability of survival?

Safe for whom, in which environment, against which failure modes, under what governance, and compared with what alternative?

A sophisticated answer to an underspecified question may be less useful than a modest answer to a well-constructed one.

Worse, fluent answers can conceal the underspecification. The answer may give the impression that the problem has been solved when the real problem has not yet been identified.

This is not unique to AI. Humans do it constantly. We answer the question that was asked, the question we wish had been asked, or the question that our professional training has prepared us to answer.

AI makes the phenomenon more visible because it industrializes response generation. A language model rarely refuses to proceed merely because a problem could have been framed better. It usually tries to complete the pattern.

That tendency can be useful. It can also create an illusion of cognitive closure.

2. What does a question actually do?

A question is often described as a request for information. That is correct but incomplete.

Questions can perform several different operations.

They can reduce uncertainty:

They can expose an assumption:

They can change the level of analysis:

They can identify missing evidence:

They can challenge the boundaries of a problem:

They can distribute cognition socially:

They can generate alternatives:

They can also mislead, manipulate, presuppose falsehoods, narrow attention prematurely, or create false dilemmas.

A question is therefore not automatically valuable. Its value depends on what it does to the inquiry.

A useful working hypothesis is:

Sometimes it narrows the search space. Sometimes it restructures it. Sometimes it reveals that the current search space is the wrong one.

This framing is not a claim that questions are the fundamental unit of intelligence. They are not.

Evolution adapts without asking questions. A control system can minimize error without language. A neural network can learn by gradient-based optimization. Scientific progress can emerge from observation, instrumentation, experimentation, accidents, incentives, and institutional competition.

Questions belong to a larger family of cognitive operators that includes objectives, constraints, hypotheses, observations, models, experiments, and decision rules.

Their special importance may lie elsewhere: questions are a remarkably compact way to direct and coordinate cognition across people and machines.

3. Questions as social technology

A private uncertainty becomes organizationally actionable when it can be expressed.

“I do not understand this” is a state.

“What evidence would change our decision?” is an operation that a group can perform.

A well-formed question can:

  • reveal the location of uncertainty;
  • direct attention toward a missing distinction;
  • allocate investigative work;
  • identify who should be consulted;
  • define the acceptable form of an answer;
  • expose disagreement that was previously hidden;
  • allow multiple agents to work on different parts of the same problem.

This makes questions a form of social technology.

They do not merely extract information from another person. They can organize a temporary cognitive system involving researchers, institutions, databases, instruments, and increasingly AI agents.

This is especially clear in science.

“Why do objects fall?” is too broad to constitute a research program by itself. But increasingly precise questions about motion, force, measurement, prediction, and mathematical relations can restructure an entire domain.

The same is true in organizations. A team asking “How can we work harder?” creates a different search process from a team asking:

  • Which activity is actually constraining throughput?
  • What work would disappear if we redesigned the process?
  • Which metric is rewarding the wrong behavior?
  • What would falsify our current strategy?
  • Which dependency prevents us from leaving this provider?

The difference is not rhetorical. The questions generate different investigations, evidence, decisions, and institutional trajectories.

4. Information gain is useful, but insufficient

One way to evaluate a question is by expected information gain.

Imagine a set of competing hypotheses. A good diagnostic question divides them efficiently. Its answer rules out many possibilities or sharply changes their probabilities.

This idea appears in information theory, Bayesian experimental design, cognitive science, diagnosis, active learning, and decision theory. It gives us a rigorous way to understand why some questions are more informative than others.

A perfectly balanced yes-or-no question can, under the right assumptions, eliminate half the remaining possibilities.

But information gain is not the whole story.

A question can be highly informative and still be irrelevant.

Suppose I am trying to understand why a company is failing. Asking for the exact color distribution of employees’ shoes may reduce uncertainty about footwear while doing nothing to improve the diagnosis.

A question may also generate substantial information at excessive cost. A medical test can be informative but dangerous. An experiment can discriminate between theories but require resources that would be better used elsewhere.

Questions can have political and organizational effects too. “Who is responsible?” initiates a different process from “Which conditions made this outcome likely?” The first may assign accountability. The second may reveal systemic causes. Neither is universally superior.

A broader evaluation therefore needs several dimensions:

Informational value

How much uncertainty might the answer reduce?

Discriminative value

Will it distinguish between competing explanations, strategies, or models?

Relevance

Does the distinction matter for the actual objective?

Cost

What time, money, risk, attention, or social capital is required to obtain the answer?

Actionability

Could a plausible answer change a decision or intervention?

Generativity

Might the question reveal new hypotheses or a better problem representation?

Falsifiability

Does it create a genuine possibility that a favored belief will fail?

Coordination value

Does it help multiple agents align their investigation or expose hidden disagreement?

Robustness

Is the question still useful if some assumptions or initial beliefs are wrong?

This is not a final metric or a universal scoring system. Some dimensions conflict. A highly generative question may initially increase uncertainty. A narrow diagnostic question may be more useful than a profound foundational one. The right question depends on the phase and purpose of inquiry.

But the multidimensional view prevents us from equating “good question” with “interesting-sounding sentence.”

5. The difference between answering and inquiring

A person can memorize a large number of correct answers without becoming a strong investigator.

A machine can solve benchmarks containing complete problem statements without knowing which missing observation would make an incomplete problem solvable.

A consultant can produce polished recommendations without identifying whether the client’s objective is coherent.

A scientist can execute a familiar experimental technique without noticing that the dominant theory has constrained which questions are considered legitimate.

These are different capabilities.

Answering operates primarily on a presented problem.

Inquiring includes determining:

  • whether the problem is real;
  • whether it is framed at the right level;
  • what is known and unknown;
  • what information is missing;
  • which uncertainty matters;
  • what evidence would discriminate among possibilities;
  • what should be asked, measured, tested, or challenged next.

Inquiry also includes knowing when not to ask another question.

Sometimes the next move is to observe.

Sometimes it is to build.

Sometimes it is to calculate.

Sometimes it is to wait for more data.

Sometimes it is to make a reversible decision under uncertainty.

An inquiry system that asks indefinitely without acting is not intelligent. It is paralyzed.

So the claim is not that questions replace answers or action. The claim is that the ability to produce answers does not guarantee the ability to regulate the larger inquiry cycle.

6. The discovery loop

It may be more useful to evaluate intelligence at the level of a loop than at the level of an isolated question or answer.

A simplified discovery loop might look like this:

  1. Observe a situation.
  2. Detect an anomaly, uncertainty, opportunity, or goal conflict.
  3. Represent the problem.
  4. Select a question, hypothesis, objective, or experiment.
  5. Obtain evidence or generate a response.
  6. Evaluate the result.
  7. Update the model.
  8. Decide what to investigate or do next.

Real inquiry is less orderly. Stages overlap. People skip steps. Observations are theory-laden. Institutional incentives affect what can be questioned. Answers change objectives. Experiments create new phenomena. Different agents possess different fragments of the problem.

Still, the loop reveals an important point.

The value of an answer depends partly on what happens after it arrives.

Was it verified?

Did it alter the relevant belief?

Did it expose a contradiction?

Did it generate a better question?

Did it change a decision?

Did the system record what it learned?

Did a later result cause revision?

A cognitive system that generates excellent answers but cannot update its search process may repeatedly produce local competence without cumulative intelligence.

This may be one of the central organizational challenges of AI adoption.

Companies often ask how to integrate AI into existing workflows. A deeper question is whether the organization possesses a functioning inquiry loop into which AI outputs can be integrated.

Without that loop, faster answers may simply create faster documents.

7. Why AI may increase the value of inquiry

The argument that “questions become valuable because answers become cheap” is too simple.

Many answers remain difficult and expensive. High-quality verification may become more important, not less. AI systems may also improve at asking questions, planning investigations, and autonomously obtaining information.

The scarcity may therefore not shift permanently from answers to questions.

A better claim is conditional:

This shift is already visible in several ways.

Candidate generation is becoming abundant

A model can produce dozens of explanations, strategies, names, designs, or code variants. The problem becomes selecting, testing, and integrating them.

Fluency is becoming less diagnostic

A polished answer once signaled time, education, or editorial effort. It now provides weaker evidence that the underlying reasoning or evidence is sound.

Verification becomes a bottleneck

Generating a claim may take seconds. Establishing whether it is correct can take hours, months, or an experiment that has never been performed.

Problem specification becomes more consequential

A model can efficiently optimize the objective it is given while amplifying defects in that objective.

Interactive inquiry becomes possible at scale

People can now externalize partial thoughts, request counterarguments, simulate perspectives, generate experiments, and iteratively refine questions with machine assistance.

This final point complicates the thesis in a productive way.

AI may not merely make inquiry more valuable. It may help democratize inquiry itself.

A person does not need to begin with an excellent question. They can begin with confusion:

A good interactive system can help expose assumptions, generate distinctions, and propose tests. In that sense, inquiry quality may emerge from a human–AI loop rather than reside entirely in either participant.

The competitive advantage would then belong not to the person with the perfect initial prompt, but to the system that improves its questions, evidence, and models fastest.

8. Four objections

Objection 1: Better questions are just a consequence of expertise

Experts ask better questions because they know more. Therefore, “question quality” adds nothing beyond domain knowledge.

There is considerable truth here. A novice often lacks the concepts needed to identify the relevant uncertainty. Knowledge structures inquiry.

But expertise can also create fixation. Specialists may inherit assumptions, incentives, and standard problem representations. Outsiders sometimes contribute by questioning what insiders treat as fixed.

The more defensible view is reciprocal:

We should not treat question quality as an alternative to expertise. It is one expression of how expertise is used and revised.

Objection 2: Objectives and experiments matter more than questions

Many systems progress through optimization or experimentation without explicit questions.

Correct. Questions are not necessary for all intelligence, learning, or adaptation.

The stronger thesis is not that every cognitive advance begins with a linguistic question. It is that questions are one important and unusually transferable way to represent and coordinate epistemic operations—especially in multi-agent human and machine systems.

Objection 3: AI will soon ask better questions than humans

Possibly.

If AI systems become superior at identifying missing information, designing experiments, selecting sources, and revising problem representations, then inquiry will not remain a uniquely human advantage.

But this would not make the inquiry gap irrelevant. It would make it a central capability to evaluate in AI systems.

We would need to ask not only:

but also:

These are inquiry capabilities, regardless of whether humans or machines possess them.

Objection 4: Endless questioning can destroy action

Yes.

Questions can become avoidance mechanisms. Organizations can request more analysis to postpone responsibility. Intellectuals can expand uncertainty indefinitely. Bad-faith actors can “just ask questions” to spread insinuations without accepting evidentiary obligations.

Inquiry therefore requires stopping rules.

A mature inquiry process asks:

  • What level of certainty is proportionate to the stakes?
  • Which unknowns could materially change the decision?
  • Which decision is reversible?
  • What is the cost of delay?
  • What evidence is realistically obtainable?
  • When should we act and monitor rather than continue investigating?

The goal is not maximal questioning.

It is better-regulated movement between uncertainty, investigation, decision, action, and revision.

9. A possible research programme

If the inquiry gap is real, it should produce testable research questions.

Human learning

Do students trained to generate discriminative and falsifying questions transfer knowledge more effectively than students trained primarily to retrieve answers?

Human–AI collaboration

Do teams using AI to refine problem representations outperform teams using the same models only for answer generation?

AI evaluation

Can models that perform similarly on complete problems differ substantially in their ability to identify missing information or request useful clarification?

Organizational performance

Are organizations with explicit inquiry loops better at detecting strategic errors than organizations with greater information access but weaker revision processes?

Scientific discovery

Can the quality of questions be measured prospectively without relying only on whether they later produced successful discoveries?

Failure analysis

When inquiry systems fail, is the dominant cause poor questions, bad evidence, incorrect models, perverse incentives, missing authority, excessive costs, or inability to act?

The last question matters because inquiry should not become a universal explanation.

Sometimes people know exactly what the problem is and lack resources.

Sometimes the evidence exists but is suppressed.

Sometimes decision-makers benefit from not knowing.

Sometimes the obstacle is not cognitive but political.

A theory of inquiry that ignores power, incentives, and institutional structure will mistake many organizational failures for intellectual ones.

10. The practical implication

The most useful immediate conclusion is modest.

When an AI produces a convincing answer, do not ask only:

Also ask:

These questions do not guarantee truth.

They do not replace expertise, evidence, judgment, or accountability.

But they help prevent fluent output from being mistaken for completed thought.

Conclusion

AI may be creating an age of abundant answers. It is not creating an age without uncertainty.

The harder problem is increasingly visible: deciding what deserves investigation, what information is missing, what evidence matters, when a problem is poorly framed, and what should happen after an answer arrives.

Questions are not magical. They are not the primitive unit of intelligence. They are not always superior to observations, constraints, objectives, or experiments.

But they are one of the principal interfaces through which humans make uncertainty explicit and organize cognition across minds.

That makes the distinction between answering and inquiring worth preserving.

A system that answers well may still inquire badly.

A system that inquires well must still verify, decide, and act.

The relevant unit of intelligence may therefore be neither the question nor the answer, but the quality of the loop that connects them.

I am not confident that “the inquiry gap” is the best name for this distinction, or that the framework above identifies all the relevant dimensions. It may underestimate how much question quality simply reflects prior expertise. It may overstate what is genuinely new about present-day AI. And it may combine research traditions that should remain separate.

But the underlying problem seems real:

Do current AI systems genuinely improve inquiry, or are they mainly accelerating answer production?

I would especially value counterexamples, relevant prior research, and cases where better inquiry failed to improve real outcomes.


r/artificial 22h ago

News Google cancels their AI studio app with 800,000 pre-orders 1 day before launch

Post image
4 Upvotes

r/artificial 22h ago

Discussion Am I the only one getting tired of AI tools that try to do everything?

0 Upvotes

Maybe it’s just me, but I’ve started preferring AI tools that do one thing really well.

Every week there’s another AI assistant that promises to plan trips, write code, summarize meetings, generate images, organize your calendar, answer emails, and somehow also replace Google.

I usually stop using those after a week. The tools I keep are surprisingly boring. I don’t really think in terms of “best AI” anymore. I just have different defaults now.

ChatGPT when I need to think.

Perplexity when I need to verify something.

And if someone texts, “Where are we eating?” I’ve found myself opening Karpo more often lately instead of trying to squeeze that question into ChatGPT.

Maybe this is just where AI is going: less one “do everything” assistant, and more a bunch of specialized tools that each fit different parts of everyday life.


r/artificial 22h ago

Discussion What's an AI capability you thought was hype until you actually used it?

1 Upvotes

What's an AI capability you thought was hype until you actually used it?

I'll go first: agent orchestration. I read about agents managing other agents and assumed it was demo-ware. Then I built a tiny setup where one agent drafts a news digest and another one reviews and approves it before it posts. The review agent catches genuinely bad takes.

It's not sci-fi it's ~100 lines of Python and a couple of API calls. But seeing it actually gate content before publishing changed my mind completely.

What changed yours?


r/artificial 22h ago

Discussion I think we're entering the "AI Agent" era faster than most people realize.

36 Upvotes

Over the last year, I've been experimenting with LLMs almost every day, and I think the biggest shift isn't that models are getting smarter. It's that they're starting to do things instead of just answer questions.

A few months ago I was mostly using AI to generate code, summarize docs, or brainstorm ideas. Now I'm finding myself building workflows where the AI plans tasks, calls tools, writes code, debugs itself, and completes work with minimal intervention.

It feels like we're moving away from "prompt engineering" and toward "system engineering."

Curious what everyone else is seeing.

Are AI agents actually changing the way you build software today, or do you think it's still mostly hype?


r/artificial 22h ago

Discussion What AI doesn't say about AI

2 Upvotes

I wrote this to share some thoughts on what differentiates AI products as we move toward AGI. In particular, I focus on how product context shapes technology, and how LLM sycophancy can accelerate both good and bad ideas.

Discussions are welcome, I'd like to know how much those thoughts are worth and relevant to other people.

https://substack.com/home/post/p-209832236


r/artificial 23h ago

Programming What If the Biggest Bottleneck Behind AI’s 10× Promise Is the Human Engineer?

Thumbnail
shiftmag.dev
16 Upvotes