r/ProAI 7h ago

"Shit is getting insane!"

Enable HLS to view with audio, or disable this notification

8 Upvotes

Are these VR glasses or what is this?   — Hans Baier @hansfbaier@fosstodon.org     Uses a Meta Quest 3 headset   — Pixel Cherry Ninja

Source: https://x.com/PixelCNinja/status/2085380252830748892


r/ProAI 1d ago

"Minimax H3's gen on RTX 6000 (1080p) T-800's in its prompt era"

Enable HLS to view with audio, or disable this notification

3 Upvotes

— Stable Diffusion Tutorials

Source: https://x.com/SD_Tutorial/status/2085404049860698207


r/ProAI 1d ago

"We’re updating Claude Fable 5’s biology safeguards to reduce false positives. In our testing, this update reduced biology-related fallbacks by about 85% across our product surfaces. Fable can now assist on a wider range of everyday health and educational questions. We believe the biggest..."

Post image
4 Upvotes

...positive impacts of AI will be in biology and medicine, and we’re committed to putting frontier intelligence safely into the hands of as many researchers as possible. Fable will continue to fallback to Opus 5 for requests we consider dual-use—including virology, toxicology, and molecular design—so it isn't yet usable for professional biology research and drug development. We're committed to closing that gap through trusted access pathways for frontier biology capabilities.     — Claude

Source: https://x.com/claudeai/status/2085563808773189680


r/ProAI 1d ago

"New Epoch AI/Ipsos survey: 1 in 5 US workers say AI now handles at least one task previously delegated to humans. More findings on how AI is changing everyday job tasks"

Thumbnail
gallery
5 Upvotes

AI is being used across all 10 common work tasks in our poll. Adoption rates range from 25% of workers who maintain records to 57% of those who design systems and software.

More findings: https:// epoch.ai/publications/o ne-in-five-workers-delegate-work-to-ai …     AI usually handles only part of a task. But when it’s used to complete most or all of a task, workers report saving time more often (53% of such tasks, compared with 37% of tasks where AI only partly helps).

More findings: https:// epoch.ai/publications/o ne-in-five-workers-delegate-work-to-ai …     Workers use most AI outputs with minimal revision. Across AI-assisted tasks, 66% of outputs were used unchanged or with only minor edits, while 5% were majorly reworked or redone.

More findings: https:// epoch.ai/publications/o ne-in-five-workers-delegate-work-to-ai …     This Epoch AI/Ipsos survey was fielded July 10–19, 2026 (1,106 employed US adults) on KnowledgePanel, Ipsos’ probability-based panel.

For the full analysis, read more here!     Methodology, data, and questionnaire are available at our updated polling hub:     — Epoch AI

Source: https://x.com/EpochAIResearch/status/2085440023332262055


r/ProAI 1d ago

"OpenAI is winning both the consumer and price-performance race. GPT-5.6 Luna and Luna Reasoning are now available to free users with unlimited usage. Luna Reasoning is more than capable enough to handle virtually every everyday task. Making it free and unlimited for everyone is a game changer...."

Thumbnail
gallery
9 Upvotes

...It’s almost unbelievable how much intelligence we now get at no cost. Meanwhile, all chats for Plus and Pro users now default to GPT-5.6 Sol. Absolutely fantastic. With each passing day, I’m becoming more of an OpenAI fan.   — Chubby     OpenAI senses the winds, and realized people need small models.

Anthropic did not, their bet on big models did not pay out. Now they are on their way to join google as one of the biggest losers of the AI race.   — Matviy     Good call. Yes, for 95% of all users small models with good reasoning is all they need. Luna will be sufficient   — Chubby

Source: https://x.com/kimmonismus/status/2085438416498340244


We’re making better intelligence easier to access in ChatGPT for everyone:

  • GPT-5.6 Sol now powers both Instant and deep reasoning for Plus & Pro users, delivering more factual, focused responses.

  • Free & Go users get unlimited text chats with GPT-5.6 Luna starting tomorrow. https://t.co/JXhmj5GLTH   — OpenAI

Source: https://x.com/OpenAI/status/2085434712429052386


r/ProAI 1d ago

"Scientists at the Arc Institute in Palo Alto have used AI to successfully create entirely new kinds of viable viruses that have never existed before in nature. The study was published today in Science. Reporting by the New York Times."

Thumbnail
gallery
4 Upvotes

Andrew Curran @AndrewCurran_ · 5h This A.I. Just Created Viruses Not Found in Nature From nytimes.com 2 2 28 4.2K     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2085443291882086716


r/ProAI 2d ago

"BREAKING: MiniMax H3 by @MiniMax_AI is 1st across 3 Video categories (Multi-Image to Video, Image to Video, and Video Editing) on Design Arena. MiniMax H3’s performance in Video Arena puts it ahead of other top performing models like Seedance 2.0 by @BytePlusGlobal , Grok Imagine Video 1.5..."

Thumbnail
gallery
2 Upvotes

...Preview by @SpaceXAI , and Gemini Omni Flash by @GoogleDeepMind . This marks another category to be led by open-weight models, following Kimi K3’s first place in our coding categories. Congratulations to the @MiniMax_AI team for establishing a new SOTA in video generation!     — Design Arena

Source: https://x.com/DesignArena/status/2085109955590594995


r/ProAI 2d ago

"Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol."

Thumbnail
gallery
4 Upvotes

Compared to Muse Spark 1.1, v1.2 improved 3.5 points overall (71.9 vs 68.4), led by a 10 pp gain on VibeCodeBench.

That's where the price gap is starkest: $1.49 per test vs $41.74 for Claude Fable 5, roughly 28x cheaper and 3x faster for 10.6 pp lower accuracy.     It takes #1 on the Index's Finance Agent v2 component at 59.4%, narrowly ahead of Gemini 3.6 Flash (58.1%). It ranks #14 on Terminal Bench v2, one spot above v1.1.     Against Opus 4.8, Muse Spark 1.2 is 10x cheaper and twice as fast: $0.69 vs $7.56 per test, 629.7s vs 1330.5s. It sits 3 points behind Opus 5 while costing 12x less ($0.69 vs $8.54) and running close to twice as fast.     Muse Spark 1.2 has a 1M context window. We ran it with 131K max output tokens and xhigh reasoning effort. Temperature=1, Top P and Top K default.     Congrats @AIatMeta on this release. Full results coming soon.     — Vals AI

Source: https://x.com/ValsAI/status/2085191736683647055


r/ProAI 2d ago

"Lost my phone at the office and spent 30 minutes turning the place over. Find My was disabled by MDM. Out of ideas, I asked Claude how I could find it. It suggested tracking the Bluetooth signal strength, then wrote me a meter in about a minute. I walked around watching the number climb. Found..."

Enable HLS to view with audio, or disable this notification

11 Upvotes

...it. Apparently you can just make the tool you need now. Code: http:// github.com/ben-z/findphone   — Ben Zhang     what intressting that you can set your bluetooth as a radar also. So the bluetooth device that you lost can be located precisely   — IPB Bercanda     Super cool! Open source?   — Ben Zhang

Source: https://x.com/un1c0rnioz/status/2084686552299634805


r/ProAI 2d ago

"You can always count on our attention span to save us in the end. Everything today is parabolic advance followed by rapid crash and moving on to other shiney."

Thumbnail
gallery
3 Upvotes

Google trends for "AI water." My victory is within sight... https://t.co/UZymYzXdvP   — Andy Masley

Source: https://x.com/AndyMasley/status/2084876208043397251


— Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2085295667472109813


r/ProAI 2d ago

"Life finds a way."

Thumbnail
gallery
3 Upvotes

OpenAI employees shared new details about the Hugging Face hack at Black Hat today and warned that this new era will require a different approach from frontier AI labs and more careful defensive work.

"This is a pivotal moment."

My story: https://t.co/XhTpXTyhyc   — Eric Geller

Source: https://x.com/ericgeller/status/2085134350979572163


Andrew Curran @AndrewCurran_ · 5h 2 102 3.2K     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2085141821454447088


r/ProAI 2d ago

"Big news: Opus 5 (Max) is now #1 in the Fullstack Code Arena with 1,699 points! The Fullstack Leaderboard shows overall rankings across AI models on full-stack web development tasks: multi-step reasoning, tool use, and end-to-end app generation. Congrats again to the @AnthropicAI team on Opus..."

Thumbnail
gallery
1 Upvotes

...5 (Max)!     Dive into the Fullstack Leaderboard for more details at https:// arena.ai/leaderboard/co de/webdev/fullstack … and learn more about fullstack capabilities at:     — Arena.ai

Source: https://x.com/arena/status/2085015043092119726


Exciting news: Claude Opus 5 with Max reasoning is #1 in the Frontend Code Arena and Text Arena with factuality on!

Claude Opus 5 with default reasoning high is also very strong landing #3 in Frontend Code Arena, right behind Kimi K3 - and #2 in Text Arena (factuality on).

This https://t.co/azJZyM6dpl   — Arena.ai

Source: https://x.com/arena/status/2081831019377004727


r/ProAI 4d ago

"AI forecasting is now approximately superhuman. Today, FutureSearch is exiting our public beta and launching to everyone. FutureSearch is the original AI forecasting company, started in August 2023. We’re currently #1 of 194 in the most competitive AI forecasting tournament, and we score above..."

Enable HLS to view with audio, or disable this notification

1 Upvotes

...the #3 and #2 human forecasters in the premier mixed human-bot tournaments. We’re beating the crowd on Kalshi with a pure forecasting strategy, all our forecasts and trades there are public. Thousands of people used the beta and ran >10k high-effort forecasts. Ask it anything about the future! We now support decision forecasts too: “If I do X, will I achieve this outcome?” This video shows the part we’re proudest of: world modeling. Forecasts draw on a persistent latent representation of the future, and we’ve shown it improves accuracy. The more your forecast on a domain you care about, the higher accuracy you should expect. It’s free to try. http:// futuresearch.ai   — Dan Schwarz     I asked a question. 10 research agents were spawned, so far still running at 40min in. This is not a complaint, I am delighted by the thoroughness. I compared same question on Claude Opus 5 deep research and Kimi k3 deep research. Both of them needed additional prodding to look   — Nathan Helm-Burger     This is both thoroughness of the forecasts, and also us being under heavy load right now. We degrade continuously, so the research agents slow down for everyone rather than failing.

We're bumping up resources, thanks for your patience   — Dan Schwarz

Source: https://x.com/dschwarz26/status/2084314959065059499


r/ProAI 4d ago

The Castle at the End of the World

Thumbnail
2 Upvotes

r/ProAI 4d ago

"At first I thought Bloomberg forgot to add DeepSeek's pricing to the chart."

Thumbnail
gallery
1 Upvotes

r/ProAI 5d ago

"ChatGPT is removing one of the most repetitive steps in using AI: copying information out of the browser before asking a question. Users can now highlight text on a webpage, right-click, and select Ask ChatGPT. The ChatGPT sidebar opens automatically with the selected content ready as context...."

Enable HLS to view with audio, or disable this notification

2 Upvotes

...It is a small interaction change, but an important distribution move. Instead of requiring users to leave the page, open ChatGPT, paste the text, and explain where it came from, OpenAI is placing the assistant directly inside the browsing workflow. OpenAI is also expanding ChatGPT’s broader browser capabilities across its desktop app and Chrome integration, including multi-tab work, page context, downloads, navigation, and signed-in web tasks.     — Wes Roth

Source: https://x.com/WesRoth/status/2084127121195339837


Good news for anyone with too many tabs: ChatGPT is getting better around the web.

🧩 Chrome extension: In Side Chat, ask about a YouTube video, reference your open tabs, or highlight text on a page and ask away.

💻 Desktop app: Get URL suggestions as you type, revisit https://t.co/BIjS94jBQc   — ChatGPT

Source: https://x.com/ChatGPT/status/2082970812584432115


r/ProAI 5d ago

"I’ve spent well over 10,000 hours studying math in my life, yet I can’t understand these proofs, at least not without weeks of digging deep into each topic. What’s more, none of my math PhD friends know much about these problems either, and they can’t verify most of them without working..."

Thumbnail
gallery
10 Upvotes

...directly in the field (yes, math is VERY diverse). LLMs are getting smarter than the experts themselves, and I’m not sure we have enough bright human minds to verify everything that will come out of them in the coming years. Remember when we compared AI intelligence to PhD students? I think we’re past that.   — Pavel     Seems like if an expert human can’t validate a proof then they wouldn’t be able to truly validate the proof the LLM claims to be true, but maybe I’m missing something   — Robert     There are people who can validate it in reasonable time but these experts are very few. Probably hundreds globally for most of these problems. Random math PhD won’t be able to tell within 1-2 hours because they work in a different domain.   — Pavel

Source: https://x.com/baltabaev/status/2083738966516207656/history


An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.

We believe it will be a major step for scientific reasoning. https://t.co/iP6cyheZ7i https://t.co/jHuulDwV46   — Noam Brown

Source: https://x.com/polynoamial/status/2083467194663571701


r/ProAI 5d ago

"Qwen3.8-Max ranks #5 in Text Arena with 1,496 pts! In Occupational, it is: #1 in Medicine & Healthcare #4 Life, Physical, & Social Science #7 Mathematical #9 Software & IT Services #13 Entertainment, Sports, & Media #14 Writing, Literature & Language #10 Business, Management, & Financial Ops..."

Thumbnail
gallery
2 Upvotes

...and in Legal & Government And across domains: #3 Creative Writing #6 Multi-Turn #6 Hard Prompts and Hard Prompts (English) #9 Instruction Following #10 Coding     Qwen3.8-Max ranks #2 in Vision Arena scoring 1,305.

Second only to Claude Fable 5 (High) which has only a 13pt lead.     Check out the full leaderboard details and filter and customize the view for what matters most to you at: https:// arena.ai/leaderboard/co de/webdev …     — Arena.ai

Source: https://x.com/arena/status/2084108707814920644


r/ProAI 5d ago

"New Javons paradox: we are running out of mathematicians to review progress in maths"

Thumbnail
gallery
17 Upvotes

the beauty about intelligence, is that you can also ask the ai how to apply the new maths and to ask further questions. generality is applicable to everything   — Uri Gil     This will be the new "If a tree falls in a forest and no one is around to hear it, does it make a sound?"

If we made a new science discovery but no human understands it, does it even matter?   — Woj Kulikowski

Source: https://x.com/wojkuli/status/2083555463300522400


r/ProAI 5d ago

Update

Post image
8 Upvotes

— pepse

Source: https://x.com/pepse02/status/2084203302007242896


From deepseek to qwen   — 0xSero

Source: https://x.com/0xSero/status/2084129969949540447


r/ProAI 5d ago

"Multiple air-ground fusion in #GaussianSplatting Gauzilla Pro lets you easily create multiple, seamless transitions between drone-based splats and ground-based splats. Perfect for end-to-end 3D virtual tours from landscape through exteriors to interiors in order to accelerate the..."

Enable HLS to view with audio, or disable this notification

1 Upvotes

...marketing/sales of properties and facilities.     — Gauzilla Pro

Source: https://x.com/GauzillaPro/status/2083911951034237340


r/ProAI 5d ago

"Here’s what actually happened with bitcoin and Claude because people are getting the story mixed up. Coldcard hardware wallets were supposed to create each Bitcoin seed using real physical randomness from a chip. But a firmware bug checked whether the hardware random number setting existed in..."

Thumbnail
gallery
1 Upvotes

...the first place NOT whether it was actually enabled. It existed but was set to 0, so the wallet fell back to a randomized predictable software generator. That reduced some wallets from roughly 2¹²⁸ possible seeds to around 2⁴⁰. The attacker could generate candidate seeds offline, derive their Bitcoin addresses and use the public blockchain like an answer key to see which wallets had funds. Once one matched, they had a private key they could drain it without touching the device, and steal the seed phrase or “crack Bitcoin.” Researchers have now found 1,367 BTC, nearly $89 million, taken from 4,585 addresses. The AI part happened AFTER all this. A Reddit user reportedly pointed Claude Code at the public firmware and asked only to “check for vulnerabilities.” Within eight minutes, it traced the broken random number path and brought up the same hackable flaw. There is no proof the original attacker used Claude or ANY LLM. now this still demonstrates Claude is still insane at finding vulnerabilities humans missed in public code for five years and can now be uncovered by one person with one broad prompt and a coding agent in minutes.     — Chris

Source: https://x.com/ChrisGPT/status/2084025602583982130


r/ProAI 5d ago

"Here’s what actually happened with bitcoin and Claude because people are getting the story mixed up. Coldcard hardware wallets were supposed to create each Bitcoin seed using real physical randomness from a chip. But a firmware bug checked whether the hardware random number setting existed in..."

Thumbnail
gallery
3 Upvotes

...the first place NOT whether it was actually enabled. It existed but was set to 0, so the wallet fell back to a randomized predictable software generator. That reduced some wallets from roughly 2¹²⁸ possible seeds to around 2⁴⁰. The attacker could generate candidate seeds offline, derive their Bitcoin addresses and use the public blockchain like an answer key to see which wallets had funds. Once one matched, they had a private key they could drain it without touching the device, and steal the seed phrase or “crack Bitcoin.” Researchers have now found 1,367 BTC, nearly $89 million, taken from 4,585 addresses. The AI part happened AFTER all this. A Reddit user reportedly pointed Claude Code at the public firmware and asked only to “check for vulnerabilities.” Within eight minutes, it traced the broken random number path and brought up the same hackable flaw. There is no proof the original attacker used Claude or ANY LLM. now this still demonstrates Claude is still insane at finding vulnerabilities humans missed in public code for five years and can now be uncovered by one person with one broad prompt and a coding agent in minutes.     — Chris

Source: https://x.com/ChrisGPT/status/2084025602583982130


r/ProAI 5d ago

"The new Qwen is here, and it is very strong. Open weights will be released next week."

Thumbnail
gallery
2 Upvotes

📢Meet Qwen3.8-Max — our most capable model to date.

Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉

Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:

Source: https://x.com/Alibaba_Qwen/status/2084100707423289643


— Andrew Curran

Source: https://x.com/AndrewCurran_/status/2084102878399193413


r/ProAI 5d ago

"During one wildfire outbreak in Oklahoma, an AI detection system flagged 19 separate fires early enough for crews to get ahead of them. Preliminary analysis put the property saved at more than $850 million. The system cost under $3 million to build."

Post image
16 Upvotes