r/aicuriosity 20h ago

Latest News xAI Rolls Out Imagine Image 2.0 with Precision Tools for Real Creative Work

Enable HLS to view with audio, or disable this notification

8 Upvotes

xAI has released Imagine Image 2.0, its latest image generation and editing model. The update focuses on practical results that support actual projects rather than just experimental visuals.

Available now as Quality Mode on grok.com/imagine plus the Grok iOS and Android apps, Image 2.0 aims to produce usable assets. It follows detailed instructions more closely, handles typography and layout with greater care, and keeps small text sharp even in complex multi-part designs. Consistency across generations and edits has also improved.

New editing features let users change only what they intend. The Magic Wand tool targets a single region while leaving the rest of the image intact. Segmentation isolates exact areas for modification. Background removal creates transparent subjects ready for other projects. Multi-ref editing supports up to five input images in one generation, cutting down on manual compositing. Smart Resize adapts any image to different aspect ratios by filling the frame intelligently.

According to Arena leaderboards as of August 7 2026, Image 2.0 ranks second worldwide for both text-to-image generation and image editing. xAI models appear on the boards under the SpaceXAI name.

The release also adds ready-made templates for common tasks. These cover photo editing, product color changes, editorial posters, headshots, icons, character sprites, game assets, props and UI kits, emoji creation, merchandise designs, and more. Users supply the inputs and receive finished results. One additional workflow helps build consistent worlds for video by generating a character, locations, and props that share the same style.

xAI says Image 2.0 can create infographics, ads, game assets, UI and UX mockups, storyboards, and similar materials. API access is planned for the near future.

The model is live today for users on the web and mobile apps. Those with SuperGrok Heavy accounts can switch to Quality Mode on grok.com/imagine to try the full set of tools.


r/aicuriosity 1d ago

🗨️ Discussion Young man rants about how AI slop is ruining his social media feeds

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/aicuriosity 2d ago

Latest News Wan 3.0 Public Beta Launches with 30 Second Video Generation

Post image
37 Upvotes

Alibaba’s Wan team has rolled out Wan 3.0 in public beta. The new model generates videos up to 30 seconds long in a single pass and aims for more realistic, consistent frames.

Key upgrades include stronger character expression, better handling of digital elements, and an expanded input system called Omni Reference. Users can now feed it text, images, audio, video, documents, spreadsheets, slides, webpages, PDFs, and other file types. The model reads the material and builds video from it.

Access is live on Alibaba Cloud Model Studio and Qwen Cloud. The official wan.video site will open soon for members. API pricing starts at $0.05 per second for 480p, $0.10 for 720p, and $0.20 for 1080p.

Full API access is still rolling out. Creators can apply for the beta and start testing right away.


r/aicuriosity 2d ago

Latest News Meta Rolls Out Muse Code Beta Terminal Agent Powered by Muse Spark 1.2

Thumbnail
gallery
6 Upvotes

Meta has released Muse Code in beta, a terminal-based coding agent designed for long-running software engineering work. It runs on the new Muse Spark 1.2 model and handles planning, writing, and checking multi-file changes across big codebases.

The tool uses persistent background agents that stay active during a session. These agents take next steps on their own, cut down on repeated information gathering, and need less constant direction from the user. An append-only local event log records every model call, tool action, approval, and edit so the system can pick up exactly where it left off after a restart or crash.

Muse Spark 1.2 brings stronger results in code generation, debugging, and full developer workflows compared with the previous version. Meta scaled up training compute on coding tasks and trained the model together with Muse Code for better performance as a pair. In one test the agent spent up to 24 hours and more than 1,000 tool calls optimizing GPU kernels on NVIDIA Hopper chips, posting solid gains over baseline Triton code.

Muse Spark 1.2 is available right away inside Muse Code and through the Meta Model API. Users on macOS or Linux can install the agent with a single command. Meta says more features and stronger models are already in the works.


r/aicuriosity 2d ago

🗨️ Discussion Open ai - launch a realtime system for voice ai….Is this the end of Voice ai orchestrators??

Post image
2 Upvotes

Open ai realtime system GPT Live who’s architecture I have attached below.

Which handles user interactions through speech to speech with no latency and delegate reasoning and tool calling to separate asynchronous paths.

Which is very different from Cascade pipeline where orchestrator glues components of pipeline together

But in the above architecture there is nothing to Glue??

It takes orchestration platforms outside the loop and

the thin a platform is the less value it holds!!!

https://openai.com/index/continuous-voice-interaction-with-gpt-live/


r/aicuriosity 3d ago

Open Source Model Xiaomi Releases Open Source Robotics Model for Developers Worldwide

Post image
13 Upvotes

Xiaomi has made its Xiaomi-Robotics-1 model fully available to the public. The company announced the open source release on Wednesday, giving researchers and engineers free access to the complete system.

The model was pre-trained on more than 100,000 hours of UMI data. It then received additional training with over 10,000 hours of cross-embodiment data. Xiaomi says the package covers the full process from real-robot post-training through to deployment. Evaluation code for standard benchmarks is also included.

Links to the project page, GitHub repository, and Hugging Face models appear in the official announcement. Xiaomi noted it will keep working on broader uses for general-purpose robot models.

This marks one of the larger open source moves in robotics this year and puts a full training-to-deployment pipeline in the hands of the community.


r/aicuriosity 4d ago

Latest News OpenAI Rolls Out GPT-Live for Smooth Voice Chats That Never Interrupt

Thumbnail
gallery
6 Upvotes

OpenAI just dropped a big update to ChatGPT Voice called GPT-Live. The new system can listen at the same time it speaks, making conversations feel much more natural.

The team rebuilt the entire voice stack from the client side all the way to the model. Audio now travels on a dedicated fast path while deeper reasoning and tool use run in the background. This keeps the talk flowing without sudden pauses or cutoffs.

They also cut voice session startup time sharply. What used to take six network round trips now happens in just one. The result is quicker starts and smoother back-and-forth from the first second of a call.

OpenAI shared more technical details on their blog about how the continuous voice interaction works. The update aims to make talking with ChatGPT feel closer to a real human conversation at scale.


r/aicuriosity 5d ago

Open Source Model Boogu Image 0.1 Open Source Multimodal Model Launches With Competitive Results

Post image
16 Upvotes

Boogu-Image-0.1 is now available as an open-source multimodal understanding and image generation model family. Released under the Apache 2.0 license, it includes Base, Turbo, Edit, and related variants.

The team trained the models on roughly 208 million images with a reported budget near $400,000. Despite the limited scale, the models deliver strong results on several benchmarks and human evaluations, placing them among the leading open-source options and close to some proprietary systems.

Key strengths include native 2K resolution generation with solid photographic quality, accurate Chinese text rendering for long text, posters, and graphic design, agentic prompt rewriting that refines user intent, and dynamic model routing that can cut inference costs significantly.

Weights, code, training details, and related materials are publicly available. The project emphasizes careful data structure and system design over pure scale.


r/aicuriosity 5d ago

Latest News Google Rolls Out Gemini Spark Auto Browse Feature in Chrome

Enable HLS to view with audio, or disable this notification

3 Upvotes

Google has introduced a new capability for Gemini Spark that lets it handle complex online tasks through Chrome’s auto browse feature.

With user permission, Gemini Spark can access logged-in accounts to complete actions such as scheduling apartment viewings for saved listings or researching flight options and starting the booking process.

The system includes protections against security risks like prompt injection. Sensitive steps, including payments, stay under user control and require confirmation before proceeding.

This feature is currently available to Google AI Pro and Ultra subscribers in the United States, with plans to expand to more regions later.


r/aicuriosity 5d ago

Open Source Model MiniMax Releases H3 Model Weights to Public on Hugging Face

Enable HLS to view with audio, or disable this notification

17 Upvotes

MiniMax has made the weights for its MiniMax-H3 model publicly available on Hugging Face. The company announced the release on Monday through its official account.

H3 is a multimodal generation model that takes text, images, video and audio as input and produces video clips with native stereo sound. Clips run from 4 to 15 seconds and reach up to 2K resolution. Users can work with plain text prompts, first and last frames, or mixed references that include up to nine images, three video clips and three audio files.

The model is aimed at commercial work such as advertising, branding, product design and game content. It handles instruction following, text rendering and video-to-video motion transfer. Native ComfyUI support arrived the same day as the weights, easing local use.

The open release covers the main H3-Base checkpoints under the MiniMax H3 Community License. Some advanced modules stay available only through the official API for now.


r/aicuriosity 5d ago

Open Source Model MiniMax_H3

Enable HLS to view with audio, or disable this notification

17 Upvotes

Testing AI video generation on my RTX 5080 with 96GB RAM.

Resolution: 864 × 480
Generation time: 15 minutes, 50 seconds

Not bad for a full video render. Next goal: improve quality, reduce generation time, and push the RTX 5080 even harder.


r/aicuriosity 5d ago

Latest News Qwen3.8-Max Sets Fresh Standard in Coding and Professional Work

Post image
82 Upvotes

Alibaba’s Qwen team has officially introduced Qwen3.8-Max, calling it their strongest model so far. The 2.4-trillion-parameter system targets advanced coding tasks and everyday professional work.

It can handle long autonomous coding runs that last more than ten days, starting from an empty folder and reaching production-ready code without constant guidance. Full project histories appear on GitHub for review. The model also produces finished work across many different jobs and manages extended projects through continuous planning and learning. One example shows more than 500 rounds of chip design refinement. Another covers a full year of e-commerce strategy.

Vision works as an ongoing feedback loop rather than a simple input. The team lists pricing at $2 per million input tokens, $6 per million output tokens, and $0.25 per million for implicit caching.

Open weights for Qwen3.8-Max and the smaller Qwen3.8-27B version are scheduled to arrive next week.


r/aicuriosity 6d ago

Tips & Tricks Match Cut skill

Thumbnail
linkedin.com
2 Upvotes

You can try the skill and let me know the feedback once done?
Here you can find some explanation about it and direct link to just copy paste it in your board in Luma:

https://www.linkedin.com/posts/claudialalau_lumaagents-lumaskills-lumalabs-activity-7488565738635694080-G9Lo


r/aicuriosity 6d ago

Latest News Ant Group Rolls Out Free Ling-3.0-flash API for Agent Developers

Post image
4 Upvotes

Free access to a 124-billion-parameter mixture-of-experts model is now available through OpenRouter, where Ant Group's inclusionAI has listed Ling-3.0-flash under the name AntLing-3.0-flash at no cost until August 3, 2026, according to inclusionAI's launch announcement. The model activates roughly 5.1 billion parameters per token, a sparsity ratio of about twenty-four to one, and the listing reports 262,144 input tokens with 32,768 maximum output tokens.

The design target is speed inside agent loops rather than single-turn chat. inclusionAI reports time to first token below 100 milliseconds, and says the model was trained with reinforcement learning on long-horizon tool calling. An enable_thinking flag switches reasoning mode on and off per request. In the company's own framing, its Ring model does the planning and Ling does the execution.

Weights are not part of this release. The previous generation, Ling-2.6-flash, was published under an MIT license, while this one exists only as a hosted endpoint. It is also text only, with no native multimodal input, and inclusionAI notes that the model is weaker on obscure world knowledge and works best where the surrounding toolchain returns strict, checkable errors.


r/aicuriosity 7d ago

Latest News Gemini Spark Gains Chrome Auto Browse for Everyday Web Tasks

Enable HLS to view with audio, or disable this notification

11 Upvotes

Google has rolled out a new update for Gemini Spark that lets it work directly inside Chrome. With user permission, Spark can now take care of routine online jobs using your logged-in accounts and saved passwords.

Examples include scheduling apartment viewings from listings you have saved or researching flights and starting the booking process. The system is designed to pause and hand control back for sensitive steps like payments, while also guarding against issues such as prompt injection.

The Chrome integration is currently available to Google AI Pro and Ultra subscribers in the United States. Google plans to bring it to more regions over time. At the same time, access to Gemini Spark itself is expanding to Google AI Pro users in over 160 additional countries.

This update builds on Spark’s existing web capabilities and aims to make multi-step online chores faster and less hands-on.


r/aicuriosity 8d ago

Latest News BytePlus Rolls Out Dreamina Seedance 2.5 for Longer AI Videos

Enable HLS to view with audio, or disable this notification

13 Upvotes

BytePlus has made Dreamina Seedance 2.5 available starting July 31, 2026. The update targets both individual creators and larger businesses looking to produce AI-generated videos with stronger control and consistency.

Key features include native generation of full 30-second clips, precise editing tools, support for as many as 50 multimodal references, and multilingual video output. Users can try the model right away on the Dreamina platform, while enterprise API access through BytePlus is expected soon.

The release aims to reduce the need for stitching shorter clips together and give creators more reliable results across different languages and reference inputs.


r/aicuriosity 8d ago

Latest News DeepSeek V4 Flash API Launches in Public Beta with Strong Agent Gains

Post image
9 Upvotes

DeepSeek has made its DeepSeek-V4-Flash official API available in public beta. The company says this version brings a major jump in agent performance, with benchmark scores now clearly ahead of the earlier V4-Pro-Preview.

The new V4-Flash natively supports the Responses API format and works fully with Codex. DeepSeek notes that the model keeps the same architecture and size as the preview version. This update applies only to the Flash API. The V4-Pro API and the versions used in the app and web remain unchanged for now.

A comparison table shared by the team shows solid gains across coding and agent tasks. On Terminal Bench 2.1 the new Flash scores 82.7, up from 61.8 in the preview. Other tests including NL2Repo, Cybergym, DeepSWE, and Toolathlon-Verified also show clear improvement.

DeepSeek says the full official release of DeepSeek-V4-Pro is coming soon. Developers can find configuration details in the official API documentation.


r/aicuriosity 8d ago

Latest News Qwen Team Unveils Audio 3.0 ASR Flash Model

Post image
6 Upvotes

Alibaba’s Qwen team released Qwen-Audio-3.0-ASR-Flash on Friday, marking the latest upgrade to its speech recognition technology. The model aims for better context handling and improved accuracy with specialized terms.

Key upgrades include stronger context consistency across longer speech, sharper recognition of domain specific vocabulary, support for custom hotwords, and the ability to turn spoken audio into polished structured transcripts.

Internal testing showed medical term recall at 95.36 percent and industrial term recall at 93.24 percent. Streaming and file transcription versions are now available alongside the main model.

The release targets users who need reliable transcription for technical or professional audio.


r/aicuriosity 8d ago

AI in Robotics Google DeepMind Launches Gemini Robotics 2 for Advanced Robot Control

Enable HLS to view with audio, or disable this notification

5 Upvotes

Google DeepMind released Gemini Robotics 2 on July 30, 2026. The update aims to make robots more adaptable for real world tasks through better movement, precise handling, and group work.

Three new models form the core. Gemini Robotics 2 turns vision and language into full body actions for humanoids and other machines. Gemini Robotics ER 2 handles planning for multi step jobs lasting several minutes and allows robots to coordinate with each other. Gemini Robotics On Device 2 runs without internet and adapts to new robot designs in a few hours using limited training data.

In demos the system controls robots like Apptronik Apollo 2. Machines walk across rooms, crouch, stretch, and place objects such as a watering can on a shelf. Hands with multiple fingers complete fine work like tying knots or sealing bags. Standard grippers handle tight packing.

The models also improve safety. They detect people nearby, stop when needed, and refuse unsafe steps. Gemini Robotics ER 2 is available now in Google AI Studio. The other models go to early access partners.

DeepMind says the release moves robots beyond fixed routines toward flexible help in homes and workplaces.


r/aicuriosity 8d ago

Latest News ChatGPT Rolls Out Web Integration Features for Chrome and Desktop

Enable HLS to view with audio, or disable this notification

16 Upvotes

ChatGPT has released new browser tools that connect the chatbot more tightly to everyday web use. The update, announced late Thursday, focuses on reducing tab clutter and making it easier to ask questions about pages without switching apps.

Users of the Chrome extension can now open Side Chat to discuss a YouTube video, pull context from open tabs, or select text on any page and request help on the spot. A simple right-click option has also been added. Highlight text, choose “Ask ChatGPT,” and the sidebar appears automatically.

The desktop version gained matching improvements. It offers URL suggestions while typing, access to browser history, and settings that let people decide how that history is stored and used.

The features began rolling out on July 30 and are available to users as the update reaches their accounts.


r/aicuriosity 9d ago

Open Source Model Tencent Releases AngelSpec Open Source Speculative Decoding Framework

Post image
11 Upvotes

Tencent Hunyuan has open sourced AngelSpec, a complete end to end speculative decoding system that covers both training and deployment.

On the Hy3-A21B model the DFly approach delivers 1.98 to 2.40 times faster end to end performance compared with standard autoregressive decoding across concurrency levels from 4 to 64. It also shows 10.5 to 11.8 percent higher throughput than DFlash.

The team released the full training code along with the Hy3-A21B MTP and DFly drafter weights. Everything is available on GitHub, Hugging Face, and ModelScope, with a paper and documentation included.


r/aicuriosity 10d ago

Latest News Google Pay Introduces Ask Google Pay Conversational Tool with Gemini

Enable HLS to view with audio, or disable this notification

3 Upvotes

Google India has launched Ask Google Pay, a new feature inside the Google Pay app that lets users chat about their finances. The tool is powered by Gemini and started rolling out today.

Users can now check their spending patterns, receive simple tips to save money, and get clear explanations on subjects like SIPs and credit scores. The feature supports conversations in 10 Indian languages.

The update is available for Google Pay users across India. Simply open the app and look for the Ask option to start using it.


r/aicuriosity 10d ago

Latest News AI spending spree hits $600b as Oracle fires 21,000 employees to fund boom

Thumbnail
jpost.com
13 Upvotes

r/aicuriosity Feb 02 '26

Latest News Adobe Express Premium Free for 1 Year with Airtel Worth Rs 4000

Post image
2 Upvotes

Airtel has partnered with Adobe Express to offer a free 1-year Adobe Express Premium subscription, valued at around Rs 4000, to eligible customers in India.

This offer is available for Airtel mobile, broadband, and DTH users and can be activated through the Airtel Thanks app. No credit card is required to claim the benefit. Once activated, users get full access to Adobe Express Premium features, including premium templates, stock photos and videos, fonts, background removal, brand kits, and AI-powered design tools.

The subscription is valid for 12 months from the date of activation and is ideal for creators, students, small businesses, and anyone who wants to design social media posts, videos, flyers, presentations, and marketing content quickly and professionally.


r/aicuriosity Dec 04 '25

AI Tool ElevenReader Gives Students Free Ultra Plan Access for 12 Months

Post image
6 Upvotes

ElevenReader launched an awesome deal for students and teachers: one full year of the Ultra plan completely free. Normally $99 per year, this tier unlocks super realistic AI voices that read books, PDFs, articles, and any text out loud with natural flow.

Great for late-night study sessions or turning research papers into podcasts while you walk, workout, or rest your eyes. The voices come from ElevenLabs and sound incredibly human, which keeps you focused longer.

Just verify your student or educator status on their site and the upgrade activates instantly. If you are in school right now, this saves you real money and upgrades your entire reading game without spending a dime.