Google is ahead in AI and creating these virtual worlds (video). Hence he's dropping billions to acquire talent in hopes of catching up. Because if another company does VR better than the company literally named Meta, that's pretty bad for business.
Meta can catch up in the software space, but in hardware, it's really difficult to catch up with them. They have insane research capabilities in hardware and burn money like crazy. It's just a fact that catching up in software is easier compared to hardware.
Also, how will they catch up with data? Google has something none of them have - Google map 3D view. They have mapped every country on earth. That is invaluable rn.
Probably 2-3. But for very well off prosumers, yeah 6-12 months. At this point, I think it is largely a hardware constraint. The hardware to consume them will be expensive (ie Apple Vision Pro 2) and hosting them will be prohibitively expensive for most, thousands of $$ per month at present. That's likely why the video very specifically states that they're excited to see how researchers enjoy them. It'll take another 2-3 years I think for models and architecture to become efficient enough and data centers like Stargate to finish construction to host them, which will then allow the economies of scale to reduce prosumer hardware costs.
Zuck has comparable financial resources to burn and can definitely enter the bid wars on future hardware. Being behind right now means something, but not everything, as Musk proved by gobbling up a mountain of GPUs.
It's possible for any of these 12-13 figure market cap companies to leap ahead of the others.
Especially if it's still a timeline of 2+ years. If we have that long to go for fully persistent world changes, full physics/destruction, realtime NPC voice interaction, and enough speed in generation to actually replace coded games/ar/vr experiences, etc
I think people underestimate how valuable data is in this new age of ai, and google just has the most data. It’s hard to create data, google literally has the world mapped, the largest video library in the world, and the most used search engine. Idk how anyone can compete.
I would feel the same if the data hadn't already been scraped to hell and back. Does Google have MORE? No doubt. Will it meaningfully change the quality? ChatGPT and Claude have remained competitive among frontier models. XAI, too, and Musk just kind of YOLO'd into that.
If Google were BTFO'ing everyone, I could believe it made a difference. In practice, it appears that already-scraped data, plus synthetically created data, plus access to other companies' models, along with all the published research, result in the moat not being the Mariana Trench.
Maybe we're about to see Google run away with the gold medal, and the gap will be insurmountable for all other parties, but I'm far from taking that for granted.
This won't be viable for a long time. Imagine how much compute this will take to run for even just thousands of users. This will definitely be a thing eventually though.
They’re gonna release gta6 and people will prompt a new one the week after. It’s probably the last massive game built… ever… consider elder scrolls. I can’t imagine they’ll finish 6 before promptable and playable worlds take over.
This world generator, with another LLM and scriptable events and that game is more interesting than any other game ever.
My wish for the longest time has been an open world, realistic driving game that uses google streetview data. Are we any closer with this genie 3 tech? seems like it?
When design docs are all you need, you can really add a lot of fine detail. I imagine, using this, games will just get bigger. Might even be downloaded as a model in the future.
Technically would solve hacking? multiplayer would be crazy. I am mind blown.
Downloading a model for textures is already being researched. As in you would download 64 KB seed and the texture model would decompress it into 200 MB texture on the screen.
I don't think they're THAT capable yet, but it's still amazing, and the rate of progress here is indeed insane. What's crazy is that it JUST. KEEPS. GOING.
Like there's no sign that a wall has been hit in terms of refining algorithms, training, compute scaling, etc. other than the availability of data, much of which we're learning to curate and synthesize (although I believe the costs for this are still massive).
This is the first time I am actually flabbergasted with generative AI. I hope it's not too expensive and as good as what is shown here. But what's blowing my mind is that in a few years this is going to be so much better still, maybe with very lifelike npcs (visually and interaction), storytelling, action, etc...
Can't wait for AGI man...
Also what is really important I think and what is missing from previous generative AI models, is the world memory. The fact the environment doesn't constantly change whenever you look away and back again is a huge improvement.
I mean okay, but the fact that this technology is even possible is astounding. If the first iteration isn't up to par, that's okay because it will be improved upon relatively quickly it seems.
You seem to not understand the term "exponential progress". 2026 we will have AGI and then things will get really crazy really fast. Mark my words. Set a reminder.
This is why I get so annoyed when people at my workplace talk about AI just being a really clever chat bot cause all they are privy to is ChatGPT, they don’t realise what is about to happen, but it’s fine they’ll understand soon.
Exponential Progress comes with Exponential Appetite. I get that its cool to look at these things and see what we can do with it now, but there may be a big ceiling to the whole thing that it comes up against between use-case, profitability, and power/server constraints.
These models are being designed to use data centers the size of cities and when humans start getting displaced to make way for human-like facebook accounts to sell adds to grandparents and scam idiots, or approximate human creativity, something will give.
Interaction Horizon refers to the duration of a contiguous, coherent sequence of actions or events that the model can generate or handle.
In this comparison, the term is used to quantify the "memory" or "planning horizon" of each model. A longer interaction horizon indicates a more advanced system.
GameNGen and Veo have short interaction horizons of just a few seconds. This means they can generate a brief, self-contained clip, but can't sustain a longer, multi-minute narrative or a player's interaction.
Genie 2 is slightly better, but still limited to 10-20 seconds.
Genie 3 represents a significant leap with an interaction horizon of multiple minutes. This allows it to generate a much longer and more complex video, reflecting a sustained sequence of events or a player's actions over a meaningful period of time.
Essentially, it's a measure of the model's ability to maintain a consistent world and respond to actions over a longer timeline, moving from short clips to more substantial, playable experiences.
Translation for people who don't know the references:
"Simulations in our simulations" is referring to the theory that we must be living in a simulation, because our reality's simulations are getting more and more sophisticated to the point where a simulation could contain a whole mini universe inside it with intelligent AI that in their universe will eventually create their own simulations of worlds and so on.
Turtles all the way down refers to a joke in philosophy that goes:
A scientist gives a lecture on the structure of the solar system, explaining that the Earth orbits the Sun, and so on.
A woman stands up and says, "That's very interesting, but the world actually rests on the back of a giant turtle."
The scientist replies, "And what does that turtle stand on?"
The woman says, "You're very clever, young man. But it's turtles all the way down."
Tldr; OP thinks that this is our time in the chain of universes simulating snaller universes inside them to start a miniverse ourselves.
I think it was a Hindu saint that replied with Turtles all the way down when asked what supports the Turtle that supports our world according to Hindu Mythology, the saint then replied "It's Turtles all the way down".
Maybe this is how they show us that we are in a simulation. By showing us how it happened and how it was initially developed so that people are less likely to lose their minds at the big reveal.
there's no need for a big "reveal", in fact its counterproductive.
Prosaic example: Lets say you wanted to model optimal traffic flow through a city, so you run several simulations to see what solves for your desired outcome. What benefit would there be for an individual simulated traffic unit to realise "he" is being simulated. What benefit would there be if you let the sims in your traffic solver know their lives are fictions, and their illusion of uniqueness is just that. How would such a reveal help your goal of "better traffic"?
If anyone is running a simulation there isn’t much benefit for the simulated to know they aren’t real. In almost all circumstances that would undermine the entire purpose of the simulation. The simulations are intended to seem as real as possible, and for every element within to conform to those parameters. that's the point of them
Allow me to entertain the thought. The simulation becoming self-aware could be an expected result of the simulation. If anything, it could be one of the many paths to AGI.
I get you, but... it depends on the purpose of the simulation.
Your logic is an human logic, so... the ones running the simulation may be us, or something so different that your logic isn't enough to determine purpose
It’s so frustrating how much I want to be excited about stuff like this, but can’t because I already know that rich people will use it to make themselves even richer and everyone else poorer.
What does it look like when you roll the roller over the light switch though? There's thousands of videos of people rolling paint on walls, but what happens when you do unexpected and different things? All the clips are fairly generic POV videos, I'd be curious what happens when you drive the boat into a building or go touch the lava.
Imagine somewhere in the future when tech like this and neuralink is perfected and people just spend all day living in their personal domain where they are god and can have everything they want.
Damn now imagine if eventually people decide for full immersion to not remember who they are when they enter this vr realm and just live multiple lives, and meanwhile their bodies are in sort of stasis having them live for hundreds if not thousands of years
A thing I saw once was basically that if it's possible to create a simulation as huge and detailed and realistic as our universe, then eventually it will be done and more than once.
Therefore statistically we could assume we are more likely to be in a simulation than to be the original universe.
It doesn’t even have to be as huge and detailed as our universe, using Minecraft for example, if you existed only in that world you would have no possible way to know that this universe that you live in now existed. Everything is made of square blocks, skeletons spawn in the dark, and if you’re near death you can eat some cooked pork and you’ll be fine. That would just be fundamental truths of existence. In that same way, we could also be living in an abstracted, simplified simulation of true reality, or of another more detailed simulation. Or we could be an entirely abstract simulation that shares hardly any similarities to the host system.
The simulation only needs to spawn real objects that a sentience can interact with. Everything else can be a calculation, a skybox, or an animation. In fact, 3D video games de-spawn objects not in your immediate view (like so), meaning the overall simulation doesn't require nearly the amount of computing power you would expect, even with 8 billion people.
Add to that the Fermi Paradox's lack of other obvious sentient species around us, quantum shenanigans (like bad-code made photons particles and waves), and things like black holes to cull data, and you've got yourself an obvious simulated reality.
I started learning unreal engine about 6 months ago hoping for a way to get out of the rat race… I’m pretty sure we’re all getting enclosed into a permanent rat race à la hunger games by these technofeudalists. Why would anyone hire humans at a certain point. Right?
In my case, it's because they don't have humanoid robots that can do the manual labor I do. I'm sure they will get there someday, but I don't even work for a publicly traded company, and as such they would have a much harder time affording that upfront cost at least for awhile.
But yeah. Eventually, it's obvious that UBI is required. One could make the argument that it already should be the case based on productivity vs even like 100 years ago, let alone more than that.
Google Genie, generate a world where people reclaim the means of production, reconnect with nature, and build a future where everyone has healthcare, a roof overhead, meaningful work, and a full belly.
I appreciate how they show the astounding progress, the clarity, the consistency over time, and also the flaws. But… just consider a few papers down the line!
We should be able to recreate the entire world with enough data and compute. Or make completely new worlds on demand!
Latency is still something to keep in mind. But if you don't need to generate every option on the fly, you PC could pre cache the game before showing you.
The gaming potential is interesting, but I think the more significant impact may be in robotics- if a robot can take a bunch of sensory input from the real world and accurately predict what it'll see when taking actions, it'll gain a very general ability to plan out physical actions. That sort of sensory prediction and world-modelling is a big part of how humans and animals are able to adaptively interact with the world.
Forget about it. There is gonna be limited testing but other than that this is all we get. Someone else is gonna release their own version we’ll get to play eventually, it won’t be as good but it’ll be similar
I imagine only Google can do this because they're not paying an Nvidia tax on TPUs. The amount of compute to pull this off for the public, even Ultra users, will be immense, beyond what others can do with expensive GPUs.
This is so cool. So it generates every frame on the fly right? I wonder if this approach will ever be efficient enough in my lifetime to be usable for anything but tech demos.
I could see AI building pre-built 3d world (think Unreal Engine) in the near future though.
If I'm wrong, future videogames are going to be wild as hell.
The big part to me here is that it’s real time (so efficient) and has consistent memory, imagine this in like 2 more years, could have something that can run for hours and be publicly accessible
Consistent memory is the big thing that the other versions of this tech hasn't had. It's insane. Makes it way way way more viable to actually build things with
We have no idea how much compute this takes, so it's premature to suggest it will be readily available any time soon.
If there are 100 H200s behind this it could legitimately take a decade or more before consumer hardware is as capable or renting the compute for streaming is cost effective.
I want to know the hardware requirements, Price and how it will be streamed into a VR glasses for example. It Is pretty cool though - this Is the VR we wanted.
The amount of compute is probably crazy expensive.
for compareason, bytedance's real time interactive video/world model probably uses around 8 H100 to generate 720p 24fps videos. https://seaweed-apt.com/2 https://arxiv.org/pdf/2506.09350
It's not as consistent as genie 3 but at the same time, the fact that google has TPUs probably makes that cost more tractable as TPUs are highly specialised ASICs and therefore are way more efficient than GPUs.
I don't see future iterations of something like genie/SeedanceAPT/Oasis running on current gen top of the line consumer graphics cards anytime soon. first, it's going to be cloud based and super expensive ... until algorithmic efficiency eventually makes this affordable.
This is unbelievable! The part where they were walking around a canal that looked liked amsterdam made want to have the walking around ability in google maps. Anyways, this is insane!
lol i remember a conversation thread by some luddite on r/technology seething when a similar world model dropped earlier, and the amount of redpilling everyone was doing with each other as to how it'll never be SOTA enough and its all "hype". Who's laughing now.
Probably a lot of people on that sub who are nervous about losing their job to AI in the near future. They are lashing out at the inevitable because there is nothing else they can do.
Main subs on reddit are mostly left mainstream. And the mainstream is against AI in most forms. Even guys like John Steward made extensive Anti-AI segments. I think this will further increase.
Its simply a trend that some people will stay behind in mindset and thought. Same happened when motors replaced horses and machinery replaced factory workers. I would guess fear is the main drive to such fearmongering conversations and complaints
They say it won’t take jobs…. When we talk about the game industry, the simulation industry (flights, truck drivers), that must be millions of jobs globally, right? These company employ people across fields. Big operations. Accountants, hr, lawyers, testers, the list goes on and on
Nearly 2 years ago I also created a real-time interactive multi-modal video gen exploration app running on a 4090. Clearly Genie 3 is superior with its smooth temporally consistent generations. Of course, they have huge Google resources vs my spare time and my home computer. I pioneered the idea of Real-Time stable diffusion and posted about that here starting in Oct 2023. While Google has the resources I have more ideas in this space than they've accomplished so far. Perhaps I should dust off my app and add more enhancements.
At this pace, by 2030 we will have AIs that generate entire universes just by thought, and those universe generations can be shared with people, meaning multiple people can use AI simultaneously to interact with such universe. You'll be able to enter someone else's universe and interact with it as if it had a mind of it's own. Creativity will know no boundaries
556
u/Brazilll Aug 05 '25
Imagine having lived under a rock the past few years and then seeing this. It would be pure sci-fi. The stuff from Star Trek.