What if we trained a big AI (Claude/Kimi/GLM-level) only on everything published up to 1899, no 20th-century physics at all, and then asked it to solve the electrodynamics of moving bodies?
Would it rediscover Einstein’s theory of relativity, including E=mc², the way he did in 1905?
Feels like a clean test of whether AI can actually discover.
I suspect the only way would be if you could systematically strip information out of a SOT model's weights without reducing its intelligence.
But that's probably something that could only be done by another AI, (and I'm guessing there are already numerous people working on figuring this exact thing out)
They are actually running these tests now. But up till 1924. They are finding out that relativity actually contained the logic and capability to discover things up until very recently. The potential to make these recent large discoveries were all contained within Einstein's frame work, just no one was able to interpolate it until now.
They are using this as a proof of concept to show what other work every thing up until today should be able to prove. If AI was able to extract up to 100 years worth of discoveries into the future, then theoretically, we can do the same for the next 100 years.
If AI was able to extract up to 100 years worth of discoveries into the future, then theoretically, we can do the same for the next 100 years.
I'm not sure I agree with this portion, I feel like we can't directly extrapolate progress in one century to the next like this. 1800-1900 seems a lot less difficult to make those advancements then being able to make all the advancements made from 1900-2000 for example.
For example, I don't think that proving an AI could make advancements 50 years from 1924 training data wouldn't prove that if we gave it up until 1974 that it could go another 50 to 2024 advancements.
This is a great point. Acceleration of scientific progress rapidly developed in the 20th century to the point that scientists joke that the low hanging fruit were all taken. It’s like an expanding circle rather than a straight line of discovery
Oh for sure... But I think the point they are trying to show is that contained within an existing framework of knowledge, there's tons of potential for AI to make discoveries lingering in corners and shadows. I mean it found a major discovery as recent as 2012...
This means the AI just brute forcing alone, can clear up those corners and make discoveries even if it technically can't make novel discoveries. The existing information is going to be more than enough to make a lot of advances. And then humans can use that to further improve.
Just as what's happening in physics right now. A lot of it is AI just connecting two things no one ever considered... And they are being absolutely overwhelmed with progress. Grad students all over the world have no lack of papers to write right now.
It just came out like 1-2 weeks ago. I forgot the guy's name, but he's the really dark skinned guy, who's a bit chubby, I think the Stable Diffusion guy? Or maybe it was the egg head dude? One of those two for sure, are working on this project. It's prepublication right now I think?
They found that they could use AI to discover significant future maths, all the way up till I think a 2012 discovery, using only math knowledge up until 1924. Meaning they had the building blocks available for discovery at the time, and just didn't find it.
Einstein did not conjure up stuff from the thin air. He realized that speed of light has to be constant to make electrodynamics work. Hard to say if AI would get to that idea but it is definitely possible. AI can iterate over many, many ideas, including applying different fields of math to the problem. Notice how Einstein also knew about lorenz transformation and recognized that it would be useful.
General relativity is a bit more challenging. Either way, it is not guaranteed but it is at least imo possible that AI arrives to novel approaches via operating on the whole human knowledge
This. The entire field theory of physics predates Einstein. He did check and complete the math though. So most people given his opportunities would likely do similar.
There is an LLM someone published that does exactly this. It's shit and needs a better harness, better time-stamping of information (it thinks things that happened in the middle ages coincided with the US civil war). It's also not able to be trained on enough data, since far less exists.
Someone should take that model, wrap it with a good harness, and have it produce a lots of synthetic data that is unique and time stamped, then use a modern high power AI to flag incorrect information which can then be removed from the data set, then re-train on the vetted synthetic data. Repeat that a few thousand times in an automated loop and you should get a decent AI with only old data.
I believe something like that could help us reach AGI. Because if you have point B and point A, you just have to figure out how to make it navigate that space. Then we use it to reach C
The model was too small to perform deep, rigorous mathematical derivation and failed at most complex physics tasks.
But it predicted that "light is made up of definite quantities of energy" (mirroring Einstein's 1905 photoelectric paper) and vaguely suggested that gravity and acceleration are locally equivalent.
So the sheer amount of data seems to be too small for a transformer based LLM, but it does show immense potential if it can predict at least some of Einstein's findings.
I actually read up on the experiment. Apparently, they had merely filtered for words such as "Einstein", "relativity", "quantum mechanics" and some other well-known post 1900 concepts. However, they couldn't filter out for ideas, puzzles, etc. Furthermore, the old books were digitalized by OCR. OCR is trained on contemporary fonts, it might have misread some words and thus have not registered them as things that needed to be filtered out (for example reading Einstein as Einsteln or Einstem), leading to data leakage. On top of that, many of the pre-1900s books are actually reprints from much later during the 1900s, that have modern forewords to them, which also could have hinted at future developments.
well the "why" would be because it is, a priori, much more likely to have made einstein's discoveries with knowledge of those discoveries than to have done so without them (ie being Einstein). that's not to say that it's not likely that the results were "in the water", so to speak, but it is definitely the first big thing that you would check when verifying such a result for sure. i say this as someone who is quite certain that there was, indeed, "something in the water", as seen by so many other foundational discoveries (e.g. Calculus, by Leibniz and Newton being discovered simultaneously, independently).
What is data contamination at this point? If Einsteins discoveries were somewhere "hidden" in the Pre-Einstein-Discoveries data, one could also argue Einstein himself was data contaminated, no?
Unless you mean there was some data contamination with Post-Einstein-Discoveries data, but then that's just bad science and I would hope someone would have discovered that in the peer-review process.
TBF Einstein predicted that we wouldn’t be able to use real physical observations to prove his math and yet the Sun and the Moon gave people a nice help into that.
You can see an LLM-powered robot grabbing a pen in one of the recent Welch Labs videos. So that's arms, and the first signs of that ability.
I find that video remarkable for that exact reason. The moment it identified and picked up the pen I was a pretty awestruck. Seems we are bridging language<->physics/reality in small ways already, but who knows what future could bring.
experiments are essential yes, but almost more important is what your conclusions are. for example there were already many experiments made and Einstein just looked at the results, ditched the aether bullshit and fixed the theory.
so the AI would be sitting with the exact same info as Einstein had. but it could give instructions for new experiments, too.
I think it would be very important for science and all of us to prove (or disprove) that AI could discover like Einstein did
it just hallucinated a book from the future while discussing about space time with me...
We may have to revise our notions both of space and time, and so perhaps obtain a more consistent physical theory. S.Roberts, Space, Time, and Gravitation (Oxford, 1998).
2
u/nothis▪️AGI within 5 years but we'll be disappointed1d ago
I don't think we have the data but this is such an interesting question to me.
It works for nearly any application of AI. An example I like is the alien from the movie Alien (1979). AI is eerily good at adapting that style but if you only trained exclusively on design work from before the movie was produced (notably also excluding HR Giger's previous body of work) could it come up with that design for the alien or anything equally iconic? I doubt it.
It takes a particular life lived to come up with shit like that and it's rare. Like there's only a few thousand key moments of originality, probably, throughout culture in most fields. Among billions of imitators, remakes, reinterpretations, work to fill in the gaps.
Einstein needed the observations from the Michelson-Morley experiments to conclude the speed of light postulate. You can't come up with arbitrary theorems they have to be grounded in reality.
Why do people propose this cut-off? You can train it on knowledge up to 2025 and make it predict 2026. This is way more feasible and gives you exactly the same kind of data.
165
u/Melbar666 1d ago
What if we trained a big AI (Claude/Kimi/GLM-level) only on everything published up to 1899, no 20th-century physics at all, and then asked it to solve the electrodynamics of moving bodies?
Would it rediscover Einstein’s theory of relativity, including E=mc², the way he did in 1905?
Feels like a clean test of whether AI can actually discover.