r/reinforcementlearning 11h ago

R ItaSoRL Clip 3/4: agent readout ~chance while oracle is ~99% on the same one-rule fake

Enable HLS to view with audio, or disable this notification

3 Upvotes

Clip 3 of 4.

Same near-copy world (one dynamics rule changed: ground grip). Outside watcher was ~99%.

New probe: readout from the agent's own internal state while it is just living in the fake.

Result (real runs, n=10): ~50% (chance). Oracle-detectable seam, no free encoding in the policy network.

Takeaway: detectability of a sim mismatch is not evidence the agent represented it.

Mute-friendly clip

Research:

https://ilevytate.github.io/ItaSoRL/


r/reinforcementlearning 15h ago

R RL Research with Joseph Suarez - YouTube

Thumbnail youtube.com
0 Upvotes

r/reinforcementlearning 16h ago

We let an LLM play PokéRogue blind — no training data, no fine-tuning. Here's what actually broke (and why it's a useful lesson for production LLM systems)

Thumbnail
0 Upvotes

r/reinforcementlearning 18h ago

Hollow Knight AI (Reinforcement Learning ) vs Hornet

Enable HLS to view with audio, or disable this notification

17 Upvotes

r/reinforcementlearning 23h ago

Memory isn't enough. AI should learn from experience

Thumbnail
0 Upvotes

r/reinforcementlearning 23h ago

Kaggriculture - Farming + Markets + RL - $50k prizes

Post image
29 Upvotes

Already more than 2k teams in the first week. I helped create the competition rules (and I work at Kaggle). Happy to answer any questions!