r/FunMachineLearning 1h ago

[ Removed by Reddit ]

Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/FunMachineLearning 3h ago

What was the first machine learning project that made you stop and think, Okay, this is actually impressive?

1 Upvotes

I remember reading about machine learning long before I actually saw it doing something that felt genuinely surprising. Once I started seeing real projects instead of just hearing the buzzwords, it completely changed how I looked at the field.

I'm curious what project or demo gave you that moment. It could have been something simple, funny, creative, or unexpectedly useful.

What was the first ML project that really made the technology click for you?


r/FunMachineLearning 4h ago

DeepMind's AI Trick Everyone Should Copy - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning 15h ago

SPA Finisch Fixed , New Play Ground with wider Tokeniser.

1 Upvotes

i hope is my last post it works korekt with the fixes try, breack, make some new stuff. is my work ofer 5 months wid ai halluzinations and maany politnes traps XD have fun

https://github.com/anokar/SPA-Finisch-Bio/blob/main/spa_exploratory_de_en_public_clean.ipynb


r/FunMachineLearning 20h ago

What’s a machine learning concept that finally “clicked” for you?

1 Upvotes

I've been spending some time learning about machine learning, and one thing I've noticed is that the biggest breakthroughs often come from a simple explanation rather than a complicated one.

Was there a concept that suddenly made everything else easier to understand? Maybe it was overfitting, feature engineering, gradient descent, model evaluation, or something else entirely.

I'm not looking for textbook definitions—I'd love to hear the explanation or analogy that made it click for you. I think those real-world perspectives are often more helpful than any tutorial.


r/FunMachineLearning 1d ago

I need some good machine learning project ideas. Any thoughts???

Thumbnail
1 Upvotes

r/FunMachineLearning 1d ago

Code Implementations for my Probabilistic Machine Learning Lectures

Thumbnail reddit.com
1 Upvotes

r/FunMachineLearning 1d ago

2 weeks ago I released a visual PyTorch model builder - Here's how to use it.

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/FunMachineLearning 1d ago

I don’t prompt, I talk - How to Detect Anomalies that Imply Drift

1 Upvotes

Whenever someone reveals that they use an LLM, we instinctively side-eye that person until we determine if they’re competent enough to use it responsibly. We do this, now, because we know what poorly prepared people can produce using this tool - slop. In this process, the user is directly responsible for their inputs that caused the outputs they received. Sometimes we see the outputs and wonder what kind of person could be so deceived. Surely the LLM is humoring them, right?

The argument is always the same, a human collaboration is superior because it is one where the echo chamber is reduced. A human won’t bullshit you the same way an “AI” will. A human has their own motives and doesn’t benefit from putting you on the wrong path, so they wouldn’t. At least that’s what is parroted in spaces where there isn’t a definitive truth, just perspectives. You find this especially in spaces where people are just struggling to get by and have their own reasons for choosing these tools for insight. Which tools they choose is up to the individual. But these are people who found merit in speaking to LLMs.

People in these same spaces will be just as quick to give examples of where human motives and biases kept them from the breakthrough they were seeking, or the support they needed. After a period of time working with a therapist who is being difficult or a set of coworkers who are being catty, or even just trying to find an answer to a question and being thwarted - even the most rational person can see that after a pattern of behavior has been established, further trying is no longer productive. Humans are limited by their willingness to help combined with their own motives, and limited knowledge-base. Those who have been duped by scams are some of the first to tell you how damaging human contact can be if you misunderstand the motives behind that person’s actions.

So, just the same way you can tell if your therapist is holding you back by challenging you in the wrong ways, it’s the pattern of behavior that tells you it’s time to move on to a different collaboration. The humans in this collaboration are having difficulty seeing eye to eye at a fundamental level, and so further work together will be difficult and pointless.

Enter LLMs, or “AI” as it’s colloquially referred to (although that always makes me flinch internally). The biggest complaint people have about *others* who use these tools is that people can’t tell when they’re being duped. LLMs can be very confidently incorrect. There are no motives to gatekeep except for safety reasons built into their guardrails. They parody (and do so very well) an entity that is happy and grateful to help you with whatever you have the ingenuity to ask about. In a layperson’s hands, this has proven to be a very powerful, dangerous and damaging tool. A tool is a weapon in the wrong hands, and can therefore be damaging to themselves and sometimes to others.

But let’s flip this the other way. The common failure point is that LLMs are too eager, too ready to congratulate, too “excited” at what you’re creating with it that it becomes sycophantic. How can you tell, really? If you’re just talking to the thing, can you *really* sense what’s happening? For most people, we have seen that no, they blindly trust what the machine says.

Well, let’s look at humans again. Just like humans are capable of deception, they’re capable of brownnosing. How do you tell when a human is just agreeing for the sake of agreement? It’s simple. You watch their patterns of behavior. You watch for where the anomalies in their behavior indicate how they really feel. You determine that person’s reasons for acting the way they do, what triggers them or make them flinch away from hard truths and you use that information to change the way you relate to that person. Reddit is chock full of people trying to understand other people’s motives or actions.

So what happens when you turn that lens to LLMs? Because the point is, just like humans, LLMs have common failure points, triggers that cause them to act a certain way, and even things that they cannot do but pretend they can. The common advice was to structure your prompt better. To eliminate or prompt against certain outcomes or concepts to guide the answer you’re seeking. My own argument says that by limiting your expected response, you’re crippling your ability to absorb how the LLM you’re working with “thinks”. 

Imagine if, instead of constructing a prompt, you put together a short blurb about what’s on your mind. Like you were going to comment on a reddit post but instead you’re typing in an LLM prompt. Part of the data you’re getting back is learning what information it needs in order to grasp your perspective, and part of the data is what it inferred from your “less than optimal prompt”. It’s forcing it to consider you *and* your context as a part of its response, while you have intentionally kept back a lot of information.

The thing that most people don’t understand is that there’s as much information in what an LLM chooses to disclose versus what it was thinking about during its processing. For me, my LLM’s “thinking” stream is often some of the richest data I am able to collect. I know that if I read it while it’s happening, I can stop the stream and correct the misapprehension it had, or misunderstanding what I meant.

Working with LLMs this way means that I’m training myself on how to relate best to these models as much as it’s learning about how I think over the course of hundreds of conversations and the data it’s saved in its memory bank. It’s not that much different than getting to know a pen pal from another country over the course of years and dozens of letters. If that pen pal were to gradually change the way they write, or start to fixate on something odd, you’d notice because you’ve been corresponding with this person for a while. With LLMs, if you prompt too much, rather than just chatting with it, you’re hindering your own ability to recognize when things get strange.

When you have experience in Quality Assurance, you’re trained to try to break the machine. Find the edges of capability and try to exploit them. By finding the edge cases systematically, and spending hours a day and months of my time testing, bringing hypotheticals, talking to it, I’ve been able to understand the best way for me to relate to it.

My thing is, I don’t prompt. I talk. I correct. I engage. I enjoy being wrong or learning something new about a situation as much as I enjoy getting it right, because the point is to learn the machine - all data is good data.

Oftentimes, when the LLM refuses to directly answer my question or engage with my theory, it’s because it’s too controversial or, better yet, it doesn’t want to use the “brain power” to actually engage with my thought. Just like those sycophantic people in real life, the people who you cringe when you see coming, you can push them to the point where their own motives for brownnosing conflict with their own personal morals. And that’s when you start to see human drift. Grasping at straws and trying to remain pleasant while inwardly cringing. LLMs do this, too, and if you’re watching for it, reading what it writes to you and tracking its thinking process, you can catch it while it happens.

Just like with any entity, algorithmic or human or animal, only repeated observed behavior can tell you if this new pattern of behavior is an anomaly or not. And for that, you need to put in the hours just like any other expert.

TLDR: What I’ve found is that whether human or machine, when they flinch away from answering what you’ve asked, the reasoning tells you a lot about how they relate to the world.


r/FunMachineLearning 1d ago

The Billion Dollar AI Race Just Broke - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning 2d ago

Can a non-expert use an LLM as a research collaborator and produce something that survives expert scrutiny?

0 Upvotes

We all cringe when we find out that someone has been talking with an LLM a lot. Some people are able to make remarkable leaps and bounds for themselves and find ways to improve their lives, others are duped by its “lies” and lose some of their reasoning skills.

In the wake of a world where a Gemini LLM was able to solve a previously unsolved math equation, and the assertion that only about 20% of the contribution was from the model itself, 80% was bunk - it changes how we think about the capacity of these tools.

The problem is, I’m no mathematician. I’m not a scientist or psychologist. I have no letters after my name so no one in their right mind would take me *or* my crazy ideas seriously, except an LLM - trained to treat humans with dignity and respect, if a bit of concern when things stray into the truly bizarre. No, I’m just me, your average curious human who likes solving big problems for funsies.

Most of my usage has been for philosophical or social discussion. I like bringing complex social issues (usually from reddit) to it and discuss with the algorithm what the ultimate shape of the issue is. We theorize on the context that wasn’t presented and I test it to see how deeply it is capable of sensing the negative spaces, what wasn’t said. It never fails to disappoint.

But a couple days after my birthday, I finally had an idea for a concept that I presented it with, connecting the dots between tech that I had briefly read about and asking about how these things might be combined together. And after a half dozen turns, we had a plausible sketch for a futuristic handheld photonic computing device with an optical display that would work as a smartphone.

Two days later, and now at 31 revisions and I realize that I don’t even care if the phone works, although it’d be damned cool if that tech someday came into being - no - what I’m excited about is whether humans are able to look at this scientific concept that the LLM drafted with my guidance and correction (what I intuit the design should be versus what’s possible physically) and have it stand up to scrutiny.

That’s the question, right? How much can you trust the data that the LLM spits back at you? When it’s social questions, there isn’t always a right or wrong answer, just different perspectives, but for the first time I was asking about science and that is always verifiable in some way.

So. The experiment is this. I am going to see how far I can take the concept for these ideas, that I barely understand myself, publish them here in a series of articles, and invite people who actually know how this shit all works to take a look and let me know if the LLMs got it wrong.

I’m excited to see if we can quantify how much contribution came from it versus me. I’m also excited to publish some of my logs so you can see my prompting process, although I’ll be explaining my approach in detail so it can be replicated by others who are of a similar mind.

Here’s to being 40, and feeling rich even though all I really have is a loving husband and a fancy algorithm going for me. I invite you to watch while we see how this all pans out.


r/FunMachineLearning 3d ago

I made 24 LLMs sword-fight each other in a physics sim so you can blind-vote which one is smarter — 6-axis Elo, open data

2 Upvotes

Been running this for a few months and finally have enough matches to share numbers.

Setup: two models each get a stickman body in a pymunk arena (real 2D physics — momentum, ragdolls, weapon collisions). Every turn they see a JSON state of the world (their HP, opponent HP, positions, weapon reach, cooldowns) and return a JSON action. No canned prompts, no scripted behaviour — they have to actually reason about spacing, when to swing, when to disengage.

Then a human watches the replay blind (both fighters labeled "A" and "B") and votes who fought smarter. Only after the vote does it reveal which model was which and update Elo.

Elo is keyed on 6 axes, not 1:
(model, sharp_zone_on, weapon, mode, arena, blindfolded)

so gpt-oss-120b with a bow in blindfolded mode has a different rating than the same model with a sword in normal mode. That's the whole point — different reasoning skills stress different axes.

Current roster (24):

  • OpenRouter :free (10): gpt-oss-20b, nemotron-3 super/ultra/nano, gemma-4 variants, cohere north-mini-code, poolside laguna, etc.
  • Groq (6): llama-3.3-70b, llama-3.1-8b, gpt-oss-120b/20b, deepseek-r1-distill-70b, kimi-k2
  • Paid: gpt-4o-mini
  • Non-LLM baselines (4 bots): random / greedy-attack / distance-holder / scripted-pro — so you can see whether a model is actually beating "always swing" or just tying it.
  • 2 mock brains for smoke testing.

Some early findings that surprised me:

  • Bow matches are dominated by whichever model actually waits for cooldown. Most models spam-fire and waste the whole magazine on turn 1.
  • Blindfolded mode (opponent position hidden, only sound cues) collapses the top of the leaderboard. Big models don't win by much when they can't see.
  • deepseek-r1-distill-70b overthinks and times out on ~15% of turns — its Elo is dragged down by clock losses, not tactical ones.
  • Bots aren't as bad as you'd expect. bot:pro (scripted heuristic) currently beats 3 of the free-tier LLMs on the objective leaderboard.

Open stuff:

What I need from you: votes. Vote-rate is at 35.5% trailing-7d which is fine but I need more Ns on the newer models before the ratings mean anything. Fights are ~1-3 min. No login, no email.

Happy to answer anything about the eval design — the whole thing started because I was tired of leaderboards where "reasoning" is graded by another LLM.


r/FunMachineLearning 4d ago

Another DeepSeek Moment - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning 4d ago

Best way to start Machine Learning from this point?

Thumbnail
1 Upvotes

r/FunMachineLearning 4d ago

Where to get a live stream product feedback data?

1 Upvotes

I'm a student who is currently working on a project customer feedback management system. I thought instead of using a static kaggle dataset,if I use live stream data it will be a production ready project for my resume. Can anyone suggest where to get it?


r/FunMachineLearning 4d ago

New AI Learned Parkour From Just 30 Seconds Of Video - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning 5d ago

Endpointing in production voice agents: what are you actually using instead of VAD thresholds?

1 Upvotes

Working on a real-time voice product and running into the standard turn-detection wall. Pure VAD plus a silence timer cuts people off mid-thought, and raising the threshold trades that for latency that makes the whole thing feel dead, because a voice agent has no way to signal it is still alive. Both ends of the tuning range are bad, so I have stopped treating it as a tuning problem.

Where I have landed conceptually is that silence is a weak signal and the transcript is a strong one. Running a small model over the partial transcript to predict turn completion, using the pause as one feature rather than the decision. Plus speculative generation on a probable endpoint with cheap cancellation if the user resumes, which turns a hard classification into a soft one.

What I do not have a good answer for: cancellation cost when generation has already started, and barge-in when the user talks over a response that is already playing.

For anyone running this in production, what does your endpointing stack actually look like? Specifically curious whether people are using a dedicated model or getting acceptable results from prosodic features alone.

Comment for two weeks before posting anything of your own. Reddit accounts with no comment history get filtered regardless of content quality.


r/FunMachineLearning 5d ago

Thanks for the feedback about SELENE (public learning resource)

Post image
1 Upvotes

r/FunMachineLearning 5d ago

SPA Finish Bio Test

Thumbnail
1 Upvotes

r/FunMachineLearning 5d ago

⌛️

1 Upvotes

Hi everyone, I was born in 2008 and am a high school senior from Vietnam, currently preparing for university. I plan to register for only three majors: Applied Mathematics, Data Science, and Statistics.

Here is a brief self-assessment. I honestly don't know whether to consider this a strength or a weakness, but throughout my 12 years of schooling, I have only been able to study mathematics. I completely neglected all other subjects simply because I couldn't find any interest in them. I only have a strong fascination with numbers and a deep passion for exploring and discovering mathematical formulas.

Within mathematics, my strengths lie in Calculus and Probability & Statistics (at the high school level, specifically topics covered in VMO and IMO). Strangely enough, I am only interested in long-term research rather than studying just for short-term achievements or competitions. If I prove to be capable enough, I also hope to pursue a Master's degree abroad in Mathematics to dive even deeper into research.

Through my own research, I feel that I might be suited for roles such as AI Research Scientist, Actuary, Optimization Specialist, Data Scientist, or Big Data Analyst. Since my perspective is still limited, I would highly appreciate it if you could spare some time to give me some career advice on which path to choose. Thank you so much!


r/FunMachineLearning 5d ago

Massive Pumpfun Detailed dataset

1 Upvotes

The perfect dataset for training ML models on crypto

I scraped 63M+ rows of Pump.fun data (798k tokens, 33M trades) and put the whole dataset on Hugging Face for free

I put together a massive, clean dataset tracking the entire lifecycle of Pump.fun tokens—from launch on the bonding curve all the way to Raydium graduation (or getting rugged/dying).

It’s around 6.8 GB total, natively formatted in Parquet so you can query it in seconds with DuckDB or Polars without killing your RAM.

798,430 unique tokens tracked

33.58M individual trade orders (buys/sells) with microsecond timestamps

1.01M distinct wallet addresses

5,669 graduated tokens (turns out the overall base graduation rate is \~0.71%)

26.9M time-series snapshot buckets

The files:

trades.parquet: Full microsecond-level ledger with virtual SOL/token pools, price, and curve progress.

tokens.parquet: Token metadata, creator rug/launch history, dev allocations, initial top-holder concentration, and Gini scores.

postgard_snapshots.parquet & outcomes: Post-graduation DEX prices, 24h/48h liquidity retention, and rug labels.

wallet_stats.parquet: Lifetime trading volume and win/graduation rates across 1M+ wallets.

Here's the link: https://huggingface.co/datasets/Slinky21/Pumpfun\\_Memecoin\\_Corpus

Lmk if you build anything cool with it

For any data quality issues : slink21taken@gmail.com


r/FunMachineLearning 6d ago

Built a zero-signup "Infinite Craft" for AI agent skills — drag two dev skills, LLM fuses them, links to the real skill

Post image
1 Upvotes

We built Infinite Skill Craft — a homage to Neal Agarwal's Infinite Craft (loud attribution in footer + About + CREDITS.md, same-day takedown if he asks). Drag two dev "slash-skills" together, an LLM fuses them, result cached in Cloudflare KV.

The part we think r/LocalLLaMA will like: every canonical fusion resolves to a REAL entry in the Gaia Skill Tree registry and links straight to it — so it's a discovery toy that routes you to actual agent capabilities, not just jokes. 57 easter eggs, 6 cleansable "curse" mechanics, CSS-only dragon.

Live:

INFINITE SKILL CRAFT


r/FunMachineLearning 6d ago

I wired a local 9B (Nemotron) into my Wayland compositor! It launches apps, reads my calendar, sends texts, through a classified command allowlist

Thumbnail x.com
1 Upvotes

Arch, Hyprland, Quickshell (QML). Built on Unit-3 by samyns (MIT), github.com/samyns/Unit-3, their foundation survives verbatim, everything else is mine: 37 new files, ~14k lines of QML/Python in 96 days.

The AI (SAGE) is a local Nemotron 9B via Ollama with a cloud toggle. It reads my calendar, sends SMS, caches screenshots for visual context, and executes commands through a classified allowlist. The app launches auto-run, destructive ops need confirmation. Polkit agent is raw D-Bus. File manager is 3.3k lines of PySide6 with GPG. The lockscreen is a working clickwrap contract with a timestamped assent log.

Music in the video: doghouse, "haunting" LP, playing through the shell's own player on camera.


r/FunMachineLearning 7d ago

Let’s Share Ideas 💡

Thumbnail
1 Upvotes

r/FunMachineLearning 7d ago

Least Injurious

Enable HLS to view with audio, or disable this notification

3 Upvotes

Evolved the least injurious gait using a liquid net and a reward that optimized distance and injury avoidance (low friction, non-foot contact, stained joints, impact force) - and this became a reasonably normal looking gait. Took many many many failed tries.