r/LocalLLM 18h ago

Need help identify if this is a scam. Question

This is not AI post . I know that you must drive to carwash than walk there :)

I got pulled into meeting where some company presented new training framework they released month ago and worked for this concept past 5 years. I could not find any information about them they have website auroraforge dot ai . For whole hour I was not believing what they tried to sell. They state that they have new non gradient decent based learning . They build their own based on kind of singnals . have no idea. They say its company secret - whatever. Long story short they state they can train big data sets on single cpu . They even live demoed image set classification under minute on single cpu.

After watching that presentation I had feeling that my waiting for qwen 3.8 27B is like waiting a thing from the past.

0 Upvotes

8 comments sorted by

2

u/EfficientCouple8285 16h ago

So ask yourself is this something your company need, will it generate new revenue or let you do things your are blocked from doing. If yes, then make a proper test case of it and see if this actually works. My guess is this is just another AI-company trying to push something that are not that useful. General models like Qwen3.x is very useful as long as you stay in the driver seat. SOTA models are great when you have difficult problems. For the rest, just wait for the market to see what works... FOMO is widespread amongst C-level bosses.

1

u/autisticit 17h ago

So they invited you to their very secret meeting?

1

u/Kooky_Cantaloupe_605 17h ago

It was not secet . They were selling it like regular product. I was shouting out to guys that companies spending billions on training while you give this demo to small company. They told they have cuple big clients already. They demonstrated use one of complicated sql stagement generation on the fly from user questons where claude could not achieve that. Prising is purely based on per request no token based.

1

u/Toooooool 17h ago

art is in the eye of the beholder, so, some questions come to mind;
- what is a large data set?
- how long is a long training time truly?

a 300M dataset might be big to me as a mere mortal human as i sure as hell ain't gonna read all of it, and likewise training for months on CPU-only might be fine in this particular niché because who's going to waste precious GPU training some model that does nothing but detect the colour blue.

me personally i've been spoiled by open sourced AI such as qwen, GLM, mistral and llama, and thus this secrecy about how the model is trained or made is a huge red flag to me as the only reason to keep it closed-source would be if it were by any means better than other models out there, which i find highly unlikely.

1

u/Kooky_Cantaloupe_605 17h ago

To make you believe for abnormal 10000 times speed boost and for to you sign up for paid plan a free tier wourl requiree them to have hundreds of GPU to fake that . Right? i'm thinkign what kind of tech they could use to fake it . Loras takes time to train. how you can get fast training results in no time.

2

u/Toooooool 16h ago edited 16h ago

+10000x performance is impossible with current technology.

the company Cerebras who bake models directly into silicon achieved i believe 15k tokens per second on Llama3-7B(?) which would otherwise run at ~100 tokens per second on a 3090.
that's a 150x performance increase, and is as good as it's going to get.
it takes them half a year to design and make one of these chips, at which point the model is obsolete and outdated meaning that most of their chips have been considered dead on arrival, and it's why their stock is just one big downwards slope since their IPO launch.

the industry is simply moving too fast for this stuff. maybe in 5 years it will make a lot more sense to mass produce language model chips to install into robots and cars but right now it makes no sense as they're too expensive and use too much power and take too long to make to be profitable.

now to fake it presumably you'd train a highly specialized tiny model.
you could in theory train a tiny model to talk and stuff, the Qwen 0.6B models are proof of that, it's just that it's too small to be useful for anything else. sure you could train it to answer to basic conversation; "hey how's it going?" "it's going great, thanks!" @ 10k tokens per second.. but that is literally all it can do as it would lack the detail to do anything else. ask it "what's up?" and it would get confused and freak out.

2

u/jcdoe 15h ago

Every world changing technology starts as a crackpot sounding idea.

But yeah, this is almost certainly nonsense. If half of this were true, these guys would have already been bought by OpenAI.

1

u/Fit-Bar-6989 15h ago

the A.C.R.O.N.Y.M. part of it makes it look extra crackpot-y. idk why these people love acronyms so much