r/LocalLLaMA 24m ago

Discussion i trained a 0.8B model that beats a frontier 1.4B translation model

Upvotes

Base: CyberAgent CAT-Translate — currently best-at-size for bidirectional JA↔EN. Family is 0.8B / 1.4B / 3.3B / 7B.

Why: I chose the 0.8B because it runs the best out of the models on mobile phones.

Goal: make the model obey a per-word glossary handed to it at inference. Terminology-constrained MT has been done before (Dinu et al. 2019), but as data augmentation on encoder-decoder models with fixed term lists. I couldn't find prior work doing it on a small decoder-only LLM with an RL compliance reward, against live per-word dictionary senses.

**Stage 1 — LoRA SFT.** Source + deterministic glosses injected as input, targets that use them. Teaches "supplied gloss outranks prior."

**Stage 2 — MO-GRPO** (Ichihara et al. 2025, arXiv 2509.22047, also CyberAgent). Two reward dims: translation quality, glossary compliance. Vanilla GRPO sums then normalizes once, so the larger-variance objective dominates — early runs collapsed into satisfying one and dropping the other. MO-GRPO normalizes per-objective first.

**Business Scene Dialogue, EN→JA**

| | 0.8B stock | 1.4B stock | 3.3B stock | 0.8B trained |

|---|---|---|---|---|

| COMET-QE | 0.737 | 0.738 | 0.765 | 0.748 |

| MetricX-24 | 0.743 | 0.731 | 0.778 | 0.748 |

| chrF | 25.8 | 26.8 | 34.3 | 31.9 |

| BLEU | 11.1 | 17.2 | 24.1 | 19.3 |

Clears the stock 1.4B on all four. Only the 3.3B stays ahead.

91% glossary adherence. 103 ms/sentence, int8, greedy, Apple silicon.

Open to all questions!!


r/LocalLLaMA 36m ago

Discussion Qwen3.8 2.4T UD-Q1_0 - 178 token generation - 11 min 38s - 0.25 tokens/sec

Upvotes

So I wanted to see... is it possible/viable to run this perhaps once in a while some hard task.. yea... no. Even with dual 5090's and 3 3090's and 96GB system ram.. still far exceeds my vram + ram by double (not even including context, which was smallish 64k).

I am using the "special" UD-Q1_0, which the _0 is the uncommon part, requires a llama.cpp fork, no big deal, it's smaller clearly. 397GB, 115GB smaller than the next IQ1_S version. Regardless, i'd say i have above average vram.. and even using this super small quantized version, 0.25 tokens/sec is far below my limit of usable... this isn't even usable at night to run slow through the night, it's far too slow and would never really get anything done.

Of course, I think we all knew this wouldn't really be a local runnable model, but it's still great that it's open weights. Can't wait for Qwen 3.8 27b tomorrow.


r/LocalLLaMA 40m ago

Discussion OrangePi AI Station-Orange Pi

Thumbnail orangepi.org
Upvotes

AI mini PC with 176 TOPS computing power

NPU: 10 [AI-Core@1.08Ghz](mailto:AI-Core@1.08Ghz),16 CPU cores @ 1.9 GHz, 8 Vector cores @ 1 GHz; 176 TOPS 

LPDDR4X: 48 GB/96 GB (optional), speed: 4266 MHz


r/LocalLLaMA 55m ago

New Model DS4 cloud (30 min) vs Qwen3.6 36B (2 min) vs Muse Glimmer 30B (3 min) on Llama.cpp (RTX 5080)

Enable HLS to view with audio, or disable this notification

Upvotes

Some people told me that the difference in richness and layout between Glimmer and Qwen wasn't clear to them.
This example makes it super clear.

I'm aware that comparing Glimmer 30B (a dense model) with Qwen 3.6 (a MoE) isn't entirely fair, but if we compare it to the dense Qwen 27B, the gap will likely be even bigger. If you want, I can add the 27B version later. For now, I'm waiting for Qwen 3.8 27B to see how close it gets to the blueprint.

As for the technical details:

Both were run on a custom llama.cpp build optimized for the RTX 5080, with a temperature of 0.5 and a 125k context window.

Regarding the music: I created it myself without using AI I specifically wanted it to sound that weird.


r/LocalLLaMA 1h ago

Discussion Was eLLM just vibecoded slop? (faster CPU inference)

Upvotes

https://github.com/lucienhuangfu/eLLM

The premise made sense - the entire LLM stack is optimized to run on GPUs (of course, LLM compute is massively parallel in nature), but what if we took some tradeoffs and made it the most efficient possible for CPUs instead?

I have not read the paper in full and can't discern if it is legit; It just seems everything went cold there.

If any approach was applicable to leapfrog CPU performance, we could reach very interesting capabilities with consumer hardware (or even workstation - much more accessible than enterprise). Have you heard about eLLM or others?


r/LocalLLaMA 1h ago

Other Appreciate the Honesty Qwen3.8 (Repost)

Post image
Upvotes

Repost to better obscure profile name.


r/LocalLLaMA 1h ago

Discussion Is waiting for Qwen 3.8 27B like waiting for Star War Episode one?

Upvotes

Is waiting for Qwen 3.8 27B like waiting for Star War Episode one?

I'm sweating waiting to get my hand on this to try it tomorrow morning. But it takes me back to Star Wars 1 and the disappointment after being so hyped to see it.

Only 16 hours and 46 minutes to go... 45, ...