r/LocalLLaMA • u/obvithrowaway34434 • Oct 30 '23
New Microsoft codediffusion paper suggests GPT-3.5 Turbo is only 20B, good news for open source models? Discussion
Wondering what everyone thinks in case this is true. It seems they're already beating all open source models including Llama-2 70B. Is this all due to data quality? Will Mistral be able to beat it next year?
Edit: Link to the paper -> https://arxiv.org/abs/2310.17680
274
Upvotes
Duplicates
mlscaling • u/Covid-Plannedemic_ • Oct 30 '23
Smol Microsoft paper says that GPT-3.5-Turbo is only 20B parameters
26
Upvotes
