Could multiple 7b models outperform 70b models?

freehuntx@alien.top · 2 years ago

Could multiple 7b models outperform 70b models?

FullOf_Bad_Ideas@alien.top · 2 years ago

Jondurbin made something like this with qlora.

The explanation that gpt-4 is MoE model doesn’t make sense to me. Gpt4 api is 30x more expensive than gpt-3-5-turbo. Gpt-3-5 turbo is 175B parameters, right? So, if they had 8 220B experts, it wouldn’t need to cost 30x more, it would be 20-50% more for API use. There was also some speculation that 3.5 turbo is 22B. In that case it also doesn’t make sense to me that it would be 30x as expensive.

AutomataManifold@alien.top · 2 years ago

Just to note: don’t read too much into OpenAI’s prices. They’re deliberately losing money as a market-capturing strategy, so it’s not guaranteed that there’s a linear relationship between what they charge for a given service and what their actual costs are.

Cradawx@alien.top · 2 years ago

No, several sources include Microsoft have said GPT 3.5 Turbo is 20B. GPT 3 was 175B, and GPT 3.5 Turbo was about 10x cheaper on the API than GPT 3 when it came out so it makes sense.

FullOf_Bad_Ideas@alien.top · 2 years ago

Yeah if that’s the case, I can see gpt-4 requiring about 220-250B of loaded parameters to do token decoding