Why is no one releasing 70b models?

Longjumping-Bake-557@alien.top · 1 year ago

Why is no one releasing 70b models?

Armym@alien.top · 1 year ago

Do you think that finetuning models with more parameters requires more data to actually do something?

thereisonlythedance@alien.top · 1 year ago

With a full finetune I don’t think so – the LIMA paper showed that 1000 high quality samples is enough with a 65B model. With QLoRA and LoRA, I don’t know. The number of parameters you’re affecting is set by the rank you choose. It’s important to get the balance between the rank, dataset size, and learning rate right. Style and structure is easy to impart, but other things not so much. I often wonder how clean the merge process actually is. I’m still learning.