Yi-34B vs Yi-34B-200K on sequences <32K and <4K

DreamGenX@alien.topB to

LocalLLaMA@poweruser.forumEnglish · 2 years ago

Hello!

By popular demand I am planning a fine-tune of https://huggingface.co/dreamgen/opus-v0-7b on top of Yi-34B and wonder whether to use the 200K as the base.

The regular Yi-34B seems slightly better than Yi-34B-200K on standard benchmarks, but I wonder how it “feels” and whether the loss of performance on short context is worth it, given that the regular version can be used up to 32K tokens.

(Yi-34B vs Yi-34B-200K)

Did anyone try an analysis of these 2 models on various sequence lengths (<4K, <8K, <16K, etc.)?

Chat

Igoory@alien.topB
link
fedilink
English
arrow-up
1·
2 years ago
The normal Yi model can go up to 32K btw

Yi-34B vs Yi-34B-200K on sequences &lt;32K and &lt;4K

Yi-34B vs Yi-34B-200K on sequences &lt;32K and &lt;4K

Yi-34B vs Yi-34B-200K on sequences <32K and <4K

Yi-34B vs Yi-34B-200K on sequences <32K and <4K