Quantizing 70b models to 4-bit, how much does performance degrade?

ae_dataviz@alien.top · 1 year ago

Quantizing 70b models to 4-bit, how much does performance degrade?

Dusty_da_Cat@alien.top · 1 year ago

The golden standard is 2 x 3090/4090 cards, which is 48 GBs of VRAM total. You can get by with 2 P40s(Need cooling solution) and run onboard video, if you want to save some money. The speeds will be slower, but still better than running on System RAM on typical setups.