Any way to decrease inference time during long chats?(+decrease repetition without breaking things)

The_One_Who_Slays@alien.top · 1 year ago

Any way to decrease inference time during long chats?(+decrease repetition without breaking things)

Former-Ad-5757@alien.top · 1 year ago

What do you want to happen when the total chat reaches 8k? Because there the server has to make a choice it can keep adding more context so it slows down, it can simply cut off the first messages but then it will for example forget its own name, or it could for example (this is a method I use but it costs interference time as you ask a 2nd question behind the scenes) ask the model to summarize the first 4K of the context so it will retain some context and still retain speed.