X · @teortaxesTex
· X / Twitter
5.6 Sol's conclusions based on the paper > At 1M context that is approximately 4.72 TFLOP per generated token …let me just say GPT still has ways to …
5.6 Sol's conclusions based on the paper> At 1M context that is approximately 4.72 TFLOP per generated token…let me just say GPT still has ways to improveKV is probably correct, except it's not going to be BF16Zephyr: Very heavy attention backbone