Skip to content
X · @teortaxesTex · X / Twitter

5.6 Sol's conclusions based on the paper > At 1M context that is approximately 4.72 TFLOP per generated token …let me just say GPT still has ways to …

5.6 Sol's conclusions based on the paper> At 1M context that is approximately 4.72 TFLOP per generated token…let me just say GPT still has ways to improveKV is probably correct, except it's not going to be BF16Zephyr: Very heavy attention backbone