Skip to content
X ยท @teortaxesTex ยท X / Twitter

RT Jenia Jitsev ๐Ÿณ๏ธโ€๐ŸŒˆ ๐Ÿ‡บ๐Ÿ‡ฆ ๐Ÿ‡ฎ๐Ÿ‡ฑ ๐Ÿ‡ฎ๐Ÿ‡ท: Art of benchmarking 2: lower the eval scores of the reference model you compare to such that …

RT Jenia Jitsev ๐Ÿณ๏ธโ€๐ŸŒˆ ๐Ÿ‡บ๐Ÿ‡ฆ ๐Ÿ‡ฎ๐Ÿ‡ฑ ๐Ÿ‡ฎ๐Ÿ‡ทArt of benchmarking 2: lower the eval scores of the reference model you compare to such that your model appears "frontier-level champion". SOOFI S trained again the already fully open Nemotron-3-Nano. Report shows for Nemotron-3-Nano strongly lower scores than the original report