Despite not being trained to, it turns out the Pearson correlation between a models AA Intelligence Index score and its ability to generate Base64 encoded responses is 0.91
I built Encode Bench, an open benchmark that asks a model to solve a task and return the answer as a Base64 payload. The initial result…