X · @emollick
· X / Twitter
This is a wonderful visualization.
This is a wonderful visualization.Matt Henderson: what is a multimodal LLM thinking as it watches a video?Gemma 4 12B reads raw image patches, as if they were tokens. It was never trained to predict anything at these 'tokens' - but this video shows what it would predict if you did sample from its next token prediction