This is a wonderful visualization. Matt Henderson (@matthen2) what is a multimodal LLM thinking as it watches a video? Gemma 4 12B reads raw image patches, as if they were tokens. It was never trained to predict anything at these 'tokens' - but this video shows what it would predict if you did sample from its next token prediction head Video โ https://nitter.net/matthen2/status/2078818110597972114#m
This AimostAll brief summarizes the linked source so readers can scan AI developments quickly and jump to the original reporting when needed.