r/LocalLLaMA
· Communities
How to handle frame deduplication for VLMs? (Analyzing long videos with MiniCPM-V)
Hey everyone, I'm trying to use MiniCPM-V to analyze long react-style videos (where a creator is on-screen commenting on a piece of content, photo, or text shown next to them). Right now, I have the audio side figured out—using WhisperX to get word-level transcripts with precise timestamps. My next step is to inject th