We open-sourced a harness for evaluating VLMs on your own video, with traced runs
The framing that finally made VLM evaluation tractable for us is simple: decide what setup is right for your task, on your videos, at the quality,…
The framing that finally made VLM evaluation tractable for us is simple: decide what setup is right for your task, on your videos, at the quality,…
https://preview.redd.it/h40uz1bvhn9h1.png?width=808&format=png&auto=webp&s=f68d2640255989fdefa3c6e5e4a5b0e1690731f6 Motherboard is a Asus Proart Creator B850 Neo Slot 1 & Slot 2 (PCIe 5.0): These are the two main physical x16 slots. If you…
There are a lot of posts about the models and benchmarks, but I am more interested in the workflows that people use. What is one workflow…
My end goal is to have multippe small to medium models running locally for data parsing and extraction tasks, working with logs, and many data inputs…
Tell me if it's a good idea or not, I have zotac solid 5090 with 128gb RAM, thinking of selling only 5090 and getting 5 x…
Having a lot of fun using Gemma 4 as an assistant, but is growing frustrated with the poor default image resolution setting for image vision. Tasks…
submitted by /u/paf1138 [link] [comments]
Hey, I currently use my RTX 4060 8G for inference with Qwen 3.6-35B-A3B Q8 (q8 for everything weight,value,key) (for quality over speed, with CPU &DDR4 offloading)…
Spent two weeks running 40 coding prompts across 7 models, self evaluating each output on correctness, completeness, and whether I'd actually use it without editing(which to…
submitted by /u/undefdev [link] [comments]