Local LLM Peeps
I am 80% done with a harness that works for local and API but is local first. The harness has some interesting logic around multiple agents…
I am 80% done with a harness that works for local and API but is local first. The harness has some interesting logic around multiple agents…
I am trying to plan and deploy a machine that serves models for coding, Hermes, and whatever else. It's got multiple GPUs in it, and I…
Have you ever thought to yourself that sometimes things happen with AI companies in a legal way, but not in an ethically correct way? Have you…
This is in response to the common post where OP has acquired some cool hardware and is wondering what to do with it. The standard response…
Quick teaser of what I’ve been working on over the last few weeks: a streaming medical speech-to-text model that runs fully on-device. This demo is running…
Domain-Specific Small Language Models Guglielmo Iozzia Review by u/skiata I came across Domain-Specific Small Language Models (https://www.manning.com/books/domain-specific-small-language-models) by attending the author's talk at an ACM Tech…
https://preview.redd.it/iiiqwt96tn9h1.png?width=3004&format=png&auto=webp&s=f02fba9f64e27ac91b2ae4cd478842106b294366 https://preview.redd.it/47cb5u96tn9h1.png?width=3024&format=png&auto=webp&s=b1cee93477970b8b0a636c37be657fecd38ba968 https://preview.redd.it/t45iv1a6tn9h1.png?width=3018&format=png&auto=webp&s=beef94ac59
If you get a good deal on some Xeons with a lot of memory bandwidth, or a cheap GPU for home inference, that's cool, no disrespect.…
I have collected 8 Tesla T4 Datacenter Cards from a few retired VDI servers. I have one in a DEG1 and works ok on n its…
The framing that finally made VLM evaluation tractable for us is simple: decide what setup is right for your task, on your videos, at the quality,…