Skip to content
r/LocalLLaMA · Communities

model: support Longcat-Flash (need testing) by ngxson · Pull Request #19182 · ggml-org/llama.cpp

This PR should be ready for testing now. I tested with a very small (8B params) sub-model extracted from the original one. Appreciate if someone can test with the bigger model. GGUF(for testing) from PR: (Please check latest comments at bottom for updated GGUFs) https://huggingface.co/ggml-org/LongCat-Flash-Chat-GGUF/t