I'm not planning to support Ollama, as I think KoboldCpp is light and easy to use. But I will keep an eye on it to check if the integration can be done quickly. Last time I checked (a few years ago), their API was different from other software.
I haven't tested local models for quite some time, but the recently released base models to look for abliteration/finetunes are:
- Gemma 4 (12B for smaller GPUs)
- Qwen 3.5 & Qwen 3.6 (there is a Qwen 3.5 9B for smaller GPUs)
Otherwise, it's probably still going to be some Mistral finetunes.