DEV Community

#gguf

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
GGUF Quantization: Which Level Should You Use?

GGUF Quantization: Which Level Should You Use?

Comments
5 min read
How to Import LM Studio Models into Ollama (No Re-Download!) 🚀

How to Import LM Studio Models into Ollama (No Re-Download!) 🚀

Comments
3 min read
GGUF vs GPTQ vs AWQ: Which Quantization Format Should You Actually Use?

GGUF vs GPTQ vs AWQ: Which Quantization Format Should You Actually Use?

Comments
4 min read
Unsloth Releases Inkling-GGUF: A Multimodal MoE Model for Developers

Unsloth Releases Inkling-GGUF: A Multimodal MoE Model for Developers

Comments
3 min read
GnLOLot Releases MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF for Enhanced Local AI Development

GnLOLot Releases MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF for Enhanced Local AI Development

Comments
3 min read
Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX

Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX

Comments
3 min read
KAT-Coder V2.5 Local Setup Guide: GGUF, vLLM, SGLang

KAT-Coder V2.5 Local Setup Guide: GGUF, vLLM, SGLang

20
Comments
7 min read
llama-bench skipped FA on capable GPUs — b9437 corrects it