Model quantization

after @jeetnirnejak

Quantize

Nimbus-3 8B
72% smaller

WEIGHTS

4-bit / weight

4 shade levels — fewer bits store coarser values.

SIZE ON DISK4.5 GB
VRAM NEEDED6.0 GB
THROUGHPUT88 tok/s
QUALITY KEPT96%

Fits comfortably on 16 GB

6.0 GB needed · 10.0 GB free

Q4_K_M — best size / quality balance