Turkish-Gemma-9b-T1 NanoQuant 1-bit
This repository contains a NanoQuant-compressed checkpoint derived from
ytu-ce-cosmos/Turkish-Gemma-9b-T1.
Quantization details
- Method: NanoQuant
- Target linear-weight bitrate: 1.0 bit
- Sequence length: 2048
- Calibration samples: 64
- English calibration samples: 32
- Turkish calibration samples: 32
- Non-factorized epochs: 2
- Factorized epochs: 2
- ADMM outer iterations: 100
- Seed: 42
- Quantization hardware: NVIDIA A100 80 GB
Model-level knowledge-distillation tuning was performed after layer compression. The tuning loss decreased from 1.9244 to 1.7136 over 8 reported epochs.
Important
This file is a NanoQuant checkpoint. It is not a GGUF file and is not a
standard Transformers save_pretrained checkpoint.
It requires the NanoQuant code and compatible CUDA kernels for loading and
inference. It cannot be loaded directly with llama.cpp, Ollama, vLLM, or
AutoModelForCausalLM.from_pretrained().
Files
Turkish-Gemma-9b-T1-NanoQuant-1bit-en32-tr32.pt: compressed checkpoint
Base-model license
Use of this checkpoint remains subject to the license and usage terms of the original Gemma-based model.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for suayptalha/Turkish-Gemma-9b-T1-1bit
Base model
ytu-ce-cosmos/Turkish-Gemma-9b-T1