Turkish-Gemma-9b-T1 NanoQuant 1-bit

This repository contains a NanoQuant-compressed checkpoint derived from ytu-ce-cosmos/Turkish-Gemma-9b-T1.

Quantization details

  • Method: NanoQuant
  • Target linear-weight bitrate: 1.0 bit
  • Sequence length: 2048
  • Calibration samples: 64
  • English calibration samples: 32
  • Turkish calibration samples: 32
  • Non-factorized epochs: 2
  • Factorized epochs: 2
  • ADMM outer iterations: 100
  • Seed: 42
  • Quantization hardware: NVIDIA A100 80 GB

Model-level knowledge-distillation tuning was performed after layer compression. The tuning loss decreased from 1.9244 to 1.7136 over 8 reported epochs.

Important

This file is a NanoQuant checkpoint. It is not a GGUF file and is not a standard Transformers save_pretrained checkpoint.

It requires the NanoQuant code and compatible CUDA kernels for loading and inference. It cannot be loaded directly with llama.cpp, Ollama, vLLM, or AutoModelForCausalLM.from_pretrained().

Files

  • Turkish-Gemma-9b-T1-NanoQuant-1bit-en32-tr32.pt: compressed checkpoint

Base-model license

Use of this checkpoint remains subject to the license and usage terms of the original Gemma-based model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for suayptalha/Turkish-Gemma-9b-T1-1bit

Finetuned
(5)
this model