Instructions to use RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf", filename="llama3-8b-stargate-m1.IQ3_M.gguf", )
llm.create_chat_completion( messages = "No input example has been defined for this model task." )
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M
Use Docker
docker model run hf.co/RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf with Ollama:
ollama run hf.co/RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M
- Unsloth Studio
How to use RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf to start chatting
- Atomic Chat new
- Docker Model Runner
How to use RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf with Docker Model Runner:
docker model run hf.co/RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M
- Lemonade
How to use RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull RichardErkhov/scandukuri_-_llama3-8b-stargate-m1-gguf:Q4_K_M
Run and chat with the model
lemonade run user.scandukuri_-_llama3-8b-stargate-m1-gguf-Q4_K_M
List all available models
lemonade list
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Quantization made by Richard Erkhov.
llama3-8b-stargate-m1 - GGUF
- Model creator: https://huggingface.co/scandukuri/
- Original model: https://huggingface.co/scandukuri/llama3-8b-stargate-m1/
Original model description:
license: mit
STaR-GATE
This repository contains the iteration 1 meta-llama/Meta-Llama-3-8B-Instruct model from an additional experiment for STaR-GATE: Teaching Language Models to Ask Clarifying Questions. Note that this experiment is an extension and is not yet included in the most recent revision of the linked preprint. The weights contained in this repository are represented by the blue line in the left-side win-rate graph below. Note that this repository contains the weights for iteration t=1, i.e. only one iteration of self-improvement.
When prompting language models to complete a task, users often leave important aspects unsaid. While asking questions could resolve this ambiguity (GATE; Li et al., 2023), models often struggle to ask good questions. We explore a language model's ability to self-improve (STaR; Zelikman et al., 2022) by rewarding the model for generating useful questions-a simple method we dub STaR-GATE. We generate a synthetic dataset of 25,500 unique persona-task prompts to simulate conversations between a pretrained language model-the Questioner-and a Roleplayer whose preferences are unknown to the Questioner. By asking questions, the Questioner elicits preferences from the Roleplayer. The Questioner is iteratively finetuned on questions that increase the probability of high-quality responses to the task, which are generated by an Oracle with access to the Roleplayer's latent preferences. After two iterations of self-improvement, the Questioner asks better questions, allowing it to generate responses that are preferred over responses from the initial model on 72% of tasks. Our results indicate that teaching a language model to ask better questions leads to better personalized responses.
Usage
Reference the paper appendix sections A.5.2 (Figure 14: Questioner Elicitation Prompt) and A.6.2 (Figure 17: Questioner Win-Rate Response Prompt.) to see how you can prompt the model for elicitation or for final responses. All code and data for the project can be found here.
- Downloads last month
- 408
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit