UD-Q5_K_L not outputting tokens using ROCm (Strix Halo)

#23
by philtheriver - opened

I've been using this model on a Strix Halo since it was released but the PP speeds on vulkan remain a bit slow.
Switching to ROCm makes the model simply not output any tokens, I've tried many different backend versions and no change.

This is on latest llama.cpp, currently using this repo's Q5_K_XL but I was having the same on the initial poolside Q4

This does not seem to happen with the newest Q4 in the poolside/Laguna-S-2.1-GGUF repo

Sign up or log in to comment