What GPU are you using? I'm using the exact same model and quant, It's taking forever for me to load it up via llama.cpp on an RX 7900 XTX and I keep getting an error that it lost connection to the endpoint was lost mid-reply, is there something specific you did, what context size are you using?