@quimedesu
Post
Post 1 of 2
If you have a 16gb gpu this is for you, I even have a command line to you use to fit 128k context on this. 10gb qwen 3.8 27b that can code just fine. https://huggingface.co/quimmedes/Qwen3.8-27B-XYZ/blob/main/Qwen3.8-27B-Q3-XYZ.gguf 12gpu are also covered. https://huggingface.co/quimmedes/Qwen3.8-27B-XYZ/blob/main/Qwen3.8-27B-Q2-XYZ-v2.gguf
Post 2 of 2
by the way you can use LocalLLM Tool to setup it, https://github.com/TRI-Tech-Revolution-Intelligence/LocalLLM
Explanation
What it says @quimedesu is offering very aggressively quantized Qwen3.8-27B GGUFs aimed at consumer GPUs. The claim is that a ~10 GB Q3 variant runs adequately for coding on a 16 GB GPU, including 128K context with a particular command-line setup; a smaller Q2 variant is intended to cover 12 GB GPUs. A second post points to LocalLLM, a setup tool for running the model.
Context This is essentially a “large model on small VRAM” optimization pitch. Fitting a 27B model into 10–12 GB requires severe weight quantization, and 128K context additionally depends heavily on KV-cache quantization/offloading and runtime settings. The post does not provide benchmarks, the claimed 128K command, quality comparisons against Q4/Q8/FP16, or measured VRAM/speed figures.
Why it matters Potentially useful if you want the largest feasible local coding model on a 12–16 GB GPU without multi-GPU inference. Before adopting it, verify coding quality versus higher-bit Qwen3.8-27B quants, actual VRAM at 128K, prompt/decode speed, KV-cache settings, and whether the XYZ quantization method has reproducible advantages over standard GGUF Q2/Q3 schemes. The post’s central performance claim is currently unsupported by evidence included here.