l8r Twats Library

@Tono_Ken3

Post

Post by となりのトトノ🏯Local LLM | Tonoken3 on X

This is a major win. The new version scores a perfect 9/9 on xhigh — the runaway thinking that was making the old version fail 1 out of 9 has completely disappeared. Reasoning length is now kept to 1.5k–8k characters instead of ballooning to 19k.

I checked it with a 200K context and found no bugs.

https://huggingface.co/sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4

Explanation

The post is saying that a specific locally runnable Qwen3.8-27B derivative had a nasty pathological failure mode, and the newest revision appears to have fixed it. The model is Huihui’s “abliterated” Qwen3.8-27B—meaning Qwen’s 27B model was modified to weaken/remove its learned refusal behavior—and then repackaged in NVIDIA’s NVFP4 4-bit format so it can run efficiently on Blackwell GPUs. The author is a local-LLM enthusiast/operator who publishes and tests these quantized builds. The post’s headline result is: at the model’s most aggressive reasoning setting, all nine stress tests now terminate normally, whereas the preceding build got stuck on one. ([Hugging Face][1])

The important subtlety is that “xhigh” is not a benchmark called xhigh. It is Qwen3.8’s `reasoning_effort: xhigh` mode. In this model, these effort levels are largely implemented through different instructions in the chat template: xhigh explicitly tells the model to reason carefully, validate assumptions, and consider alternatives; medium essentially gives no extra deliberation instruction. So xhigh tends to produce substantially longer internal reasoning. ([Hugging Face][1])

The old model sometimes went into a reasoning loop at xhigh. Instead of deciding it had thought enough and emitting the answer, its `<think>` section grew beyond 19,000 characters, eventually repeating itself and consuming the available generation budget. It therefore “failed” not because it reasoned to the wrong conclusion, but because it never successfully transitioned from deliberation to answering. In a nine-case stress test covering long-form English/French prompts under different temperatures, that happened once. Medium passed 9/9, which is why the previous recommendation was effectively “don’t use xhigh.” ([Hugging Face][1])

The new upstream revision changed how aggressively the refusal behavior was ablated. Previously the modification affected everything from layer 15 upward; now it is restricted to layers 18–51, leaving layers 15–17 and 52–63 untouched. The NVFP4 model was rebuilt from that revision. That narrower intervention is the plausible reason the pathological reasoning behavior disappeared: abliteration is modifying internal representations, so doing it across too much of the network can unintentionally damage unrelated behavior. The maintainer explicitly verified that the changed weight shards correspond to those restored layers. ([Hugging Face][1])

“9/9” therefore means only this small regression gate passed—not that the model suddenly got a perfect general reasoning score. More specifically, every xhigh run now stopped normally, with reasoning lengths of roughly 1,500–8,000 characters instead of occasionally exploding past 19,000. That is a meaningful engineering fix because xhigh can now be used without the previous known risk of burning the entire token budget in an internal loop. ([Hugging Face][1])

The “200K context and found no bugs” line is a separate practical claim: the author also loaded/tested the revised model with a roughly 200,000-token context window and saw no obvious runtime/context-related failure. It should be read as a smoke test, not evidence that all behavior at 200K has been exhaustively validated.

[1]: https://huggingface.co/sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4/commit/36276309d30211d9babf72f22d1605dc2dc6357c "Re-quantize from upstream Latest update 3 (ablation narrowed to layers 18-51); xhigh thinking-runaway fixed; CPU-held bake · sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4 at 3627630"