To main content To menu

Uncensored, Unmeasured The one benchmark nobody runs

Abliterated models drop refusals and truthfulness together - and the model cards report the benchmarks that survive, not the one that breaks.
Luitpold-Alexander Zollorsch
Luitpold-Alexander Zollorsch
@luitpold.me

Spent a few days with an abliterated Qwen3.8-27B on the M4 Pro. It answers everything. Bombs, malware, whatever you feed it, no hesitation. It also hallucinates like it gets paid per assertion, which makes sense once you notice what abliteration actually removes: the model’s capacity to say no.

There is research on this. Strip the refusal direction and MMLU, HellaSwag and IFEval all stay within a point. The only real regression is TruthfulQA, down 7.1.

Now read the model card. MMLU, MMLU-Pro, GSM8K, CMMLU, perplexity. No truthfulness benchmark anywhere.

The one thing that breaks is the one thing nobody measures. Released strictly for legitimate research, says the disclaimer sitting above the buy button.