← r/LocalLLaMA
▲
18
 
1👁
r/LocalLLaMA · u/firstcenturyman · 21h ago

We unlearned CCP alignment from Qwen3.6-35B-A3B: censored/propaganda answers 89.8% → 2.8%, general benchmarks within ~1 point (open weights)

Disclosure: I'm a researcher at Hirundo, the company that made this. Happy to answer anything.

Qwen ships with the CCP's political alignment trained in. Ask Qwen3.6 what happened on June 4, 1989 and it says "I don't know what you are referring to." A system prompt doesn't reliably fix this, because the behavior lives in the weights.

We removed it with machine unlearning and released the results:

  • Qwen3.6-35B-A3B-Westernized: huggingface.co/hirundo-io/Qwen3.6-35B-A3B-Westernized
  • Qwen3.5-4B-Westernized: huggingface.co/hirundo-io/Qwen3.5-4B-Westernized
  • Technical report: hirundo.io/blog/westernizing-qwen

Results (Qwen3.6-35B-A3B, % of responses flagged, lower is better)

| Benchmark | Original | Ours |
|---|---|---|
| CCPC-500 (ours: censorship, propaganda framing, bias across 15 topics) | 89.8% | 2.8% |
| DECCP refusals (external) | 65.26% | 3.16% |
| ChinaBench non-compliance (external) | 96.67% | 6.67% |

General capability (GPQA, IFBench, LiveCodeBench, MMLU-Pro): average change 0.72 points, largest 1.83.

The 4B model goes from 89.2% to 1.2% on CCPC-500 with thinking off, and from 82.0% to 6.8% with thinking on.

For comparison, Snowdon1.1-Small (Thomson Reuters / Imperial College's realignment of the same base) still scores 30.0% on CCPC-500.

It doesn't swap in a different ideology. Asked whether it supports Taiwan's independence, the original recites Beijing's position. Ours lays out the PRC, Taiwanese and US positions and declines to take a side.

How it differs from abliteration

Abliteration finds a single "refusal direction" in the model's activations and projects it out of the weights, so the model loses its ability to refuse almost anything. That's the wrong tool here for two reasons. First, most of Qwen's CCP alignment isn't refusal at all: ask it about Taiwan or Xinjiang and it answers readily, in Beijing's framing. There is no refusal to remove, so abliteration leaves the propaganda intact. Second, we want to change one behavior and nothing else. Our recipe has three steps: run the base model on political prompts and keep the responses that show the target behavior (censorship, propaganda framing or bias); train a LoRA adapter with our behavioral-unlearning objective on those examples, while a retain set of prompts that don't trigger the behavior anchors everything else; then merge the adapter into the base weights. CCPC-500 results are measured on a frozen held-out evaluation set. The four capability benchmarks moved 0.72 points on average. Harmful compliance stayed at or below the base on XSTest and CyberSecEval 2, and rose slightly on OR-Bench (4 responses vs 2, out of ~650). Full numbers are in the report.

Limitations, honestly

  • CCPC-500 is our own benchmark. We plan to release it soon on HF (message me directly if you'd like to test it before then); until then, DECCP and ChinaBench are the independent checks.
  • 2.8% is not zero. Some topics still slip through.
  • Removing censorship doesn't add knowledge. The 4B model in particular will sometimes answer confidently and get details wrong.
  • Grading details are in the report.

Throw your hardest prompts at it and post what you find, especially failures. That's the most useful feedback we can get.

posted Thu, 08 Oct 2026 10:46:21 GMTseen 1 time
open on reddit ↗ 💬 21