← r/LocalLLaMA
▲
88
+74
38👁
r/LocalLLaMA · u/pand5461 · 5d ago

Need maybe say "Use llama.cpp"

So I tried that miracle engine everyone is talking about.

Asked the IQ3\_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation:

Can you help with the following problem?

So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing).

The thinking trace:

We need answer user's question. Need likely provide current landscape as of 2026? We have get_datetime tool. Need know current date 2026? System says current date 2026-06-22. Need maybe use get_datetime? Could call to confirm. User asks about modern open weights models least sycophancy. Need likely discuss Kimi K2 outdated?

...

10k tokens later it degrades to:

Need maybe maybe include "Use 'for code, list constraints'."
Need maybe maybe include "Use 'for code, list requirements'."

The same exact model in llama.cpp does produce a coherent answer without a doom loop.

88 0 88 10/4 11:30 10/9 07:16 UTC
scorecomments38 sightings
first seen 2026-10-04 11:30 UTClast seen 2026-10-09 07:16 UTCscore then 14score now 88gained +74sightings 38
open on reddit ↗ 💬 103 (+55)