Need maybe say "Use llama.cpp"
So I tried that miracle engine everyone is talking about.
Asked the IQ3\_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation:
Can you help with the following problem?
So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bug noticing).
The thinking trace:
We need answer user's question. Need likely provide current landscape as of 2026? We have get_datetime tool. Need know current date 2026? System says current date 2026-06-22. Need maybe use get_datetime? Could call to confirm. User asks about modern open weights models least sycophancy. Need likely discuss Kimi K2 outdated?
...
10k tokens later it degrades to:
Need maybe maybe include "Use 'for code, list constraints'."
Need maybe maybe include "Use 'for code, list requirements'."
The same exact model in llama.cpp does produce a coherent answer without a doom loop.