← r/LocalLLaMA
▲
0
-4
7👁
r/LocalLLaMA · u/rm-rf-rm · 29h ago

Cloudflare Clef Experience

Using llama.cpp 0.6.0 and bartowski's Q4_K_M quant for Cloudflare Clef

Running the basic example:

curl http://127.0.0.1:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "I was charged twice this month, please refund one of them.",
"questions": {
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorised use"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"]
}
}
}'


getting a lousy output below. The noul is just 0.57. When tested with Jev, as expected the noul is 0.99.

Lost all faith after such a poor result to a first basic question. Has anyone else had better luck?

{
"model": "models-gpt/cloudflare_clef-Q4_K_M.gguf",
"answers": {
"refund": {
"type": "noul",
"noul": 0.5733821642835288
},
"team": {
"type": "choice",
"choice": "billing",
"probabilities": {
"billing": 0.6760871850733896,
"fraud": 0.17096718668994926,
"technical": 0.15294562823666114
},
"confidence": 0.5141307776100844
},
"urgency": {
"type": "score",
"score": 1.1166479856366882,
"legend": {
"0": "Can wait",
"1": "Today",
"2": "Blocking the customer now"
},
"probabilities": {
"0": 0.23199687091164825,
"1": 0.4193582725400152,
"2": 0.34864485654833655
},
"confidence": 0.1290374088100228
}
},
"usage": {
"input_tokens": 334,
"output_tokens": 0
}
}

4 0 0 10/8 07:28 10/9 07:36 UTC
scorecomments7 sightings
first seen 2026-10-08 07:28 UTClast seen 2026-10-09 07:36 UTCscore then 4score now 0gained -4sightings 7
open on reddit ↗ 💬 19 (+11)