How to create my own benchmarks
Since the begin of the year I wanted to create some personal benchmarks, but I haven't seen a proper guide for it. I usually don't like making "spam posts", but since the purpose of my benchmarks is quite specific I've decided to make a post anyway. To put it short the purpose of these benchmarks is to test the models on some reasoning puzzle games with the purpose of monitoring the reasoning traces (which can be quite painful if more than 8k tokens) so it can't be anything automated where an answer like a, b, c is simply accepted. I will use magic the gathering as examples, since I was pretty good at it back then and this is not what my benchmarks are based on and I can't recommend anyone making some on it since the game has now probably around 30k cards (not sure if unique tho). My main concerns are: The Process: - I have already created a set of 20 questions, which need some refinement.- I will test all the question locally using llama cpp, llama-cli with a few modifications to track the tokens used. This doesn't change.- I can only test what runs on 16gb vram and 96 gb ram.- I will test mostly Unsloth quants, especially for non q8 quants, but I might use uncensored models as well if they are q8 I suppose. - I will use the recommended settings provided by Unsloth for each models. (I have no idea how the formatting bugged like that)- 16k context budget. I don't think I can get a higher context for every model without using lower quants. I think 12k reasoning budget, with xhigh effort.- the question will be given as the model loads using the --file parameter. That will definitely make a lot of reads on my ssd, so I might find another way, maybe just create a special command to give the benchmark file, or just use "\n" and give it as prompt and use /clear.- I will post the actual benchmarks questions/answers on a github page (or whatever we will be allowed to use at that time) after a few models pass the benchmark fully. Concerns: I originally wanted to see if the models can answer 1 time out of 10 tries and move to the next one after it answered. The justification is that I care to know if the models can eventually figure it out since it's difficult to answer these questions, is not something like 2+2=4 to require 100% accuracy, however, I believe that 10 is a bigger number, especially since qwen models casually think for 8k tokens (when they actually decide to stop), so I am not sure if 5 is a good one, or simply put how lower the number can be to actually make the benchmark "valid". should I read the entire thinking tokens? I think I should, but that might be just overkill (for me). For example I tested 2 questions from the benchmarks and there was a point when it said "there are 12 creatures" and while it was completely unrelated it started counting "3+2+1+1+3+2=11 no, let me count again..." and I've lost it at that point. The whole 8k reasoning ended with "I don't know it's 50/50 but I must give an answer", all of that while mistral (non reasoning) answered in 700 tokens. uncensored vs "original" models, not sure if the uncensored ones will produce much worse results. the models have some decent knowledge about the niche game I have based my benchmarks on, but at some point mistral 24b said "tap this creature, activate ability. Can I activate the ability twice in this variant?" or it thinks that "if I use duress on my opponent and he gets hexproof the target will bounce to the next opponent". Probably not best example, but I think I should put in the prompt that "abilities trigger only once, and it doesn't resolve if it failed", which leaves the question "how much additional information I must provide without spoonfeeding the models"? I think I should try to help the models where they are not completely aware of the rules, so I think I should refine the questions based on how the first run goes. I want to provide some stats like: how many tokens for answer, speed t/s, number of tries until correct answer (with token count for each failed question as well, separately). I will also add some observations and eventually track other things like "common sense" even if the answer was wrong. I did use the flags "-DGGML_CUDA_FA_ALL_VARIANTS=ON" and "-DGGML_CUDA_FA_ALL_QUANTS=ON" and I am not sure if this will quantize the kv cache (I don't intend to), but Qwen3.8-27B-UD-IQ4_XS runs with mproj on probably slightly more than 32k context so I am not sure if this is some architectural improvement or just me messing things around without knowing what these flags do, although I suspect the context is being put on the cpu instead. one of the first models I downloaded was qwen 3 14b reasoning, and q8 was about 15gb vram, so I downloaded q6 instead. Do these models have the "context baked in"? Like 16k or 8k context? qwen 3.8 makes it even weirder for me... llama cpp bugs. I know at least at some point the /clear cmd removed the system prompt, not sure if on the recent versions still does. It seems that whenever I used /regen mistral gave me a worse response even though from first try it gave the right answer. It could have been rng tho... If I was to post the results of these benchmarks in late 2024 - early 2025 everyone would have been ecstatic for some "trust me bro" benchmarks that try to figure out if the models can "reason", that basically have no real use case, but well, ngreedia doesn't support innovation. Anyway, the real question is if it is worth posting the results (statistics basically) with some sort of explanation. the benchmarks, since this is as I said only some "trust me bro" thing since I can't provide the way to reproduce them, and are based on what I believe to be something the models haven't been trained on, and nobody else made something like that publicly (which I've seen some people already made), which all of these combined lead to a very heavy assumption that might not actually be true (the assumption being that similar questions haven't been asked to lead the models being benchmaxed on, like the strawberry one). quantisations: it was easy at some point, but then iq came and now ud as well, which makes me wonder for gemma 4 31b UD-IQ3_XXS vs Q3_K_S, the fact that UD-Q2_K_XL has the same size as the UD-IQ3_XXS makes things even more confusing. I mean I know that the quants are there to fit in specific vram sizes, but what's the point of making multiple quants have the same size? Anyway, I decided to go with UD-IQ3_XXS, I guess I will leave this here as a rant... Is the --jinja flag important? I mean I assume that for the older models llama cpp updated their templates anyway. Observations: (this part is basically useless, but funny)- the benchmarks are based on giving information like: the current cards (hand, graveyard, battlefiend) and the events that led to the current board (or game) state. The more information I add (even if not noise, but only repetitive to explain how the things went that way, something needed for some specific questions) makes the answers worse, for both reasoning and non reasoning, making the reasoning models reason about the noise instead of focusing on the important hypotheses. - one question was something like "both players play lands, no spells or abilities, then the opponent reanimates Gigantosaurus from their graveyard". At that question qwen 3.8 with the reasoning off can't say proper answer "he cheated". When I asked "how did the creature get into his graveyard" it answered "that creature was not put into the graveyard to start with, in fact he couldn't even reanimate it." then it explains all the rules properly and then gives other answers completely avoiding to say anything along the lines "he cheated". I did more testing on steering them and eventually both models answered correctly. Also I loaded the qwen (no reasoning) script by mistake when I tried to get a translation and it answered from the first try, but it had to ask itself the question I asked to steer them "but was that creature put into the graveyard". What I try to say by this is that all these "but wait " questions are only for them to find the right question that will steer them towards the answer. - I guess the last one. We are all aware of the caveman reasoning patterns and other behaviours they took from us, for example the models start using abbreviations like "ETB = enters the battelfield", "P1 = Player 1" and so on, but if "ETB" and "enters the battlefield" both have the same amount of tokens (I didn't test) isn't it non beneficial for them to use that wording? Also seen gemma saying "BUT IT DOESN"T MAKE SENSE!" all in caps, and one more thing I've forgot, but I am not sure if they benefit from the behaviours they have learnt from us. I will be honest, I have slacked a lot on these benchmarks, mostly because I enjoyed more the "diffusion models" since I can run most of them especially the fp8 ones. I am aware that the intelligent LLMs start from dense 30b or even 70b (if any nowadays lol) and I originally tried to go for 24-48gb vram but things didn't work that way for me. All these things combined made me very disappointed in llms. I mean they are useful, I've got what I wanted from them on my daily usage, but I've got a bit bigger plans with them, and with the current hardware limitation (or availability) it just demotivates me to do anything with them. Even for this post it took me 2 weeks to decide on eventually posting it after saving the original draft.