Some recent decision models <=3B on internal benchmarks vs fine-tuned embedding model
Before anyone has the change, yes, I know that the classifier models may be able to better generalize. But in my case with a dataset of \~10k rows and general task routing based on a sleuth of customer service inquiry/interaction to the appropriate service/task, it basically covers the whole range of what I can think of and what I could find online already. Very surprised from my results to see large decision models perform worse than smaller ones. (especially the liquidai ones)
|Rank|Router|Accuracy|Macro-F1|Balanced accuracy|Mean latency|
|:-|:-|:-|:-|:-|:-|
|1|Qwen Embedding 0.6B Baseline|0.7865|0.7604|0.8422|—|
|2|Jiwo-0.8B|0.7027|0.6771|0.7989|56.30 ms|
|3|D1-Omni-600M|0.7054|0.5998|0.5803|16.68 ms|
|4|D1-3B|0.5568|0.5915|0.7425|26.00 ms|
|5|Laya|0.6081|0.5295|0.5914|27.71 ms|
|6|Decider-2B|0.4811|0.5160|0.7055|59.39 ms|
|7|Decision 2.0 Sol 2B|0.4149|0.4722|0.6955|60.21 ms|
Any tips, follow ups and criticisms well appreciated from the community!