Fly to Taipei. Round-trip from Orlando: $1,081. Go to the largest retailer in Taiwan to Spend NT$129,990 ≈ US$4,093. Hang out in Taiwan for two weeks. Eat good food. Touch international grass. Fly home and flex on r/LocalLLaMA.
Fly to Taipei. Round-trip from Orlando: $1,081. Go to the largest retailer in Taiwan to Spend NT$129,990 ≈ US$4,093. Hang out in Taiwan for two weeks. Eat good food. Touch international grass. Fly home and flex on r/LocalLLaMA.
Disclaimer: no AI was used whatsoever to write this post
Cautionary tale about chasing cheap tokens.
exposé: https://kendell.dev/blog/crofaifalse/
reaction by nahcrof, announcing the shutdown of the service: https://x.com/nahcrof/status/2099552389434900643 - now deleted, archive picture: https://i.imgur.com/teOQngH.png
NahCrofAI (crof.ai, nahcrof.com) was an inference provider which had all the latest models at the cheapest price, often significantly below the lowest alternative on OpenRouter. The owner claimed that they are running custom inference engines that allows them to offer tokens for dirt cheap, and other providers are suffering from "skill issues", that's why they are so expensive.
In reality:
kimi-k3 are sold at $2/$10 in/out, but instead routed to GLM 5.3 Flash via OpenRouter, representing a 13.3x multiple on input, and 20x multiple on outputgreg-2-ultra routes to GLM 5.2, greg-1-mini routes to Qwen 3.5 9B. greg-2-super, greg-1, greg-1-super routes to Kimi K2.7 Code. All of these at a significant markup compared to the actual model being served. CrofAI admits in DMs that his claims of the greg family being made by him is a lie.CrofAI responded to the exposé by announcing the shutting down of their service; after their failure to provide their own inference, they promise to provide one last thing: a refund to those asking.
UPDATE: around 4:30 AM UTC of Sept 15, the owner published a now-deleted blog post (archive image) writing under the fake pretense that it's his "team" authoring it, stating all of CrofAI founder's claims "were written under a lot of stress, and they described the situation as worse it was", and that a new team is taking over, with the service being resumed in 2 weeks.
At the same time, the CrofAI twitter account was also supposedly "taken over" by the team, starting each twitter reply with "Hey, Nathan here", stating the founder is stepping back and a "team" is taking over everything. This fake pretense act only lasted a few hours, and scared either by the public not buying the Nth fake story of the pathological liar that CrofAI is, or by the public's replies reminding him that what he committed is numerous counts of wire fraud, he has now deleted all his online presence: nahcrof.com and crof.ai return 404, Twitter page is deleted, /r/CrofAI sub is now private.
Here is another image of the owner admitting that he was defrauding customers for the entire 2 year operation of his service, then begging the investigator to help him cover his tracks and not expose him
EDIT: Commenters pointed out that NahCrof is 4chan in reverse. The owner's Discord name was "Devious Flimflam". Flimlam is defined as "deception, fraud". Looks like it was a deliberate scam operation from the get-go, and the owner's age was among the many lies.
I cannot stress this enough: if you bought any credits (even if you used them up) you are entitled to a full refund for every transaction as the victim of fraud. Open a chargeback with your bank for every transaction made. If you used their API, assume that everything was logged and is currently being mined for personal information and API keys to sell on the black markets. Rotate your keys, change passwords, get a new debit/credit card.
Two months ago I announced I had found a new dynamic quant method called Voodoo Quant which was SOTA for the most aggressive quant levels on some smaller Qwen3.5 GGUF models. I kept the methodology private at the time, but I've seen too many requests for dyn quants for various models lately, so I decided to give my method to the community since I don't have the time to scale this into something that could do it justice. Hopefully it will also inspire some researchers to find out more about it and improve it as I am just scratching the surface.
Here is the new toolset so you can now make your own dynamic quants: https://github.com/curvedinf/voodoo-dyn-quant
Many postulated on what method I was using, and its actually fairly simple and elegant: I found a way to use gradient descent to optimize the per-tensor quant layout.
What is a Dynamic Quant? Some model formats, namely GGUF, support quantizing (compressing) each tensor (set of weights) with a different quant level. Static quants make static selections of certain types of tensors having a set quant level. Dynamic quants make a different quant selection for each tensor of each checkpoint size.
How does Voodoo Quant work? Voodoo Quant runs all the quant levels of a model at the same time, for every tensor, and lets gradient descent pick which ones optimize loss the lowest for a given target filesize. Technically speaking, this is done by an epoch of training which freezes all candidate quant weights (as provided by conversion directly from llama.cpp's underlying library, gglm) and only trains a single scalar gate per tensor per quant level. The scalar gates of a tensor represent which quant levels are most optimal. Over time a tau level is annealed that helps the training freeze into singular predominant quant selections for each tensor instead of mixtures. Softmax is used so all quant levels receive gradient, even when a selection is mostly frozen. The quant selections are trained on a diverse calibration dataset. The training is then measured with a loss function which finds the KL divergence of the mixed-quant logits versus the reference BF16 checkpoint, rewarding a lower KLD, while also rewarding getting closer to a provided filesize target. This info should get you started on understanding what is going on, and for more details you can dive into the source!
What does the repo have? A complete set of tools to train your own dynamic quants using this methodology. It is currently set up for Qwen, but it can be adapted quickly for any model arch.
How does UD 3.0 compare? Unsloth Dynamic 3.0 is a proprietary methodology that unsloth has not revealed any details of (by the way, people were criticizing me for not revealing my methodology, but unsloth had been doing that for years!). However, we do know it is very good. In my testing, UD3 is better than VQ at high to mid quant levels, but VQ is better at aggressive levels. As far as I can tell, UD 3.0 is an advancement of static analysis techniques that are currently defacto. Static analysis means the weights of a model are analyzed in various ways using statistics and static functions, sometimes tuned by repeated runs benchmarking KLD and other metrics. Voodoo Quant is the first method to my knowledge that uses a backwards pass and gradient descent to choose per-tensor quant levels. Using GD to optimize quant levels requires a much more powerful system than static analysis, but technically speaking is more efficient at maximizing performance because it compares the equivalent of many more iterations of benchmarking runs than is reasonably possible via SA.
How well does Voodoo Quant work? This is a research grade project, and is not studied at larger model sizes. At smaller model sizes it is shown to be exceptional, as in the charts above, especially at the lowest quant levels which can benefit from more complex/diverse quant selections. I used research level control for my testing, but I don't claim that VQ has been studied to a scientific level of proof of effectiveness. A lot is still left to learn about how well it works, so I hope to see more research in this direction. I don't believe there are many dynamic quant open source projects out there, so I hope the community can use this to improve local models, and especially for low VRAM machines.
Why open source now? I have like a dozen irons in the fire for various other projects, and this is just sitting there when it could be used by the community. I have made many open source projects for 20 years, so its nothing new.
Peace!
Maybe some of you know but I didn’t see any post about this. Apple just made available their AFM model on MacOS 27 natively. Just run fm chat in a terminal.
Disclaimer: I’m an open weight person. I prefer open models and ecosystem, but I’ll still open the discussion.
Did you test them? Build using them? Are these models good?
I feel like this is still a huge step in the direction of local AI that a company like Apple does this and release hardware optimized models.
So what do you think?
EDIT: Sorry for the unclear title. This model is UkisAI's Swift-Qwen3.8-27B, not a new version of BottleCap AI's 3.6-ThinkingCap. All credit goes to UkisAI for making great fine-tune, and I made this post to celebrate their work. I meant no disrespect by mentioning another model in the title.
I doubt I'm in the minority here when I say I love Qwen models, but the overthinking is a major timekiller. It was bad in 3.6-27B, and it's worse in 3.8. I know there are some who say, "well that's how it achieves such a good performance/size ratio"... But now there's some definitive proof that's not the case: UkisAI's Swift-Qwen3.8-27B!
This model seems to be inspired by Qwen3.6-27B ThinkingCap, which was the version of 3.6-27B I used as a daily driver before switching to the 3.8 series. For those of you who haven't heard of it, ThinkingCap is a fine-tuned version of 27B that uses about 40% less tokens to accomplish comparable benchmarks and general performance as the original model. It's one of those fine-tunes that actually works. I used it daily for months without any issues, and it saved me countless hours.
I had been waiting and hoping that they would release a similar version of 3.8, because it is so slow, despite its impressive performance, but so far none has been forthcoming. However, it looks like UkisAI also enjoyed that model, and took it upon themselves to deliver a sequel. They identified "reasoning-marker tokens that ... trigger overthinking in Qwen’s reasoning rollouts" and penalized them using RL, resulting in fewer overthinking errors. They also employed "a transfer component derived from BottleCap AI's ThinkingCap-Qwen3.6-27B". The end result is an average of 30-50% fewer reasonign tokens for the same quality outputs on a number of benchmarks (see the model card for all of them).
This claim is quite impressive, and I have independently verified their claims and the quality of the model in my own use cases and in coding benchmarks using Aider as an eval suite (with Q8_0 for both models):
|Metric|Swift-Qwen3.8-27B|Qwen3.8-27B|
|:-|:-|:-|
|Pass1 (%)|30.8|27.1|
|Pass2 (%)|75.7|77.6|
|Well-formed diff (%)|98.1|99.1|
|Completion tokens|7,301|12,547|
|Seconds/case|750|1,481|
|Total tokens/solve|12.1k|19.3k|
As you can see, their claims hold true -- Swift accomplished an equivalent success rate in approximately half the time, using 63% of the tokens! This is a huge win for 3.8-27B users, because of course decode drops off more and more the longer the response gets, which is why the time is halved even though the tokens are closer to two-thirds of 3.8-27B.
Anyways, my posts tend to get excessively long so I'll cut it off here, I was just really excited after finishing my eval suite on this model and wanted to share.
Hey r/LocalLLaMA,
We’ve released our full ShapeLearn GGUFs for Qwen 3.8 27B.
TL;DR
Lite held up very well. As we expected.
We released ShapeLearn-Lite quants a couple of days after Qwen arrived: less optimization, targeted sanity checks, full benchmarking after release.
Then Unsloth v3 arrived with lower KLD at several comparable sizes. Lite looked overtaken, until the task results came in. Three of six Lite models made the quality/speed frontier against twelve Unsloth v3 models in our RTX Pro 6000 comparison. Pretty good for an impatient release. Full ShapeLearn now pushes that frontier further.
Which brings us to KLD.
Unsloth Dynamic V3’s UD-IQ3\_S had \~20% lower KLD than our similarly sized smallest Lite model, but scored 95.55% versus Lite’s 97.33% of BF16’s aggregate benchmark score.
Closer token distributions did not mean better task performance. KLD is useful to avoid a quant that has fallen over the edge, but it isn’t a quantization leaderboard.
That distinction is the subject of our paper on KLD and quantization fidelity, recently accepted for publication to the EMNLP 2026 Industry Track. We also released blog post version of the paper a few weeks back.
We benchmarked this release on RTX 6000 Pro Blackwell, RTX 5090, RTX 4090, RTX 3090, RTX 4080 and RTX 5060 Ti. The benchmarks we used to measure quality are: GSM8K for math, IFEval for instruction following, MMLU for general knowledge, LiveCodeBench V6 for coding, Multi-IF for multi-turn and multilingual instruction following, ACEBench for tool use and agentic tasks (both thinking and instruct), Multiple HumanEval for coding (thinking) and BFCL V4 for tool calling and agentic tasks (thinking).
If you want to dive deeper or choose the best model for your use case, the blog has the complete results across all tested GPUs, along with the methodology, model sizes, and full legend.
I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be faster) through 100K tokens, with context of up to 240K.
If you want the repo, it is here:
https://github.com/JakeATX/llamAmpere
I recommend running with this quant, which is \~ 4 K M quality but considerably faster (technically, a 3 K XL upgrade)
https://huggingface.co/jakeatx/Qwen3.8-27B-ATX-IQ4\_XS-M-GGUF
If you want the deep dive on how it is so much faster (80% vs the near comp at 200K!), at more context, there is a long form article here.
https://x.com/JakeKAllDay/status/2095646450138874095?s=20
Running faster than API speeds on my 3090 (for 27b at least) has genuinely been a step change in the utility of the model + card for me. I hope you enjoy it!
I am a beginner, Took a while to get started, get everything right.
This setup is native not container. Still not sure if I did this right, or if I can tune this more.
Environment=HF_HUB_OFFLINE=1 Environment=VLLM_LOGGING_LEVEL=INFO Environment=VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 Environment=PATH=/home/suryakiranc/vllm/.venv/bin:/usr/local/cuda/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin Environment=CUDA_HOME=/usr/local/cuda ExecStart=/home/suryakiranc/vllm/.venv/bin/vllm serve unsloth/Qwen3.8-27B-NVFP4 \ --served-model-name unsloth/Qwen3.8-27B-NVFP4 \ --safetensors_load_strategy prefetch \ --tensor-parallel-size 4 \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_xml \ --enable-auto-tool-choice \ --gpu-memory-utilization 0.91 \ --kv-cache-dtype fp8 \ --max-num-batched-tokens 16384 \ --speculative-config '{"method": "mtp", "num_speculative_tokens": 3}' \ --mm-encoder-tp-mode data \ --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' \ --max-model-len 1000000 \ --host 0.0.0.0 \ --port 8000
So I would like to run qwen 3.8 27b locally for my ai agents, maybe even in parallel with other small LLMs (such as qwen3.5 9b, oss 20b etc..).
What is the best hardware to do this? Not a video card but I mean as “mini pc ai”.
Thank you 🙏
For years, the JVM has watched the AI revolution from the bench. Every model, AI framework, every breakthrough, built with/for Python.
jinfer is an inference engine built for the JVM from first principles: chat, vision, audio transcription, embeddings, reranking, and TTS. No Python runtime, no ONNX, no sidecar process, no wrappers; the whole stack is built for the JVM:
It integrates with Spring AI and LangChain4j, and has first-class support for GraalVM Native Image.
Where things stand: this is an early release. CPU is the main target today, and is already competitive with llama.cpp. GPU support via jota is in progress.
Runnable examples + benchmarks: https://qxotic.ai
Jinfer (Apache 2.0): https://github.com/qxoticai/qxotic/tree/main/jinfer
PS: I'm behind it and also the author of llama3.java (2024) and gemma4.java