Not on main yet I think (as of ~2:50PM UTC on 2026-08-26) – there’s a link to the PR for it in my other comment though. Unsloth’s fork has that integrated (they submitted the PR). I wouldn’t be surprised if something lands quickly in main, but this is a new architecture so may take a bit for people to figure out how to get the most out of it – bunch of discussion about e.g. SSD offloading for the ngrams and stuff like that in the github thread.
I’m pretty excited about this one, is it supported by llama.cpp main yet?
Not on main yet I think (as of ~2:50PM UTC on 2026-08-26) – there’s a link to the PR for it in my other comment though. Unsloth’s fork has that integrated (they submitted the PR). I wouldn’t be surprised if something lands quickly in main, but this is a new architecture so may take a bit for people to figure out how to get the most out of it – bunch of discussion about e.g. SSD offloading for the ngrams and stuff like that in the github thread.
llama.cpp support for it got merged a few hours ago 🏆
Nice! Hopefully I can figure out how to actually get it to load tomorrow… (It keeps getting OOM-killed when I try.)