As of Oct 5, llama.cpp v0.6.0 ships GLM-5.3-Flash 320B support, the Clef decision model with text and vision, and a /v1/systemone API covering five decision models.
llama.cpp released v0.6.0, headlined by support for the GLM-5.3-Flash (GLM5-Next) 320B text+vision hybrid model and the Clef decision model with both text and vision. The release ships a new /v1/systemone server API covering five decision models: laya, julia-1, lev, openjev (with vision) and kev. Qwen4Exp gets high-quality support with MTP speculative decoding, about a 1.5x decode speedup on DGX Spark. Other changes include a new llama_batch_ext API for mixed token/embedding inputs, an overhauled Web UI with a Hugging Face Hub data layer, Metal F16 KV flash attention and ggml v0.26.0.
Items older than 72 hours never go in Top stories and show their original date.