llama.cpp

Product · 1 update · last updated Oct 5

Where it stands

As of Oct 5, llama.cpp v0.6.0 ships GLM-5.3-Flash 320B support, the Clef decision model with text and vision, and a /v1/systemone API covering five decision models.

Timeline · newest first
  1. Oct 5
    Official

    llama.cpp v0.6.0 adds GLM-5.3-Flash 320B, Clef decision models and a /v1/systemone API

    llama.cpp released v0.6.0, headlined by support for the GLM-5.3-Flash (GLM5-Next) 320B text+vision hybrid model and the Clef decision model with both text and vision. The release ships a new /v1/systemone server API covering five decision models: laya, julia-1, lev, openjev (with vision) and kev. Qwen4Exp gets high-quality support with MTP speculative decoding, about a 1.5x decode speedup on DGX Spark. Other changes include a new llama_batch_ext API for mixed token/embedding inputs, an overhauled Web UI with a Hugging Face Hub data layer, Metal F16 KV flash attention and ggml v0.26.0.

    Source: ggml-org
How we label
Official
An institution announces something that has already happened, through its own channel or a founder or executive account.
Statement by …
A named person's views, predictions, plans or comments, including remarks reported by media when the original is unavailable.
Reported by …
Only media or another third party is saying it. We name the outlet.
Unconfirmed
Circulating online (posts, comments, leaks). No company, named person or outlet has confirmed it.

Items older than 72 hours never go in Top stories and show their original date.

DailyCitedRSSAboutPrivacy