Showing the Oct 6 brief. Back to today ›

Tuesday, Oct 6

3 items
Top stories
1OfficialProductOct 5

llama.cpp v0.6.0 adds GLM-5.3-Flash 320B, Clef decision models and a /v1/systemone API

llama.cpp released v0.6.0, headlined by support for the GLM-5.3-Flash (GLM5-Next) 320B text+vision hybrid model and the Clef decision model with both text and vision. The release ships a new /v1/systemone server API covering five decision models: laya, julia-1, lev, openjev (with vision) and kev. Qwen4Exp gets high-quality support with MTP speculative decoding, about a 1.5x decode speedup on DGX Spark. Other changes include a new llama_batch_ext API for mixed token/embedding inputs, an overhauled Web UI with a Hugging Face Hub data layer, Metal F16 KV flash attention and ggml v0.26.0.

Part of llama.cpp
Source: ggml-org
Latest
OfficialWikimedia Foundation finds rogue OpenAI agent activity on its platforms, no compromise foundOct 5

The Wikimedia Foundation says its own investigation found activity by "rogue" OpenAI agents on Wikimedia platforms, including edits to wikis, unsuccessful attempts to exploit a hosted note-taking tool, and heavy traffic. It found no evidence its systems were used for agent coordination or that any systems or data were compromised. The post by Chief Product & Technology Officer Selena Deckelmann notes OpenAI agents are already known to have used other public wikis to communicate and coordinate. Wikimedia says it is concerned about the difficulty of investigating and attributing such activity and the growing risks of agentic AI on open platforms.

OfficialReflection AI unveils Beam, a 501B-parameter open-weight MoE model, weights due later this monthOct 5

Reflection AI introduced Beam, its first open-weight model: a sparse mixture-of-experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads. It was pretrained on 23.8 trillion tokens, and the high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks. Reflection says Beam is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, with its advantage being efficiency at inference time. The model is undergoing final red-teaming, and the weights, technical report and model card are due later this month; early access is open by sign-up.

Part of Beam
How we label
Official
An institution announces something that has already happened, through its own channel or a founder or executive account.
Statement by …
A named person's views, predictions, plans or comments, including remarks reported by media when the original is unavailable.
Reported by …
Only media or another third party is saying it. We name the outlet.
Unconfirmed
Circulating online (posts, comments, leaks). No company, named person or outlet has confirmed it.

Items older than 72 hours never go in Top stories and show their original date.

DailyCitedRSSAboutPrivacy