Showing the Oct 10 brief. Back to today ›

Saturday, Oct 10

11 items · Updated 12:38 PM ET
Top stories
3OfficialCompanyOct 9

Ai2 replaces its priority-based GPU scheduler with time budgets and fair-share allocation

Ai2's AI infrastructure team says it replaced a priority-based scheduler for its clusters of thousands of NVIDIA H100, B200 and B300 GPUs with a system of GPU time budgets, hierarchical fair-share allocation and a time-slicing contract. With requests for two to three times more GPUs than it has, Ai2 says the change turned decisions about how much compute each research project gets into a transparent budgeting process while keeping clusters fully occupied.

Part of Ai2
4Reported by TechCrunchCompanyOct 9

a16z partner Olivia Moore says consumer AI needs revenue beyond subscriptions and API charges, per TechCrunch

Andreessen Horowitz partner Olivia Moore, who this week released a report on the top 100 consumer AI apps, told TechCrunch she sees a large opportunity in consumer AI if companies can earn money in ways other than subscriptions and token usage. Her report finds ChatGPT still far ahead, with smaller players such as Suno and ElevenLabs showing staying power, and lists about half a dozen consumer categories AI has not yet touched.

Part of a16z Top 100
Source: TechCrunch
5Reported by The VergeProductOct 9

Text-message AI agent Instinct holds its own against Muse and Dots in hands-on testing, per The Verge

Instinct, a startup AI agent that launched in August invite-only with no marketing and became one of the most talked-about agents for its text-message interface and chores like booking DMV appointments, now faces similar agents Muse and Dots from far larger companies, The Verge reports. In the reviewer's testing, Instinct handled online tasks about as well as the two Big Tech rivals.

Part of Instinct
Source: The Verge
Latest
OfficialRecraft publishes a guide to building character reference sheets from one image in Recraft StudioOct 10

Recraft published the second part of its character-design series, showing how to turn a single hero image into a turnaround, an expression sheet and a pose sheet in Recraft Studio. The guide says a reference sheet keeps a character on-model across new images, avoiding the usual drift in haircut, clothing details and colors when an image model is asked for the same character from another angle.

Source: Recraft
UnconfirmedDeveloper says Aplomb 1 tops Hugging Face's Typed Decisions benchmark for probabilistic decisionsOct 9

A developer says their model Aplomb 1 ranks first on Typed Decisions, a Hugging Face benchmark that gives a model one piece of unstructured text and five typed questions about it at once. According to the post, Aplomb 1 has the best KL and Brier scores on the leaderboard, ahead of a 27B model, and the highest accuracy among models under 6B. The claims were not independently checked.

Source: r/LocalLLaMA
UnconfirmedUser tests Mellum2.1-12B-A2.5B as a local coding agent and finds it usable but weak on one-shot projectsOct 9

A user ran Mellum2.1-12B-A2.5B at Q8 locally with the Pi coding agent and llama-server on five one-shot tests. According to the post, results were mixed, with poor output on a pelican SVG, a browser OS and a Minecraft clone, though the model was described as surprisingly usable overall. The results were not independently checked.

Source: r/LocalLLaMA
UnconfirmedDeveloper shares a 24KB stand-alone HTML language model that generates short stories in the browserOct 9

A developer shared a 24KB self-contained HTML page that runs a tiny language model in the browser and generates consistent short stories from random seeds. According to the developer, it can exceed 60 tokens per second even on a smartphone, though story quality depends on lucky seed draws. The claims were not independently checked.

Source: r/LocalLLaMA
UnconfirmedMaintainer of a widely used open-source local MCP server releases version 2Oct 9

The maintainer of an open-source MCP server that began as a Reddit side project 18 months ago says it has grown into one of the most popular MCP servers and has just released v2. The project runs locally and is open, and the maintainer says they already struggle to keep up with pull requests and issues. The project was not named in the material available, and the claims were not independently checked.

Source: r/LocalLLaMA
UnconfirmedUser runs a Qwen3.8-Flash-Next variant at 20-30 tokens per second on 12GB VRAM with a custom Strata forkOct 9

A user says they ran Qwen3.8-Flash-Next-GSQ-RCO-Abliterated at IQ3_S on 12GB of VRAM and 32GB of system RAM with NVMe offload, reaching 20-30 tokens per second decode (39-45 at Q2) and up to about 90,000 tokens per second prefill at 131k context. The setup uses a heavily modified, experimental fork of Strata. The results were not independently checked.

Source: r/LocalLLaMA
How we label
Official
An institution announces something that has already happened, through its own channel or a founder or executive account.
Statement by …
A named person's views, predictions, plans or comments, including remarks reported by media when the original is unavailable.
Reported by …
Only media or another third party is saying it. We name the outlet.
Unconfirmed
Circulating online (posts, comments, leaks). No company, named person or outlet has confirmed it.

Items older than 72 hours never go in Top stories and show their original date.

DailyCitedRSSAboutPrivacy