Meta Muse Glimmer brings local AI agents to consumer GPUs

Meta Muse Glimmer brings local AI agents to consumer GPUs

Meta is releasing Muse Glimmer beneath an Apache 2.0 licence for native AI brokers that may run on a client GPU.

The corporate’s  Superintelligence Labs has launched the 30-billion-parameter mannequin’s weights on Hugging Face. Meta says builders can use it for native coding, operate calling, native brokers, and LLM-as-a-judge analysis.

The discharge targets an operational constraint going through AI groups: cloud-hosted fashions want community entry and central infrastructure. Meta as a substitute pitches Muse Glimmer for workloads that require an on-device mannequin, together with private brokers with entry to schedules, messages, recordsdata, and different non-public context.

Meta Muse Glimmer leads a number of agent activity benchmarks

Meta’s benchmark assessments put Muse Glimmer forward of Gemma4-31B and Qwen3.6-27B on 5 of eight general-agentic benchmarks. The mannequin scored 75.5 on MCP Atlas. Gemma4-31B reached 54.2, and Qwen3.6-27B recorded 62.5.

DeepSearch QA follows an identical sample. Meta stories a rating of 74.6 for Muse Glimmer, in opposition to 61.7 for Gemma4-31B and 71.1 for Qwen3.6-27B. The provided announcement identifies each benchmarks as assessments of an agent’s capability to work inside scaffolds and full multi-turn requests.

The mannequin scored 23.5 on τ²-Banking. Gemma4-31B recorded 15.1. Qwen3.6-27B reached 16.7.

Muse Glimmer additionally posted 47.6 on WildClawBench, forward of Gemma4-31B’s 37.6 and Qwen3.6-27B’s 43.2. Its GAIA2 outcome reached 43.3, in contrast with 36.4 and 40.0 respectively.

Different agent scores favour Qwen3.6-27B. Meta’s desk provides that mannequin 1,141 on GDPval-AA, in opposition to Muse Glimmer’s 953 and Gemma4-31B’s 811. Qwen3.6-27B additionally led SkillsBench with Expertise at 46.6, the place Muse Glimmer recorded 44.3.

OSWorld-Verified produced the biggest hole on this group. Meta stories 75.6 for Qwen3.6-27B. Muse Glimmer reached 65.9, and Gemma4-31B scored 58.5.

These assessments measure constrained duties. They don’t reveal how an area agent will behave after an organisation connects it to its personal recordsdata, calendars, messaging techniques, or inside instruments.

Coding outcomes cut up between Muse Glimmer and Qwen

Muse Glimmer’s coding outcomes present a narrower comparability. The mannequin led SWE-Bench Professional with a rating of 51.2. Meta stories 36.9 for Gemma4-31B and 50.2 for Qwen3.6-27B.

SciCode produced an in depth outcome. Muse Glimmer scored 43.6, marginally above Gemma4-31B at 43.4. Qwen3.6-27B recorded 39.8.

Qwen3.6-27B led two different coding evaluations. It scored 77.2 on SWE-Bench Verified, in contrast with Muse Glimmer’s 76.0. TerminalBench 2.1 gave Qwen3.6-27B a rating of 60.7; Muse Glimmer reached 51.7, and Gemma4-31B posted 43.4.

A neighborhood coding agent does greater than produce code. It wants a scaffold that decides which repositories, terminals, take a look at environments, and instructions the mannequin could entry. Meta says Muse Glimmer helps OpenClaw and different agent-orchestration patterns, with customized scaffolds coated in its developer documentation.

An organisation evaluating the mannequin for software program work ought to outline the instructions and repositories accessible to the agent earlier than measuring activity success. The provided materials describes retry coaching for failed instrument calls. That behaviour requires controls over repeat makes an attempt, particularly the place a instrument can alter supply code or invoke an exterior system.

Multimodal scores favour Qwen in most assessments

Muse Glimmer accepts interleaved textual content and pictures via a devoted notion encoder. Meta says this design lets brokers interpret screenshots, charts, and paperwork as a part of a dialog.

The benchmark chart places Muse Glimmer forward on Charxiv Reasoning. Its rating reached 78.8, in opposition to 77.7 for Gemma4-31B and 78.4 for Qwen3.6-27B.

Qwen3.6-27B led ScreenSpot Professional with 76.1. Muse Glimmer recorded 75.4, and Gemma4-31B scored 75.9. The identical mannequin led OmniDocBench v1.5 at 77.8, in contrast with Muse Glimmer’s 75.8 and Gemma4-31B’s 72.5.

MMMU Professional produced smaller variations. Meta lists Muse Glimmer at 74. Qwen3.6-27B reached 75, and Gemma4-31B posted 73.

These outcomes matter for groups contemplating brokers that act on visible interfaces. A screenshot-reading mannequin can interpret what it sees, but native testing should nonetheless cowl permissions, show layouts, doc codecs, and errors returned by related instruments.

Security figures present decrease reported assault success than Qwen

Meta additionally stories two safety-related evaluations: CI Reminiscences and Siren AgentDojo. The chart makes use of totally different measures for every take a look at.

On CI Reminiscences, Meta lists a violation fee of 26.4 for Muse Glimmer and a protection rating of 64.8. Gemma4-31B recorded a violation fee of 12.1 with protection of 53.0. Qwen3.6-27B posted a violation fee of 53.4 and protection of 66.9.

The Siren AgentDojo outcome makes use of assault success fee and utility. Meta provides Muse Glimmer an assault success fee of 28.4 and a utility rating of 94.2. Gemma4-31B scored 25.6 on assault success fee, with utility at 90.8. Qwen3.6-27B recorded 40.3 and 92.7.

Common reasoning outcomes add context to agent claims

Muse Glimmer led 4 of six general-capabilities-and-reasoning assessments in Meta’s comparability. It scored 77.0 on IFBench. Gemma4-31B recorded 76.0, and Qwen3.6-27B reached 70.8.

The AIME 2026 rating was 94.7 for Muse Glimmer. Meta stories 89.2 for Gemma4-31B and 94.1 for Qwen3.6-27B. On AA-LCR, Muse Glimmer reached 80.0, forward of 68.3 and 73.3.

The mannequin additionally led Beam 128K at 65.1. Qwen3.6-27B scored 63.0. Gemma4-31B recorded 58.2.

Gemma4-31B led GPQA Diamond with 85.7. Muse Glimmer scored 83.5, adopted by Qwen3.6-27B at 84.2. Gemma4-31B additionally took the highest rating on Humanity’s Final Examination, Textual content No Instruments, at 23.6; Muse Glimmer reached 22.0.

One mannequin doesn’t lead each take a look at. Meta’s outcomes as a substitute present Muse Glimmer competing intently with two equally sized fashions throughout a blended set of agent, coding, visible, security, and reasoning evaluations.

Reminiscence limits form the native deployment design

Meta says a full-precision 30-billion-parameter mannequin would require greater than 55 GB of reminiscence. Muse Glimmer as a substitute makes use of roughly 4-bit weight quantisation, decreasing the language mannequin to beneath 20 GB.

That allocation leaves reminiscence for a KV cache. The mannequin additionally wants room for its notion encoder and a speculative-decoding drafter. Meta targets a 24 GB or 32 GB reminiscence envelope for these parts.

The corporate says the DFlash-based drafter proposes blocks of tokens for the primary mannequin to confirm in parallel. Meta says this speeds technology in contrast with customary token-by-token output and retains similar output high quality. The provided put up doesn’t embody token-per-second figures, immediate sizes, energy knowledge, or concurrency outcomes.

Meta examined its Ok-Quant-17GB model with the quantised DFlash drafter on MacBook M4-Max {hardware}, MacBook M5-Max {hardware}, and an RTX-5090. It describes the ensuing expertise as appropriate for fluid dialog and real-time agent interplay.

The general public weights can be found via Hugging Face. Meta says integrations with llama.cpp, MLX, and ExecuTorch will arrive within the coming days.

See additionally: Alibaba assessments new enterprise mannequin for Qwen open-source AI

Need to be taught extra about AI and large knowledge from business leaders? Try AI & Big Data Expo happening in Amsterdam, California, and London. The excellent occasion is a part of TechEx and is co-located with different main know-how occasions together with the Cyber Security & Cloud Expo. Click on here for extra data.

AI Information is powered by TechForge Media. Discover different upcoming enterprise know-how occasions and webinars here.