Meta is back with Muse Glimmer: local, agentic, multimodal, and open source
2026-08-20 · Hugging Face
Meta Releases Muse Glimmer: A 30B Local, Agentic, Multimodal, Open Source Model
Meta has introduced Muse Glimmer, a 30-billion-parameter multimodal model distilled from its predecessor Muse. Released under the Apache 2.0 license, it is specifically optimized for local agentic workloads, making it ideal for privacy-focused applications, cost reduction, and local experimentation.
Key Features
- Local-first design for privacy and reduced inference costs
- Agentic capabilities optimized for coding agents, document analysis, and personal assistants
- Native multimodal understanding for both images and videos
- Fully open source under Apache 2.0
The model is positioned for privacy-aware use cases including local coding assistants, document intelligence, and Claw- or Hermes-style personal agent setups.
Benchmark Performance
Muse Glimmer demonstrates competitive and often superior results compared to Gemma4-31B and Qwen3.6-27B (Thinking Mode).
General Agentic Benchmarks
| Benchmark | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B |
|-----------|------------------|------------|-------------|
| MCP Atlas | 75.5 | 54.2 | 62.5 |
| DeepSearch QA | 74.6 | 61.7 | 71.1 |
| τ³-Banking | 23.5 | 15.1 | 16.7 |
| WildClawBench | 47.6 | 37.6 | 43.2 |
| GAIA2 | 43.3 | 36.4 | 40.0 |
It also performs strongly on agentic coding tasks such as SWE-Bench Pro (51.2) and SWE-Bench Verified (76.0), as well as multimodal benchmarks including Charxiv Reasoning (78.8), ScreenSpot Pro (75.4), and MMMU Pro (74).
The model maintains solid safety metrics and excels in general reasoning benchmarks including AIME 2026 (94.7), GPQA Diamond (83.5), and AA-LCR (80.0).
Model Architecture
Muse Glimmer is a dense 30B model composed of:
- A 2B ViT-style Perception Encoder for vision
- A 28B parameter text decoder
An optional speculative decoding drafter implemented on DFlash is also provided. This module can substantially accelerate generation for structured outputs like code, at the cost of additional memory.
Text Decoder
The language model employs a hybrid attention pattern repeated 13 times for a total of 52 layers:
- Three consecutive sliding window attention layers (2,048 tokens) with Rotary Position Embeddings (RoPE)
- One full attention layer using NoPE (no positional embedding)
It also features:
- Gated Grouped-Query Attention: Each KV head is shared across 16 query heads, reducing KV cache size by 16×
- Q-K normalization + extra query scaling: RMSNorm is applied to queries and keys, followed by a scaling factor on queries to stabilize attention logits and control softmax temperature
Perception Encoder
Unlike typical small vision encoders in other VLMs, Muse Glimmer uses a substantial 2B-parameter ViT-like model based on Meta’s Perception Encoder architecture.
Images are patchified into shape `2 frames × 3 channels × 14 × 14`, projected linearly, and augmented with interpolated absolute positional embeddings. The encoder consists of 50 layers with GELU MLPs and follows the same `(Window Attention × 3, Full Attention)` pattern, using 2D RoPE.
After the transformer blocks, a pixel shuffle operation merges 2×2 neighboring spatial tokens, reducing token count by 4× while preserving channel information. The features are then projected into the text decoder’s embedding space.
Video Processing
Videos are processed frame-by-frame at 2 frames per second, with a maximum of 96 evenly sampled frames. The processor inserts timestamped placeholders such as `Time: 0.0s <|video|> x N`, which are replaced by the final video embeddings before the projection layer.
Ecosystem Integration
Meta has delivered day-0 support across the Hugging Face ecosystem. The model is compatible with:
- transformers
- llama.cpp
- vLLM
- Inference Endpoints
To get started, upgrade the libraries:
pip install --upgrade transformers accelerate
Both the main multimodal model and the speculative decoding drafter are supported via `AutoModelForMultimodalLM` and `AutoProcessor`.
Muse Glimmer is now available on the Hugging Face Hub for immediate local deployment and experimentation.