Skip to content
Muse Atlas
Explore Muse
Open weights

Muse Glimmer

AI on your hardware, close to your files and tools.

Concept illustration of Muse Glimmer

Your environment. Your configuration.

Muse Glimmer is Meta's open-weight model: its model files are available to download and run yourself. It understands text and images, is distilled from Spark—trained using the larger model—and is released under Apache 2.0 for workflows on your own hardware.

What Glimmer is good at

Keep inference local

Inference means running the model to produce an answer. Here, it runs on your hardware.

Work with private files

Choose which local documents and source files a workflow can access.

Connect local tools

Use your own tool loop and execution controls.

Shape your deployment

Choose a runtime and manage the model yourself.

Review code inside your own workspace.

Follow the model, files, tools and review stages of a local workflow.

Illustrative scenario

Your computer

  1. 1

    Model

    Load Glimmer in a compatible runtime—the software that runs the model—on your hardware.

  2. 2

    Files

    Make only the selected local files available to the workflow.

  3. 3

    Tools

    Control tool execution and network access in your own environment.

  4. 4

    Review

    Inspect the result locally. Local inference alone does not isolate every tool.

Illustrative workflow. Actual access depends on configuration.

What hardware do you need?

Start with the model variant and runtime. Memory must hold the model weights and working context; longer inputs and concurrent tasks need more headroom. Quantization stores model values at lower precision to use less memory, with tradeoffs in quality and speed.

Where it fits in Muse

Spark and Glimmer at a glance
ConsiderationSpark Glimmer
RunsHosted serviceYour hardware
Typical fitComplex reasoning and tool-based workLocal files and self-managed workflows
You provideContext, tools and permissionsHardware, runtime, tools and controls
TradeoffData handling depends on the hosted serviceHardware needs and capability limits

Real examples

Technical architecture

Unlike Spark, Glimmer is downloaded and run on your own hardware using a runtime, the software that runs the model (for example vLLM, SGLang, llama.cpp or ExecuTorch). It is not called through Meta Model API.

Your request passes to Glimmer, context and planning, local browser, file and code tools, then result and review. External services require permissions.
Conceptual flow. The model proposes work; software executes tools and enforces configured controls. Steps may repeat.

Specifications and limitations

Size class
~30B multimodal
Context length
131,072+ tokens in the model card; practical capacity depends on runtime and available memory
License
Apache 2.0
Availability
Open weights · local runtimes
Serving
Not via Meta Model API

Keep in mind

  • Smaller than Spark — not a drop-in replacement for its capabilities on complex tasks
  • Requires suitable hardware and serving setup
  • You own security, updates and evaluation for production use