Keep inference local
Inference means running the model to produce an answer. Here, it runs on your hardware.
AI on your hardware, close to your files and tools.

Muse Glimmer is Meta's open-weight model: its model files are available to download and run yourself. It understands text and images, is distilled from Spark—trained using the larger model—and is released under Apache 2.0 for workflows on your own hardware.
Inference means running the model to produce an answer. Here, it runs on your hardware.
Choose which local documents and source files a workflow can access.
Use your own tool loop and execution controls.
Choose a runtime and manage the model yourself.
Follow the model, files, tools and review stages of a local workflow.
Load Glimmer in a compatible runtime—the software that runs the model—on your hardware.
Make only the selected local files available to the workflow.
Control tool execution and network access in your own environment.
Inspect the result locally. Local inference alone does not isolate every tool.
Start with the model variant and runtime. Memory must hold the model weights and working context; longer inputs and concurrent tasks need more headroom. Quantization stores model values at lower precision to use less memory, with tradeoffs in quality and speed.
| Consideration | Spark | Glimmer |
|---|---|---|
| Runs | Hosted service | Your hardware |
| Typical fit | Complex reasoning and tool-based work | Local files and self-managed workflows |
| You provide | Context, tools and permissions | Hardware, runtime, tools and controls |
| Tradeoff | Data handling depends on the hosted service | Hardware needs and capability limits |
Unlike Spark, Glimmer is downloaded and run on your own hardware using a runtime, the software that runs the model (for example vLLM, SGLang, llama.cpp or ExecuTorch). It is not called through Meta Model API.
