Felix Pinkston
Aug 10, 2026 11:02
Meta’s Muse Glimmer 30B, a 30B parameter open AI mannequin, now runs effectively on AMD {hardware}, enabling highly effective native AI purposes.

Meta’s newest AI mannequin, Muse Glimmer 30B, is setting a brand new customary for native inference, and AMD {hardware} is positioned because the prime platform for operating it. This 30-billion-parameter dense mannequin, launched beneath the Apache 2.0 license, is optimized for agentic and coding workloads, giving builders a sturdy software for native AI purposes with out counting on cloud infrastructure.
The mannequin is particularly designed to leverage AMD’s Ryzen™ AI Max+ processors and Radeon™ AI PRO R9700 graphics playing cards. Benchmarks point out spectacular efficiency: as much as 24 tokens per second on a Ryzen AI Max+ 395 processor and as much as 53 tokens per second on a single Radeon AI PRO R9700 GPU. These figures include dFlash enabled and had been examined utilizing open frameworks like llama.cpp, additional highlighting the flexibility of the system.
Muse Glimmer’s open-weight structure helps native workflows, a essential benefit for privacy-sensitive purposes. In contrast to cloud-first AI options that may introduce latency and safety dangers, operating the mannequin domestically retains knowledge and operations beneath direct management. This aligns with the rising development of decentralized AI adoption, the place builders prioritize on-device options for quicker, safer, and extra cost-efficient purposes. Early studies from the Unsloth neighborhood recommend the mannequin requires about 18GB of RAM for inference, making it accessible for high-end shopper PCs and workstations.
Why Muse Glimmer 30B Issues
The 30B mannequin is a serious leap for agentic AI. It’s designed to deal with complicated, multi-step workflows that demand context retention, software utilization, and adaptive decision-making. With assist for 128K+ context, Muse Glimmer may energy purposes like coding assistants, workflow automation, and personal AI instruments. Its coaching additionally emphasizes security, with mechanisms to withstand oversharing and immediate injection assaults.
Builders can get began with Meta’s LM Studio for straightforward mannequin deployment on AMD {hardware}. Programs with 32GB+ Variable Graphics Reminiscence (VGM) or VRAM are really helpful, permitting customers to run the mannequin in minutes. For integration into current purposes, the Lemonade platform affords a light-weight API that simplifies deployment whereas protecting knowledge native. This embeddable binary method opens the door to quite a lot of real-world use circumstances, from analysis instruments to offline assistants.
Open-Supply and Business Flexibility
Muse Glimmer’s Apache 2.0 licensing provides builders intensive freedom. They will modify, redistribute, and combine the mannequin into industrial purposes with out restrictive phrases. Group contributors have already launched GGUF quantized variations on Hugging Face, enabling extra environment friendly inference on a variety of {hardware}.
This launch positions Meta as a pacesetter within the push for native AI options. By specializing in agentic workloads and prioritizing privateness, Muse Glimmer 30B fills a essential hole available in the market. It’s particularly related as companies and builders look to transition from cloud-dependent AI to on-device options that provide higher management and decrease working prices.
What’s Subsequent?
Muse Glimmer 30B is a part of AMD’s imaginative and prescient for the “Agentic PC,” a platform that strikes past remoted AI options to host persistent intelligence. With additional software program optimizations and ecosystem growth anticipated, the efficiency of Muse Glimmer on AMD {hardware} will possible enhance additional. For builders and enterprises looking for to discover the following section of native AI, this mannequin and its {hardware} companions signify a compelling start line.
Picture supply: Shutterstock
