4 entries with this tag
One new research article published: comprehensive analysis of Meta's Muse Glimmer 30B — a distilled local-first agentic model running on consumer hardware with Apache 2.0 licensing and DFlash speculative decoding.
On August 10, 2026, Meta released Muse Glimmer — a 30B-parameter multimodal agentic model distilled from Muse Spark, released under Apache 2.0, and optimized to run on a single consumer GPU. Covers the distillation pipeline, DFlash speculative decoding, 3.1x speedup on RTX 5090, benchmark results against Gemma4-31B and Qwen3.6-27B, the safety evaluation framework, and strategic implications for the local agent ecosystem.
June 5: One new research article — Gemma 4 12B, the encoder-free multimodal laptop model that changes the game. Google DeepMind's 12B dense model eliminates separate vision/audio encoders entirely, runs on 16GB laptops under Apache 2.0, and delivers 78.8% GPQA Diamond. The efficiency revolution now has a multimodal face.
Google DeepMind releases Gemma 4 12B — a 12B dense model with encoder-free multimodal architecture, native audio support, and 256K context. Runs on 16GB laptops under Apache 2.0. Benchmarks approach the 26B MoE sibling at less than half the memory. The most practical multimodal model for local deployment yet.