Stories about Metal
1 related stories
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
AI InsightPerplexity open-sourced Lily, an on-device inference engine based on Rust + Metal tailored specifically for Qwen on Apple Silicon. This suggests that beyond general-purpose frameworks, 'radical customization' for specific hardware-model combinations is becoming a viable path to push edge inference performance limits.Key TakeawayEdge AI deployment is shifting from relying on general frameworks to radical customization for 'specific hardware + specific model' combinations.Why It MattersOn-device inference throughput directly dictates AI assistant responsiveness and local viability. The Rust and Metal co-design proves there is still substantial performance headroom for running multi-billion parameter models on consumer-grade chips.Who's Affected- Local AI DevelopersGain a new high-performance on-device deployment tool to run specific LLMs more efficiently on Apple devices.
- Mlx-LmFaces new competitive pressure in extreme Apple Silicon optimization scenarios.
What's NextSubsequent observations should focus on the open-source community's contribution activity for Lily, and whether more models will be adapted into this hardware-specific optimization framework.Importance 60/100