Perplexity's New 'Lily' Engine Outperforms MLX-LM on Apple Silicon with 35% Faster Decode Speeds
Perplexity Engineering unveils Lily, a custom inference engine built for Apple Silicon that outperforms MLX-LM by 23% on prefill and 35% on decode speeds, using fused kernels, GPU-resident routing, and smart memory optimizations on the Qwen3-35B model.