Perplexity Claims pplx-embed-v2-late Embedding Models Hit 92.4% on MADQA
Summary
Perplexity releases pplx-embed-v2-late, a family of late-interaction embedding models for joint text and image retrieval. The company claims the 9B version scores 92.4% on MADQA, while the 0.6B model matches systems five times larger on ViDoRe (V3). The models, distilled from an 18B teacher, trained on 186 million query-document pairs in 46 languages.
Key Points
- Perplexity releases pplx-embed-v2-late, a collection of late-interaction embedding models handling text and image retrieval in a shared embedding space across model sizes.
- The family includes 0.6B and 9B models distilled from an 18B teacher, trained on 186 million query-document pairs from 594 datasets in 46 languages.
- Perplexity claims state-of-the-art retrieval results, with the 9B model reaching 92.4% accuracy on MADQA and the 0.6B model matching models with five times as many active parameters on ViDoRe (V3).