Perplexity has released two multimodal late-interaction embedding models, pplx-embed-v2-late-0.6b and pplx-embed-v2-late-9b, for text, image and visual-document retrieval. Unlike dense retrievers that compress an input into one vector, the models retain a 128-dimensional vector for each token and use a ColBERT-style MaxSim scoring function. That design increases storage and scoring cost but lets different parts of a query match different parts of a document.
The models are designed to search rendered document pages directly, reducing dependence on OCR or text extraction for PDFs, slides, screenshots, charts and tables. Perplexity trained an 18B teacher and distilled it into the two released sizes using token-level representation alignment. The resulting checkpoints share the same embedding space, so a corpus indexed offline with the 9B model can be queried online with the smaller 0.6B model.
Perplexity reports that this asymmetric configuration improves average retrieval quality by 1.6 percentage points over using the 0.6B model for both indexing and querying on its 72 domain-specific text tasks. On ViDoRe V3 image retrieval it reports 63.5% nDCG@10 versus 62.3% for the symmetric 0.6B setup. The 9B model used on both sides remains stronger. These measurements come from Perplexity's evaluation and should not be treated as independent replication.
Both checkpoints are available on Hugging Face and support sentence-transformers and transformers. The practical significance is architectural: teams can spend more compute once during indexing while keeping live query encoding comparatively light, including local or edge deployments. The trade-off is a larger multi-vector index and more expensive candidate scoring than conventional dense retrieval.