pplx embed preview encodes the document before the chunk
The pplx embed preview is a document-level embedding model, and the model card does not print the score the newsroom leads with. One page is a warning label. The other is a claim of a new retrieval mark, built mostly from the previous generation.

in this block
The pplx embed preview is the Hugging Face checkpoint perplexity-ai/pplx-embed-v2-context-9b-preview. The card says a document is passed as a list of chunks, the chunks are encoded together so each embedding reflects its surrounding context, and one embedding comes back per chunk. CryptoBriefing, in a story dated Sept. 30, 2026, says Perplexity claims a new state of the art for document retrieval with this model. The card that loaded here does not print a benchmark table. Those are different kinds of pages, and this note does not promote the newsroom's older numbers into a score for this file.
TL;DR - The model card calls the release a preview. Weights, embeddings, and the interface may change without backward compatibility, so embeddings from this preview should not be mixed with a later release. - The card lists 2048 dimensions, Matryoshka sizes of 1024 and 2048, native unnormalized INT8 embeddings, mean pooling, and no free-form instruction, only fixed query and document prefixes. The card's size line says 8B parameters. The repo name says 9b. - CryptoBriefing's numeric comparisons, including an 81.96 percent ConTEB nDCG@10, are about the February 26, 2026 v1 context model, not a score printed for this preview.
What actually happened
CryptoBriefing's lead is the company's claim. It says pplx-embed-v2-context-9b-preview sets a new state of the art for document retrieval because it encodes an entire document once, so each piece carries the context of the whole. The problem it describes is ordinary chunking. A slice that says a company raised its forecast is useless if the embedding has forgotten which company the document was about. Isolated chunk embeddings are, in the article's image, one page torn out of a novel. The preview tag, the article says, marks an early release rather than a finished product, and Perplexity is positioning the model as the next step after its first generation of contextual embedders.
The model card is more specific about the interface and less specific about the crown. Queries and documents use different methods. The card says to call encode_queries for queries and encode for document chunks. The model was trained with separate query and document prefixes, and encoding a query with encode silently degrades retrieval quality. Like pplx-embed-context-v1, it natively produces unnormalized int8 embeddings. Compare them with cosine similarity, or set normalize_embeddings to true and use a dot product.
The spec table on the card is one row. Dimensions are 2048. Matryoshka representation learning sizes are 1024 and 2048. Quantization is INT8. Instruction is "No," with fixed prefixes. Pooling is mean. Usage requires transformers 5.4.0 or newer, plus torch, numpy, safetensors, and tqdm. The model uses custom code, so it loads with trust_remote_code set to true. The example encodes a list of documents, each a list of chunk strings, and returns one array per document. A three-chunk document comes back with shape (3, 2048). Queries are passed as single-chunk rows.
Defaults in that same card: batch size 32, embeddings not normalized, NumPy arrays rather than tensors, no progress bar, and no forced device unless you pass one. To get 1024 dimensions, take the first 1024 values of each unnormalized embedding and normalize after the slice. The card says other truncation sizes were not trained. The downloads line on the card read 145 for the last month. The tensor type listed is F32. The size line says 8B params. Anyone calling the pplx embed preview a 9 billion parameter model is reading the repo name, not that size line. This article does not resolve the gap. It reports both.
What the newsroom measured instead
Most of the numbers in CryptoBriefing belong to the previous family, and the article marks the date. On Feb. 26, 2026, Perplexity released pplx-embed-v1 and pplx-embed-context-v1, each in 0.6B and 4B sizes. On ConTEB, pplx-embed-context-v1-4B scored an average nDCG@10 of 81.96 percent. The article explains nDCG@10 as whether the relevant results show up in the top 10 and whether the best ones land near the top. voyage-context-3 scored 79.45 percent on the same measure. A model from Anthropic scored 72.4 percent. Those three figures are the v1 comparison. They are not on the v2 card opened here.
The same section says Perplexity reported recall gains over Qwen3 embeddings on internal web-scale tests of more than 1 billion production pages, and that Perplexity ran those tests itself. Training for v1, as the article tells it, started with diffusion-based techniques on about 250 billion multilingual tokens, then a multi-stage contrastive setup with bidirectional attention. The v1 models support a 32K context, native INT8 and binary quantization, up to 32x storage savings by the company's account, and no instruction prefixes. API pricing for the v1 lineup runs from $0.004 to $0.05 per million tokens, depending on variant and size. None of those sentences is a price or a context length for the preview checkpoint. The card does not list a price, a token limit, or a binary mode.
CryptoBriefing's closing caveat is the responsible part of the piece. It says some of the striking comparisons, including the Qwen3 recall gains, come from Perplexity's internal benchmarks, and that a preview may change before any general release. That warning matches the card's own compatibility note. What the newsroom does not supply is a public nDCG, a recall figure, or a cost for this 8B-or-9b file. A reader who pastes 81.96 percent under the v2 name is writing a result neither page contains.
Read as an interface, the pplx embed preview is still a real change relative to "embed each chunk alone." The card's encode call takes the whole document's chunks in one list. That is the mechanism CryptoBriefing describes in prose and the card specifies in code. Mechanism is not a leaderboard win. State of the art, in the Sept. 30 story, is the company's claim about retrieval, not a number this article can check.
What a reader should not collapse
Do not collapse the name. "9b" in the repo and "8B params" on the card are both printed. Until a third page counts the tensors, the honest line is that the two labels disagree.
Do not collapse v1 and v2. The 0.6B and 4B models, the 32K window, the $0.004 to $0.05 prices, the 81.96 percent, and the 250 billion token training line are CryptoBriefing on the February release. The preview card's distinctive lines are the 2048-wide INT8 vectors, the 1024 Matryoshka slice, the split between encode and encode_queries, trust_remote_code, and the ban on mixing embeddings across versions.
Do not collapse a preview with a stable index. The card says later weights may break compatibility. If you embed a corpus now and swap the checkpoint later, the old vectors and the new vectors are not one space. That is a storage problem, not a slogan.
For another embedding release with its own scorecard, see the Cohere Embed 5 note. For a translation model that is not a retriever, see the Bilibili translation note. Neither one is a license to treat 81.96 percent as this checkpoint's ConTEB result.
What to do as a reader (not a trade)
This is not investment advice. There is no token on either page. If you are deciding whether to index documents with the pplx embed preview, freeze three choices the card already forces: encode versus encode_queries, 2048 versus a normalized 1024 slice, and whether you can afford to rebuild the index when the preview moves.
Then ask the question the pplx embed preview card does not answer with a table. Which public set, which nDCG, and whose harness? If the only numbers you can cite are from Feb. 26, say Feb. 26. Internal recall on a billion pages is the company's run, as the newsroom notes, not an independent audit.
The primary page is the Hugging Face model card. The second page is CryptoBriefing's Sept. 30 story. Use the card for the interface and the compatibility warning. Use the newsroom for the company's state-of-the-art claim and for the v1 figures it actually prints.
Not financial advice. DYOR, ser.