Close Menu
    Facebook X (Twitter) Instagram
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Facebook X (Twitter) Instagram
    Fintech Fetch
    • Home
    • Crypto News
      • Bitcoin
      • Ethereum
      • Altcoins
      • Blockchain
      • DeFi
    • AI News
    • Stock News
    • Learn
      • AI for Beginners
      • AI Tips
      • Make Money with AI
    • Reviews
    • Tools
      • Best AI Tools
      • Crypto Market Cap List
      • Stock Market Overview
      • Market Heatmap
    • Contact
    Fintech Fetch
    Home»AI News»Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence
    Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence
    AI News

    Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence

    October 1, 20266 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email
    binance

    Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines. Each chunk is embedded with the full document in view. The real change is the training signal. The model learns to retrieve the answer along with the context needed to verify it, not one ‘gold passage.’ P

    Is it deployable? Yes, as a self-hosted preview. Weights are on Hugging Face under the MIT license. Loading requires transformers>=5.4.0 with trust_remote_code=True. It is not yet on the Perplexity API. The model card notes that weights and interface may change without backward compatibility.

    Why the gold passage falls short

    RAG systems split long documents into chunks. A chunk often depends on an entity, heading, or definition stated elsewhere. Contextual models address this with late chunking. The document is encoded in one pass, then pooled per chunk.

    Training, however, usually marks one gold chunk per query. Every other chunk becomes a negative, including the sentences that make the answer checkable. Perplexity lists 3 more problems. Binary labels give a coarse signal. LLM annotation cost grows linearly with dataset size. Labels are also tied to one chunking strategy.

    How the training works

    The teacher is Perplexity’s query-aware context compression model. It reads the query and document together and scores every token.

    binance
    • Chunk relevance: the mean of the top n token scores inside each chunk.
    • Soft target: a temperature-scaled softmax over chunks in the positive document. Chunks in other documents get zero.
    • Distillation loss: forward KL divergence between teacher and student distributions.
    • Document loss: InfoNCE, where a document scores as its best chunk, inspired by ColBERT’s MaxSim.

    Each batch samples a random chunking strategy. Chunks are separated by a learned <|chunk_sep|> token and mean-pooled. The teacher runs only during training, so inference adds no latency or storage.

    The model starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048 dimensions. Matryoshka training also supports 1024 dimensions. Quantization-aware training enables native int8 embeddings. The release is a model soup of several checkpoints. Training used roughly 430 datasets covering over 50 languages, with no ConTEB data.

    Interactive explainer

    pplx-embed-v2 Contextual Embedding Explainer

    An interactive walkthrough of Perplexity’s contextual embedding preview: why isolated chunks fail, how teacher distillation replaces the single gold passage, and what the reported numbers mean.

    01Why context
    02How it trains
    03Reported results
    04Storage math

    Three lease files share the sentence “Monthly rent is …”. Only one belongs to 5 Park Avenue. Switch modes and run retrieval.

    QUERYWhen does 5 Park Avenue’s lease end and what is the current rent?

    Isolated sentencesContextual (late chunking)

    Run retrieval

    Pick a mode and press Run retrieval.

    Illustrative example modeled on the lease scenario in Perplexity’s post. Scores are for explanation only, not model outputs.

    A context compression model acts as teacher. It scores every token for the query. Those scores are pooled per chunk (mean of the top n tokens) and turned into a soft target, instead of a one-hot gold label.

    1 Teacher scores tokens2 Top-n mean per chunk3 Softmax target4 Student matches via KL

    Gold-chunk label (one-hot)

    Change the boundaries: the same token scores re-aggregate without re-annotation. That is the “flexible chunk boundaries” property Perplexity describes. Token scores here are illustrative.

    context-bench (2,099 queries, 38,894 documents, 2,458,072 sentence chunks, exhaustive ranking). Numbers below are as reported by Perplexity at K = 10.

    pplx-embed-v2-context-9b-previewvoyage-context-4 (derived from reported gap)

    Replay

    Voyage values are computed as Perplexity’s figure minus the stated gap (14.4 and 5.0 points). Other Voyage metrics appear only in Perplexity’s chart and are not shown here.

    Contextual embeddings store one vector per chunk, same as a normal chunk index. Cost depends on vector size. Perplexity reports that 1024-dim int8 (1 KB) slightly exceeds voyage-context-4 at 2048-dim float32 (8 KB) on its chunk-retrieval suite.

    Chunks

    100,000
    1,000,000
    2,458,072 (context-bench)
    10,000,000
    100,000,000

    1024-d2048-d

    int8float32

    0

    vector storage (vectors only, not full index)

    Chunk-size sensitivity (64 to 512 tokens)

    81.0% to 79.9%

    Bytes = dimensions x bytes per value. Sensitivity is mean nDCG@10 across 74 MTEB tasks, as reported by Perplexity.

    context-bench: a new benchmark

    context-bench is built and privately held by turbopuffer to limit training contamination. It holds 2,099 queries over 38,894 documents in 21 domains. Sentence chunking yields 2,458,072 chunks. Median target document length is roughly 6,100 tokens. Queries test 12 contextual capabilities, from pronoun resolution to table structure. Perplexity says the model was submitted blind.

    Metrics are Document@K, Answer@K, Evidence Recall@K, and All-Evidence@K. Every model is ranked exhaustively against all chunks, so index settings play no role.

    Results

    • context-bench at K = 10: 45.5% answer recall, 40.6% evidence recall, 31.1% all-evidence recall.
    • Document recall: 15.2% at K = 1 and 61.6% at K = 10.
    • vs voyage-context-4: ahead by 14.4 points on answer recall and 5.0 on evidence recall at K = 10.
    • ConTEB: highest average nDCG@10 among models shown. pplx-embed-context-v1-4B wins NarrativeQA, and Nemotron-3-Embed-8B wins COVID-QA.
    • General retrieval: best average on query-to-chunk tasks. Slightly behind voyage-context-4 on query-to-document.
    • Storage: 1024-dim int8 (1 KB per vector) slightly beats voyage-context-4 at 2048-dim float32 (8 KB).
    • Chunk size: average score moves from 81.0% to 79.9% between 64 and 512 tokens.

    Comparison with the closest competitors

    Featurepplx-embed-v2-context-9b-previewvoyage-context-4pplx-embed-context-v1-4BNemotron-3-Embed-8BChunk embeddingsContextualContextualContextualIndependent per chunkAccessOpen weights, MITHosted API (Voyage, MongoDB Atlas)Open weights, MIT; Perplexity APIOpen weights, OpenMDW-1.1Parameters9B per blog (Hugging Face lists 8B)Not disclosed (MoE backbone)4BAbout 8BDimensions2048, 10242048, 1024, 512, 2562560 (Matryoshka)4096, sliceableQuantized outputNative int8int8, uint8, binary, ubinaryint8, binaryFloatContextEvaluated up to 32,768 tokens32K per request; 120K with auto-chunking32K32,768Auto-chunkingNoYesNoNoPriceSelf-host$0.12 per 1M tokens; first 200M freeSelf-host or APISelf-host

    Sources: Voyage docs, model cards linked above. Checked September 30, 2026.

    Key Takeaways

    • A token-level teacher replaces the single gold-chunk label.
    • The model retrieves answers plus supporting evidence in one chunk index.
    • context-bench Answer@10 is 45.5%, 14.4 points above voyage-context-4.
    • Open MIT weights ship today; Perplexity API access is still pending.

    Check out the technical details and model weights. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

    livechat
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Fintech Fetch Editorial Team
    • Website

    Related Posts

    Who we become when we talk to machines | MIT News

    Who we become when we talk to machines | MIT News

    September 30, 2026
    Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

    Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

    September 29, 2026
    Estimating suicide risk from text | MIT News

    Estimating suicide risk from text | MIT News

    September 28, 2026
    AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

    AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

    September 27, 2026
    Add A Comment

    Comments are closed.

    Join our email newsletter and get news & updates into your inbox for free.


    Privacy Policy

    Thanks! We sent confirmation message to your inbox.

    kraken
    Latest Posts
    Ethereum Privacy Push Grows as Aztec Relaunches zk.money Wallet

    Ethereum’s Privacy Initiative Expands with Aztec’s Relaunch of the zk.money Wallet

    October 1, 2026
    person on phone leaning against outside wall with scenic view at airbnb rental property

    Is BCE Still a Good Investment? Here’s My Opinion.

    October 1, 2026
    Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence

    Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence

    October 1, 2026
    Laziest Way to Make Money with AI (If You're Broke)

    Laziest Way to Make Money with AI (If You’re Broke)

    October 1, 2026
    OpenAI's GPT Escaped Again, and it Proves How Dangerous AI Really Is

    OpenAI’s GPT Escaped Again, and it Proves How Dangerous AI Really Is

    October 1, 2026
    livechat
    LEGAL INFORMATION
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Top Insights
    Quote.Trade Releases V6 AI-Powered Dark Pool DEX & Alpha Trading League

    Quote.Trade Releases V6 AI-Powered Dark Pool DEX & Alpha Trading League

    October 1, 2026
    Bitcoin Options Traders Hedge For More Downside As Deribit Warns Of Market Weakness

    Strive Acquires 1,107 Bitcoin for $94.5M, Increasing Holdings to Over 27,400 BTC

    October 1, 2026
    synthesia
    Facebook X (Twitter) Instagram Pinterest
    © 2026 FintechFetch.com - All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.