block #0001 in --:--:--Join the pool
PENDING…
ai pending 5/6 1h ago · 6 min read

Cohere Embed 5 is generally available. The scoreboard is still Cohere's.

Cohere Embed 5 arrived on September 30, 2026, as Embed 5 Pro and Embed 5 Fast, generally available on Cohere's own surfaces and on two big clouds. The ViDoRe numbers people will screenshot are Cohere's tests, and The New Stack repeats them rather than rerunning them.

pending 5/6 — still in the mempool

Early story. Some claims here are not officially confirmed yet. We update this post as it confirms.

Cohere Embed 5 is generally available. The scoreboard is still Cohere's.
in this block
  1. What actually happened
  2. What to do as a reader
  3. What would count as a next source

Cohere made Cohere Embed 5 generally available on September 30, 2026, in two sizes of the same idea: Embed 5 Pro and Embed 5 Fast. The surfaces named with that release are the Cohere API, Model Vault, Microsoft Foundry, and Amazon SageMaker. The New Stack repeats the date, the prices, a throughput claim, and the benchmark figures.

TL;DR - On Cohere Embed 5, text is priced at $0.12 per million tokens for Pro and $0.08 for Fast. Image input is $0.40 per million for both. - Both offer a 128K context, support 100-plus languages, take text, image, or fused inputs, share an embedding space, and expose Matryoshka sizes from 2048 dimensions down to 256. - On Cohere's ViDoRe V3 average for parsed text, Pro is 85.8 and Fast is 84.5. A cross test of Fast queries on a Pro index scores 98.4 against a Pro-on-Pro baseline of 100.

What actually happened

The New Stack does not rerun those figures. The scores below are Cohere's own tests.

General availability is the status word, and the distribution list is the part a team can check without believing a chart. You are not waiting on a waitlist in the launch account. The named doors are Cohere's API, Model Vault, Microsoft Foundry, and Amazon SageMaker.

If your stack already lives in one of those, the release is a model name you can look up. If it does not, nothing in these pages says a fifth door opened.

The price card is simple enough to mis-copy, so here it is in one place, attributed to the launch and repeated by The New Stack. Pro text is $0.12 per million tokens. Fast text is $0.08 per million tokens.

Image input is $0.40 per million tokens for both. Text and image are different rates. Pro and Fast are different rates only on text. Quoting a single "Embed 5 price" flattens a card that has three numbers.

On Cohere Embed 5, the shared spec line is where Pro and Fast look like siblings. Both claim a 128K context and 100-plus languages. Both take text, image, or fused inputs.

They share an embedding space. Matryoshka dimensions run from 2048 down to 256, which is the release's way of saying you can keep a shorter vector from the same family. Shared space is the premise for the cross test later. Without it, querying one index with the other model would be a party trick instead of a product claim.

Cohere's ViDoRe V3 average, on parsed text, using RCP-nDCG@10, is Pro 85.8 and Fast 84.5. The same average in Cohere's comparison is Voyage 4 Large at 83.7, Gemini Embedding 2 at 83.2, and OpenAI text-embedding-3-large at 75.5. Pro is up 8.8 versus Embed 4 on that Cohere figure.

The New Stack repeats these scores. Repeating is not an outside rerun. Say "Cohere's test" every time the number moves to a slide.

The cross-model test is the other number that will get stripped of its setup. Across 40 datasets, Fast queries against a Pro index score 98.4, versus a Pro-query on Pro-index baseline of 100. That is a claim about mixing Fast and Pro inside the shared space, on those datasets, in Cohere's harness. It is not a claim that the models are interchangeable for every corpus you have.

Storage is spelled out with an example rather than a slogan. A 2048-dimension float32 vector is 8 KB, and the launch notes about 819 GB for 100 million chunks at that size. A 256-dimension binary vector is 32 bytes, and the same example puts 100 million chunks near 3.2 GB.

The New Stack also repeats a 2.4 times document throughput figure for Fast. Like the scores, that multiple is part of the launch claim being repeated, not a stopwatch The New Stack ran in public for this story.

What to do as a reader

Write the score's owner in the same cell as the score. Pro 85.8, Fast 84.5, and the comparison points for Voyage, Gemini, and OpenAI text-embedding-3-large are Cohere's ViDoRe V3 average on parsed text at RCP-nDCG@10. If a vendor chart shows up in a newsroom, the newsroom can still be echoing the vendor.

The New Stack is doing that here, on the September 30 release, the prices, the 2.4 times throughput line, the 98.4 figure, and the ViDoRe scores. Useful journalism, wrong thing to call independent.

Decide Pro versus Fast with the card you actually have, not with a one-point story. On Cohere's average, Fast sits at 84.5 and Pro at 85.8. Text is cheaper on Fast, at $0.08 per million versus $0.12.

On Cohere Embed 5, image is the same $0.40 per million for both. Fast is also the model in the 2.4 times document throughput claim. None of those facts tells you which index to build. They tell you which trade the launch is offering, and they tell you the evidence is in-house.

If you care about mixing models, keep the 98.4 attached to its baseline. Fast queries on a Pro index, 40 datasets, baseline 100 for Pro on Pro. A shared embedding space is why that experiment is even on the page.

It is a reason to test your own queries against an index you built with one of them. It is not a reason to assume a 256-dimension binary cut and a 2048-dimension float index tell the same story. The storage example exists so you can see those footprints are wildly different, using the figures they published, without you inventing a new one.

Check the door you will actually call. General availability on the Cohere API, Model Vault, Microsoft Foundry, and Amazon SageMaker is a distribution fact from the September 30 launch. Confirm the model ID in the console you pay, because a blog list is not your account's enablement switch.

While you are there, match the rate to text or image. The $0.40 image rate is easy to miss if you only remember the text pair.

What would count as a next source

An outside rerun of ViDoRe V3, with the same parsed-text setup and the same RCP-nDCG@10 average, would be a different article. It is not this one. A customer index with a published method would also be a different article. Until that exists, the honest caption is that Cohere measured Cohere, and The New Stack told readers the measurement again.

The Matryoshka range, the fused inputs, the 100-plus languages, and the 128K context are capability claims from the same launch. They belong in a spec table with a source line. They do not promote the 85.8 into an industry ranking.

Embed 4 is the internal comparison, at plus 8.8 for Pro on Cohere's average. Competitors in the table are there because Cohere put them there.

On Cohere Embed 5, start with Cohere's Embed 5 post and then read The New Stack's repeat of the launch. If the second page's numbers match the first, that is consistency of copying, which is worth having, and it is still not a second lab.

Readers who want sourced recaps that are already on the site can read DogeOS public testnet and PUMP token buyback burn as separate live posts.

Not financial advice. DYOR, ser.

More in the pool

all ai
gm ser

Get confirmed before the crowd

Daily block at 07:00 UTC. No spam, just the block, ser.