Bilibili translation model family goes open on Apache-2.0
The Bilibili translation model family, Index-Translate, was released on September 30, 2026. Two English reports agree on the sizes and one FLORES score. They do not agree on every other number, and this piece does not average them.
Early story. Some claims here are not officially confirmed yet. We update this post as it confirms.

in this block
The Bilibili translation model family, Index-Translate, was released by the company's Index LLM team on September 30, 2026. AI Weekly's alert is timestamped 21:27 UTC that day, which is 00:27 on October 1 in Moscow, and it was updated a day later. Pandaily's English write-up is dated October 1.
TL;DR - Both pages describe Apache-2.0 text models in 2B, 9B, and 35B-A3B preview sizes, built on Qwen3.5, covering 150 languages, with weights on Hugging Face and ModelScope. - Both pages print a FLORES COMET-22 score of 0.8794 for the 35B preview. Other scores do not match line for line, and this article does not average them. - Specialized branches cover dubbing, syllable-limited translation, and long documents. A non-preview 35B model is still described as future work.
What actually happened
Pandaily says the Bilibili translation model text checkpoints follow instructions about terminology, formatting, and spans that must be left untouched. One project-page example, which this article will not reprint, has the 9B model translate a JSON game-maintenance notice into Korean while keeping structure, star symbols, and a hashtag. The same page says a slang line about the urge to play a long-running online game was rendered as a short English sentence that kept the joke. In a fantasy document of about 32,000 tokens, where a character's name could also be read as a royal title, the long-document model kept the name stable, while the 9B model working chunk by chunk drifted.
The family extends past plain text. Index-Echo produces translated subtitles or dubbed speech that keeps the source speaker's voice characteristics. Pandaily says the packaged speech-to-speech release covers Chinese to English, Spanish, and Japanese, and English to Chinese, Spanish, and Japanese. Index-Homura pushes a translation toward a target syllable count for dubbing and subtitles. Index-NativeLong, published under model IDs named Index-Nailong, translates whole documents and tries to keep references consistent. AI Weekly describes the same three branches, with Echo as speech-to-text and speech-to-speech, Homura as syllable-controlled dubbing, and NativeLong as full-document work, and it says the 2B and 9B sizes exist for those tasks as well as for text.
Pandaily also notes a browser extension for translating pages with a locally deployed model, and a video dubbing pipeline. Bilibili's stated plan, on that page, is to release the official 35B-A3B model, open-source its benchmarks, add languages to Index-Echo, and publish larger models. The preview label on the 35B text model is therefore doing real work. It is not the finished checkpoint.
Which numbers are on which page
The shared figure is the careful one. AI Weekly says the preview lands a 0.8794 COMET-22 on FLORES. Pandaily says the 35B-A3B preview scores 0.8794 on FLORES using COMET-22. Use that number once. Do not invent a second FLORES run.
From there the pages diverge, and the divergence looks like different slices rather than a single typo. Pandaily puts the 35B preview at 76.76 on a WMT26 judge score, and the 9B model at 0.8789 COMET-22 and 75.35 on that judge score. It says some large general models score higher on WMT26 in the same table, including DeepSeek-V4.1-Flash at 83.55. On low-resource instruction following, Pandaily says the 9B model posted the best instruction score in the comparison, 0.7725, and the lowest off-target rate, 3.47 percent. Index-Homura-9B, on the team's SandGlass test, lands within 10 percent of the requested syllable count 81.92 percent of the time.
AI Weekly's card-level notes are different rows. It says low-resource pairs show an off-target rate of 2.4 percent, the lowest in that comparison set. It reports domain scores of 0.8962 COMET-22 on MuST-Cinema subtitles and 0.8681 on biomedical text, plus 77.6 percent on C-Eval and 49.1 percent on GPQA-Diamond. It describes training as 167.77 billion tokens of multilingual mid-training, with a constant stage mixing general, parallel, and monolingual data at 1:1:1 and a decay stage weighted 1:4:2 toward pivot data. Post-training, in that account, runs through specialist supervised fine-tuning and reinforcement learning, a general-translation step judged by XCOMET-XXL, an instruction step described as rubric-as-reward, parameter interpolation, and a final multi-teacher distillation. Serving is a one-line vLLM command with a shipped context window of 262,144 tokens and a recommended maximum length of 32,768.
The 2.4 percent and 3.47 percent off-target rates are both "lowest in a comparison," attached to different wordings. This article will not average them to 2.9 or pick a winner. AI Weekly itself asks what sits inside the unnamed comparison set. That question is still open on the page.
AI Weekly also says the 9B model matches or exceeds two named closed models on WMT26 judge scores and FLORES. Pandaily's table, with DeepSeek-V4.1-Flash at 83.55 against the preview's 76.76, shows that "matches or exceeds" cannot mean every large model on every column. Read each sentence against the models it actually names.
The Bilibili translation model card, as quoted by AI Weekly, already limits its own marketing: scores vary by language and task, combined constraints and low-resource directions stay harder, and a successful example does not mean every instruction will be followed. That caveat is more useful than the headline about 150 languages.
What is still unsettled
Neither page opened here is the Bilibili translation model repository. Both are news write-ups of the card and the project page. AI Weekly asks whether the Qwen3.5 backbone's license allows Apache-2.0 redistribution of the fine-tuned weights. Pandaily states that the release is Apache-2.0 and does not walk through that license question. Do not treat the license as settled because a news desk repeated the badge.
The 35B-A3B design, 35 billion parameters with 3 billion active, is AI Weekly's explanation for why inference cost can stay near the 9B model on consumer hardware. Pandaily states the size labels and does not repeat that cost sentence. Hardware cost is single-source on these two pages.
Pandaily's plan to ship a non-preview 35B model means today's checkpoint is a preview by the team's own roadmap. Benchmarks the team says it will open-source later are not open in the sense that an outside lab has reproduced 0.8794. AI Weekly says that reproduction is the near-term test.
What to do as a reader (not a trade)
This is not investment advice. A Bilibili translation model release is not a reason to trade a media stock or a token. If you localize subtitles, the practical questions are narrower.
Check whether your language pair is inside the 150 and whether you need Echo's speech path, which Pandaily limits to a short list of Chinese, English, Spanish, and Japanese directions. A text model that covers many languages is not the same product as a dubbing model with six directions. If syllable count matters, Homura's 81.92 percent figure is the team's SandGlass result, not a guarantee on your episode.
Prefer the shared 0.8794 over any blended score. If you deploy the preview, start from the recommended 32,768 context even though a longer window is shipped, because that is the setting AI Weekly says the card recommends. Keep the team's own caveat: one clean JSON example does not prove every formatting instruction will hold.
Separate model launches from the same week, including Gemini 4 Argon and Cohere Embed 5, do not test this translation stack. AI Weekly's alert and Pandaily's October 1 story overlap on the release facts and then print different score rows. Use both, and do not merge the rows.
Not financial advice. DYOR, ser.