BGE-M3
BGE-M3 is BAAI's January 2024 multilingual embedding snapshot, not a 2026 retrieval crown. Rechecked 2026-08-17: huggingface.co/BAAI/bge-m3 is still listed, MIT, XLM-RoBERTa, 568M, 1024-d, 8192 tokens, with leftover dense / multi-vector / sparse copy. Drop leftover first model to do all three / MIRACL nDCG@10 70.0 / MKQA 75.5% / beats OpenAI latest embedding as a 2026 census.
Compare text-embedding-3-large if you wanted a still-listed OpenAI v3 row, and text-embedding-ada-002 if you wanted the older billed OpenAI snapshot.
Key Features
- Previous open card: The HF page we opened still hosts the 2024 weights. Recheck that card before you quote 568M / 1024-d / 8192.
- Not a 2026 leaderboard: MIRACL 70.0 and MKQA 75.5 stay off the table unless you re-run a first-party board.
- Hybrid story is leftover: Dense + multi-vector + sparse is still on the card. Do not call it the first or only hybrid embedder in 2026.
- No invented API dollar: Self-host or recheck a host. Do not invent a BGE-M3 token price.
Use Cases
- People with a 2024 BGE-M3 bookmark who need the catalog corrected.
- People comparing leftover hybrid BGE vs billed OpenAI v3 rows.
- People who were about to paste MIRACL 70.0 into a 2026 deck.
Limitation: We did not re-embed a multilingual corpus or re-run MIRACL / MKQA.
Pricing
First-party pages on 2026-08-17.
| Piece | Price | Notes |
|---|---|---|
| BGE-M3 weights | MIT card on Hugging Face | Recheck the live card for license and files. |
| Hosted inference | Recheck the host | Do not invent a per-1M dollar. |
If a leftover blog still calls BGE-M3 the current multilingual crown, treat this page as the correction.
Getting Started
- Open huggingface.co/BAAI/bge-m3 before you download weights.
- Recheck text-embedding-3-large if you wanted a billed OpenAI v3 row.
- Recheck text-embedding-3-small if you wanted the cheaper listed v3 row.
- Do not paste MIRACL 70.0 into a deck.
Frequently Asked Questions
Is BGE-M3 still the multilingual embedding flagship?
Not as a 2026 census. It is a still-listed 2024 MIT card.
Did it beat OpenAI on MKQA?
That leftover score is out as a 2026 census.
Same as text-embedding-3-large?
No. 3-large is a billed OpenAI v3 row. BGE-M3 is an open 2024 snapshot.
Alternatives
- text-embedding-3-large: Still-listed OpenAI v3 large row.
- text-embedding-3-small: Cheaper listed OpenAI v3 row.
- text-embedding-ada-002: Older billed OpenAI snapshot.
Tips
- Call BGE-M3 a January 2024 snapshot.
- Recheck the HF card before you quote 8192 tokens or 1024-d.
- Do not invent a 2026 MTEB rank.
Conclusion
BGE-M3 is a 2024 BAAI multilingual embedding snapshot that is still on Hugging Face, not a 2026 MIRACL or MTEB crown. Start at huggingface.co/BAAI/bge-m3, then decide whether an OpenAI v3 row already covers the retrieval path you need.
Comments
No comments yet. Be the first to comment!
Related Tools
Qwen3-Embedding
huggingface.co/Qwen
Previous Qwen3 embedding snapshot. Rechecked: still listed on Hugging Face. Drop leftover MTEB #1 as a 2026 census.
NV-Embed-v2
huggingface.co/nvidia/NV-Embed-v2
2024 NVIDIA embedding snapshot. Rechecked: HF card still listed. Not a 2026 MTEB #1 census.
EmbeddingGemma
ai.google.dev/gemma
Lightweight multilingual text embedding model from Google DeepMind, optimized for on-device AI with <200MB RAM usage.
Related Insights

Grok Bot and Hermes Bot: one person finally gets a think tank and a secretariat
Grok Bot now ships with Cursor Pro+. Hermes Bot runs on a VPS. They are not smarter chat boxes. The think tank advises, the secretariat executes, and you still make the call.
Hook OpenCode Go into Codex on Windows. Do Not Open a Second Toolkit.
A ChatGPT-signed Codex desktop app still shows mostly GPT in the picker. On Windows, enable only OpenCode Go and the Grok, GLM, Kimi, DeepSeek, and MiniMax models you already pay for appear in the same selector. Keys stay local. Native GPT stays put.
What Locks Codex Is the Picker, Not the Models
You already pay for OpenCode Go, Grok, and Z.ai, but the Codex picker still shows mostly GPT. The community project codex-router does not teach another install ritual. It puts subscriptions you already bought back into the selector, keeps keys on the machine, and leaves native GPT alone.