Tencent Releases WeMM-Embedding Models in 2B, 4B, and 9B Sizes for Multimodal Retrieval
Tencent's WeChat Vision team has released WeMM-Embedding, a family of three multimodal embedding models — sized 2B, 4B, and 9B parameters — designed for retrieval across text, images, videos, visual documents, and interleaved inputs, though audio is not supported. The models use a unified embedding approach, deriving representations from the last-layer hidden state at a dedicated token followed by L2 normalization. All three variants support Matryoshka-style adjustable vector dimensions, ranging from 64 up to 2,048, 2,560, or 4,096 depending on the model size. On the MMEB-v2 benchmark covering 78 datasets, Tencent reports Hit@1 and NDCG@5 scores of 77.9, 79.2, and 80.6 for the 2B, 4B, and 9B models respectively, with the broader MMEB-v3 evaluation across 190 tasks yielding scores of 56.0, 58.2, and 59.5. The release includes inference examples, serving instructions, and evaluation code, with Tencent recommending specific library versions for reproducibility; the reported benchmark figures remain project-reported and await independent verification.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in