Skip to content

[ commentary ]

Ovis-VL-Embedding-9B puts text and documents in one search space

A 9B vision-language embedding model maps text, images, documents and video into one space, so one index can serve mixed business content.

Published · on huggingface.co · 2 min read

A 9B vision-language embedding model maps text, images, documents and video into one space, so one index can serve mixed business content.

The AI Search newsletter reported on 2026-09-27 that a vision-language embedding model, Ovis-VL-Embedding-9B, is published on huggingface.co. The page describes a model that turns text, images, visual documents, video and interleaved multimodal inputs into one representation space, so a single encoder can be used for cross-modal retrieval.

What the model does

According to the model page on huggingface.co, Ovis-VL-Embedding-9B is built from a Qwen3.5-9B backbone. The language-modelling head is removed, and the final-layer hidden state at the last non-padding token becomes the retrieval embedding. No modality-specific projection head is added, so the output width stays at 4096 dimensions. Queries and candidates are encoded separately, L2-normalised, and ranked by cosine similarity. It is a bi-encoder: no answer generation and no cross-attention between query and candidate during retrieval.

The page lists intended uses as multimodal search, multimodal RAG, image and visual-document retrieval, video search and temporal localisation, recommendation and nearest-neighbour matching. Classification labels, passages, images, documents, videos and interleaved items are all treated as candidates in the same space.

What a team would want to test

For work that involves reading incoming documents, the interesting part is the visual-document retrieval claim. A page of a scanned invoice, a tender or a contract can be indexed as an image rather than pushed through a separate OCR step first, and the same index can hold plain text. That is a possibility worth testing, not a result BARGO has measured.

The reported benchmark is MMEB-v2, 78 datasets across image, video and visual-document tasks. The page reports 81.13 overall, 1.04 points above the strongest compared baseline, with 83.96 on image and 83.06 on visual document, and 72.90 on video, which trails the strongest compared baseline by 3.05 points. These are the publisher's own reported figures.

Limitations matter for planning. This checkpoint does not support audio; the page points to Ovis-Embedding-Omni-3B for that. It produces embeddings only, not generated text. Quality depends on task-appropriate query instructions and the native preprocessing and chat template. The 4096-dimensional output and 9B backbone need more memory, storage and inference compute than the 2B variant. The page also notes that benchmark scores may not predict performance on a new domain, and recommends evaluating with representative queries, candidates and retrieval metrics before deployment.

For a team working on answers from company documents or on product data, the practical first step is a small retrieval test on your own material: a set of real queries, the documents or catalogue images they should return, and a measure of whether the right item comes back. The model is released under the Apache 2.0 licence, which makes that kind of internal trial straightforward to run.

Source: huggingface.co — BARGO’s commentary on the linked source.

Reported in the AI Search newsletter:

Which task costs your team the most hours every week? Tell us, and we will tell you whether AI can take it and what it would cost.

← all posts · RSS