What's new?
This release adds support for EmbeddingGemma 2. EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.
Text search
The feature-extraction pipeline returns normalized embeddings, so the dot product of two embeddings is their cosine similarity:
import { pipeline, matmul } from "@huggingface/transformers";
const extractor = await pipeline("feature-extraction", "onnx-community/embeddinggemma-2-ONNX", {
device: "webgpu", // or "wasm" (browser) / "cpu" (Node.js)
dtype: "q4", // see "Choosing a dtype" below
});
const query = "task: search result | query: Which planet is known as the Red Planet?";
const documents = [
"title: none | text: Venus is often called Earth's twin because of its similar size and proximity.",
"title: none | text: Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"title: none | text: Jupiter, the largest planet in our solar system, has a prominent red spot.",
"title: none | text: Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
];
const embeddings = await extractor([query, ...documents], { pooling: "mean", normalize: true });
const query_embedding = embeddings.slice([0, 1]);
const document_embeddings = embeddings.slice([1, null]);
const scores = (await matmul(query_embedding, document_embeddings.transpose(1, 0))).tolist()[0];
const ranking = scores.map((score, i) => ({ score, document: documents[i] })).sort((a, b) => b.score - a.score);
console.log(ranking);
// [
// { score: 0.854, document: "title: none | text: Mars, known for its reddish appearance, ..." },
// { score: 0.783, document: "title: none | text: Saturn, famous for its rings, ..." },
// { score: 0.752, document: "title: none | text: Jupiter, the largest planet in our solar system, ..." },
// { score: 0.684, document: "title: none | text: Venus is often called Earth's twin, ..." },
// ]Images, audio and video
For other modalities, use the processor and the model directly. Every input maps into the same embedding space, so any embedding can be compared with any other: here, text queries against an image, an audio clip and a video.
import { AutoModel, AutoProcessor, load_image, load_audio, load_video, cat, matmul } from "@huggingface/transformers";
const model_id = "onnx-community/embeddinggemma-2-ONNX";
const processor = await AutoProcessor.from_pretrained(model_id);
const model = await AutoModel.from_pretrained(model_id, { device: "webgpu", dtype: "q4" });
// The processor takes (text, images, audio, videos)
const embed = async (...inputs) => (await model(await processor(...inputs))).sentence_embedding;
// Text queries, with a task prefix (see "Task Instruction Prefixes" below)
const queries = [
"task: search result | query: cats sleeping on a couch",
"task: search result | query: a president's speech about serving your country",
"task: search result | query: a turtle swimming in the ocean",
];
const query_embeddings = await embed(queries);
// An image, an audio clip (mono, 16 kHz) and a video (1 frame per second)
const url = "https://huggingface.co/datasets/Xenova/transformers.js-docs/resolve/main";
const image = await load_image(`${url}/cats.jpg`);
const audio = await load_audio(`${url}/jfk.wav`, 16000);
const video = await load_video(`${url}/sea-turtle.mp4`, { fps: 1 });
const media_embeddings = cat([
await embed(null, image),
await embed(null, null, audio),
await embed(null, null, null, video),
]);
// Embeddings are normalized: the dot product is the cosine similarity
const scores = (await matmul(media_embeddings, query_embeddings.transpose(1, 0))).tolist();
["image", "audio", "video"].forEach((name, i) => console.log(name, scores[i].map((x) => x.toFixed(3))));
// Each input scores highest with its own query:
// cats speech turtle
// image ["0.740", "0.462", "0.508"]
// audio ["0.504", "0.763", "0.487"]
// video ["0.505", "0.500", "0.731"]To embed several items of a modality at once, pass one list per input: processor(null, [[image1], [image2]]) returns two image embeddings, while a flat list of images, processor(null, [image1, image2]), is a single input made of both images.
Full Changelog: 4.3.0...4.3.1