Models & Algorithms•SOTAAZ Lab••KR

Is EmbeddingGemma 2 Better Than EmbeddingGemma 1 for Korean Search? We Tested It on 177 of Our Own Posts

We searched our 177 published Korean/English post pairs with each post's description as the query. With Korean queries, EmbeddingGemma 1 ranked the right post first more often than EmbeddingGemma 2: 162 vs 153 posts in Korean (p = 0.022) and 164 vs 155 when the target was the English version (p = 0.035). English-to-English showed no difference. Pre-registered, text only.

Is EmbeddingGemma 2 Better Than EmbeddingGemma 1 for Korean Search? We Tested It on 177 of Our Own Posts

Is EmbeddingGemma 2 Better Than EmbeddingGemma 1 for Korean Search? We Tested It on 177 of Our Own Posts

Google DeepMind released EmbeddingGemma 2 in October. It is a 740M-parameter open embedding model under Apache 2.0. A 270M text model sits at its core, and separate vision and audio encoders load only when you need them. The model card reports text results against the first generation: multilingual MTEB 61.36 vs 61.15, and code retrieval 78.68 vs 68.76. It does not report Korean on its own.

We run a bilingual blog, so we have a ready-made Korean test set. Every post exists in Korean and English, and every post has a search description written for it. We used each description as a query and checked whether the model ranked its own post first. Before any model ran, we committed the questions, decision rules and the sentence we would write for each outcome.

The setup

  • Corpus: every post published in both Korean and English on this site, which came to 177 pairs. We took bodies and descriptions from the published versions in our CMS and dropped the title block, so no model saw the title.
  • Queries: each post's description. The right answer is the post it describes.
  • Three tasks: Korean query → Korean posts (KK), Korean query → English posts (KE), and English query → English posts (EE) as a reference. Each task searches all 177 posts in the target language.
  • Chunks: paragraphs packed into chunks of up to 1,000 characters, each truncated at 512 tokens for every model. A post's score is its best chunk's score.
  • Models: EmbeddingGemma 2 (text only), EmbeddingGemma 1 (embeddinggemma-300m), Qwen3-Embedding-0.6B, and jina-embeddings-v5-omni-nano. Each used the query and document prompts from its model card. We also ran BM25 on character bigrams (Korean) and words (English). All models ran in float32, in one environment (transformers 5.19.0, sentence-transformers 6.1.0).
  • Decision rule: the main comparison is EmbeddingGemma 2 against 1, on KK and KE. We counted how often each model ranked the right post first (Recall@1) and compared with McNemar's exact test at p < 0.05. We also report MRR@10 with a 10,000-sample bootstrap interval.

These queries are easier than real user questions. A description shares words with the post it summarizes, which favors keyword search, and every model sees that same advantage.

The result: the first generation found more Korean posts

Three dot plots of how many of 177 posts each model ranked first. Korean to Korean: EmbeddingGemma 2 153, EmbeddingGemma 1 162, Qwen3-Embedding-0.6B 162, jina-v5-omni-nano 159, BM25 153. Korean to English: 155, 164, 156, 148, 130. English to English: 160, 162, 155, 148, 157.
TaskEmbeddingGemma 2EmbeddingGemma 1Only 2 rightOnly 1 rightp
Korean → Korean153 / 177162 / 1772110.022
Korean → English155 / 177164 / 1773120.035
English → English160 / 177162 / 177460.75

On both Korean tasks, EmbeddingGemma 1 ranked the right post first more often, and both differences pass the test. MRR@10 agrees: 2 is lower by 0.027 on KK (interval −0.049 to −0.008) and by 0.028 on KE (−0.051 to −0.006). On English we can't tell the two apart. The gap is in the top slot: both generations put the right post in their top five almost every time (175 vs 176 of 177 on Korean → Korean, 175 each on Korean → English). If your pipeline passes the top few results to a model, the two will look much closer than the top-1 numbers suggest.

The card makes no Korean-specific claim, so this doesn't contradict it. The card's own numbers put most of the second generation's gain in code (68.76 to 78.68), while multilingual text barely moved (61.15 to 61.36). On this Korean test, the older model came out ahead.

We checked three things before trusting this:

  • Prompts: EmbeddingGemma 2 received task: search result | query: and title: none | text: , as the card specifies.
  • Text-only loading: on a sample of 45 inputs (5 queries and 40 chunks), embeddings from the text-only load matched the full multimodal load exactly (maximum difference 0.0).
  • Truncation: both generations tokenize identically, so with the document prompt attached the same 259 of 1,763 Korean chunks were cut at 512 tokens for both. Truncation can't explain the gap.

The other models

Against EmbeddingGemma 2, only three differences pass the test:

  • Qwen3-Embedding-0.6B was higher on Korean → Korean (162 vs 153, p = 0.049). That sits at the edge of the threshold, and the MRR@10 interval for the same pair crosses zero (−0.050 to 0.001). It truncated more Korean chunks than the Gemma models (471 of 1,763) and still scored higher.
  • jina-v5-omni-nano was lower on English → English (148 vs 160, p = 0.012).
  • BM25 was lower on Korean → English (130 vs 155, p = 0.0005), which is expected: a Korean query shares almost no characters with an English post.

BM25 tied EmbeddingGemma 2 on Korean → Korean (153 each) and came within 3 on English. That says more about our queries than about embeddings. Descriptions reuse the post's own words, so keyword search gets a lot for free. Don't read it as "embeddings don't help."

Shrinking the vectors

Both generations are trained so you can keep only the first 512, 256 or 128 numbers of each 768-number vector. Before running, we set a tolerance: a cut "stays within" if the bootstrap interval for the MRR@10 drop stays below 0.03.

MRR@10 drop from 768Korean → KoreanKorean → English
EmbeddingGemma 2, 512within 0.03within 0.03
EmbeddingGemma 2, 256within 0.03can't say (0.012, −0.011 to 0.035)
EmbeddingGemma 2, 128can't say (0.019, −0.007 to 0.045)more than 0.03 (0.077, 0.043 to 0.112)
EmbeddingGemma 1, 512within 0.03within 0.03
EmbeddingGemma 1, 256can't say (0.017, −0.002 to 0.039)can't say (0.014, −0.003 to 0.033)
EmbeddingGemma 1, 128can't say (0.027, 0.002 to 0.053)can't say (0.019, −0.006 to 0.044)

The model card's own table shows 128 dimensions losing ground too: multilingual MTEB goes from 61.36 at full size to 57.89. Our Korean → English drop for EmbeddingGemma 2 at 128 points the same way, and on this task it was larger than our tolerance. On the same task, EmbeddingGemma 1's drop at 128 was 0.019, against 0.077 for EmbeddingGemma 2. At 128 it still ranked the right English post first for 159 Korean queries, more than EmbeddingGemma 2 managed at full size (155). We didn't pre-register or test either of these two comparisons, so treat both as observations.

What we'd do with this

If you already run EmbeddingGemma 1 for Korean text search, this test gives no reason to switch for Korean. The reasons to move to the second generation are images, audio, video and code, which we didn't test. If you're choosing a small Korean embedding model from scratch, EmbeddingGemma 1 and Qwen3-Embedding-0.6B tied for the most top-1 hits on Korean → Korean (162 each). On Korean → English, EmbeddingGemma 1 was ahead on its own (164), and Qwen3-Embedding-0.6B (156) was level with EmbeddingGemma 2 (155, p = 1.0). We didn't test EmbeddingGemma 1 against Qwen directly.

The cheapest step is the one we took: build a test set from documents you already have, where each item knows its own right answer. Our RAG evaluation post goes further into scoring.

Limits

  • One corpus of 177 bilingual posts, mostly about AI models and tools, with 12 pairs on SQL and data analysis, and descriptions as queries. Real questions are shorter and share fewer words with the answer.
  • Korean text lost more to the 512-token cut than English. With the document prompt, 259 of 1,763 Korean chunks were truncated for both Gemma models, against 11 of 2,425 English chunks. Qwen3-Embedding-0.6B truncated 471 Korean chunks.
  • 177 queries can't detect small differences. Most comparisons outside the main pair came out as "can't tell."
  • The posts are public, so any model may have seen them in training. We have no reason to think that favors one model.
  • Text only. EmbeddingGemma 2's main additions are images, audio and video, and this test says nothing about them.
  • jina-v5-omni-nano is licensed CC-BY-NC-4.0. We measured it and did not redistribute it. Its loader treats strings that start with http as media links, so our runner pins every string to the text path. On a sample of 40 plain-text chunks, this produced identical embeddings (maximum difference 0.0).

Everything is in drafts/ko-embed-bench/: the plan, committed before any model ran; the runner, per-query scores and analysis; verify_checks.py and its output checks.json for the three checks above; and RUNS.md, the run history. The plan's change log records three changes, each made before the runs it affects: two to how the corpus was built, before any model ran, and one to the runner before the jina run. The newsletter covers the next measurement when it is published.