Finding notes, photos and videos on an iPhone: what we measured
Reteum searches your notes, photos and videos with AI that runs on the phone, so nothing you write is sent anywhere to be understood. When Google released EmbeddingGemma 2, we tested it against the model Reteum uses today. This is what we found, language by language.
The short answer
- Notes: our fine-tuned model stays. It finds the right note first 67.9% of the time; EmbeddingGemma 2, even after the same fine-tuning, reaches 66.6%, and it groups notes into subjects less well.
- Photos: EmbeddingGemma 2 is much better than SigLIP 2, the model Reteum uses today: 68.7% against 54.1%, and twice as good in Japanese.
- Videos: EmbeddingGemma 2 reads a video as a whole, not as a few still frames, and finds the right one first 71.5% of the time against 65.5%.
- So a coming update keeps our model for notes and the Atlas, and moves photo and video search to EmbeddingGemma 2. Both run on the iPhone.
The models
- Reteum's text model: Google's EmbeddingGemma (300M) fine-tuned by us for short personal notes in many languages: typos, feelings, broader words ("cooking" for a soup recipe) and searching in one language for a note written in another.
- EmbeddingGemma 2: Google's new model, which reads text, photos and video. We tested it as released and after our own fine-tuning, done the same way as for our current model, including averaging several fine-tuned versions ("model soups").
- SigLIP 2 (base, patch 16, 224): Google's photo model that Reteum uses for photo and video search today.
How we tested
- Notes: 503 short notes and 3,018 searches in 15 languages, written for the test, not anyone's real notes. Their topics and people are kept out of the training data. A search counts when the right note comes first.
- Grouping: 600 notes on 60 topics never seen in training, grouped automatically and compared with the true topics (adjusted Rand index; 1 is perfect).
- Photos: XM3600, 3,600 photos from around the world, each described by people in many languages. The description is the search; the photo is the answer.
- Videos: 200 clips from MSR-VTT, searched with English descriptions.
Finding notes, per language
Share of searches where the right note comes first. Our model is the one in Reteum today. EG 2 is EmbeddingGemma 2 as released; EG 2 tuned is its best version after our fine-tuning.
| Language (searches) | Our model | EG 2 | EG 2 tuned |
|---|---|---|---|
| English (974) | 76.4% | 71.3% | 73.5% |
| Simplified Chinese (290) | 64.5% | 60.7% | 62.1% |
| Japanese (280) | 64.6% | 56.8% | 61.8% |
| Korean (244) | 64.8% | 59.4% | 65.6% |
| Traditional Chinese (235) | 56.6% | 53.6% | 58.3% |
| Spanish (168) | 65.5% | 58.9% | 63.7% |
| French (140) | 70.0% | 61.4% | 66.4% |
| Arabic (139) | 58.3% | 60.4% | 65.5% |
| German (134) | 64.2% | 62.7% | 65.7% |
| Mixed languages (132) | 66.7% | 63.6% | 63.6% |
| Portuguese (126) | 65.9% | 62.7% | 67.5% |
| Russian (56) | 64.3% | 60.7% | 66.1% |
| Hindi (51) | 66.7% | 62.7% | 64.7% |
| Swahili (46) | 60.9% | 37.0% | 47.8% |
| All 3,018 | 67.9% | 63.0% | 66.6% |
The 3 Italian searches are left out of the table, as too few to mean anything; they are in the total. The fine-tuned column is the best of the nine EmbeddingGemma 2 versions we trained and blended; the others scored between 64.7% and 66.5% overall.
Finding notes, by kind of search
The same three models, on four kinds of search that are hard for any model.
| Kind of search | Our model | EG 2 | EG 2 tuned |
|---|---|---|---|
| A broader word ("cooking") | 38.0% | 28.0% | 38.4% |
| A feeling ("when I was anxious") | 14.5% | 9.6% | 10.6% |
| With a typo | 53.0% | 54.2% | 60.1% |
| In another language than the note | 83.8% | 79.6% | 79.8% |
Searching by feeling is the hardest for every model: a note rarely names the feeling it was written in.
Grouping notes into subjects
Reteum's Atlas sorts notes into subjects on its own. On 600 notes about 60 topics none of the models had seen, the adjusted Rand index was 0.499 for our model, 0.399 for EmbeddingGemma 2 as released and 0.442 at best after fine-tuning. That gap is the main reason our model stays for notes.
Finding photos, per language
XM3600, 3,600 photos. Share of searches where the right photo comes first. EmbeddingGemma 2 lets you choose how much detail it takes from a photo; 140 visual tokens was as accurate as 280 at half the work, so that is what we will use.
| Language | SigLIP 2 (today) | EG 2 |
|---|---|---|
| English | 55.0% | 59.5% |
| Japanese | 37.8% | 81.1% |
| Chinese | 43.1% | 70.0% |
| Korean | 50.1% | 65.1% |
| Spanish | 62.9% | 66.7% |
| German | 71.1% | 81.0% |
| French | 67.7% | 75.0% |
| Arabic | 40.9% | 55.3% |
| Portuguese | 57.9% | 64.2% |
| Average | 54.1% | 68.7% |
Fine-tuning EmbeddingGemma 2 for notes left photo search where it was (68.4% to 68.6%), so photos can use the model as released. XM3600 is a public set and may have been part of EmbeddingGemma 2's training, so we are also checking with our own photos.
Finding videos
200 MSR-VTT clips, searched with English descriptions.
| How the video is read | Right one first | In the top 5 |
|---|---|---|
| SigLIP 2, 4 still frames (today) | 65.5% | 84.0% |
| EmbeddingGemma 2, the same 4 frames | 69.0% | 85.0% |
| EmbeddingGemma 2, whole video, 140 tokens a frame | 71.5% | 90.0% |
| EmbeddingGemma 2, whole video, 70 tokens a frame | 70.5% | 91.5% |
Reading the whole video at 70 tokens a frame is almost as accurate and more than twice as fast, so that is what we will use.
What this means for you
- Searching notes in English, Chinese, Japanese and Spanish stays with the model that does best there, and the Atlas keeps its better grouping.
- Searching photos and videos by what they show gets clearly better, most of all in Japanese, Chinese and Korean.
- All of it runs on your iPhone. Your notes, photos and videos are never sent to a server to be searched. See the privacy policy.
More on how Reteum works: a second brain that runs on your iPhone and keeping notes from piling up.
See how Reteum works For iPhone with iOS 26 · 7 days free