New models for the animal-ml sidecar: a retrained dog detector and a new identity embedder. Both were chosen and tested on DogReID-1553, 1,553 dogs filmed by their owners on phones, which is much closer to a real Immich library than the datasets earlier releases were tested on.
What's new
Many more of your dogs end up in a correct album. We simulated 300 households using dogs the models never saw during training. Per 100 dogs:
| 0.1.1 | 0.2.0 | |
|---|---|---|
| Dogs grouped into one correct person | 15 | 47 |
| Dogs never grouped | 62 | 27 |
| Dogs merged with another dog from the same home | 5 | 10 |
| Photos in a person that belong to that dog | 95.0% | 94.1% |
On named dogs from Wikimedia Commons, the share of dog photos left out of any person drops from 74% to 28%.
The detector finds more dogs. It was fine-tuned with owner phone photos from DogReID.
- Owner phone photos with the dog found: 83% → 95%
- In-the-wild dog photos (Wikimedia Commons and MPDD) with the dog found: 85% → 88%
- Dog-free photos with a false detection: 16% → 22%. Most of these are other four-legged animals: 71% of wolf and fox photos, 11% of cat photos, and some goats, deer and sheep. People, horses, landscapes and buildings almost never trigger it.
One embedder for everyone. 0.1.x served a smaller ResNet50 model to users who picked a smaller Immich face model. 0.2.0 has a single embedder, DINOv2-B/14, whatever face model Immich uses.
Cost. The embedder takes about 130 ms per dog on 4 CPU threads, about 4× the 0.1.x embedder. For a photo with one dog, the sidecar's total work is about 1.1× Immich's own buffalo_l face pipeline. The models add about 130 MB to the image.
Upgrading from 0.1.x
0.2.0's dog embeddings can't be compared with 0.1.x's. Your existing dog faces have to be removed first so the new model replaces all of them. Dog names and merges are lost; human faces are untouched.
- Administration → Settings → Machine Learning: set the URL back to
http://immich-machine-learning:3003. - Administration → Job Queues → Face Detection → Refresh, then wait for it to finish. This removes the 0.1.x dog faces.
- Change the image tag to
ghcr.io/rtp4jc/animal-ml:0.2.0, then rundocker compose up -d animal-ml. - Set the Machine Learning URL to
http://animal-ml:3003again. - Face Detection → Refresh again, then name your dogs.
If you set DOG_MAX_DISTANCE yourself, remove it: the default is now 0.4, which is tuned for the new embedder.
The animal-ml sidecar adds your dogs to Immich's People tab.
Immich already detects human faces and groups them into people. This runs alongside it and does the same for individual dogs.
This beta is about dogs. Cats and other animals are not supported yet. A cat is occasionally detected, but that is not a goal of this release.
Before you start
Use Refresh on the Face Detection queue, not Reset. Refresh keeps every face already in your library, so names, merges and hidden people survive both adding the sidecar and removing it. Reset deletes all named faces and manual merges.
Setup
Next to your Immich docker-compose.yml, create docker-compose.override.yml:
services:
animal-ml:
container_name: animal_ml
image: ghcr.io/rtp4jc/animal-ml:0.2.0
environment:
UPSTREAM_ML_URL: http://immich-machine-learning:3003
restart: alwaysdocker compose up -d animal-mlThen go to Administration → Settings → Machine Learning and set the URL to http://animal-ml:3003. Leave Min Detection Score and Max Distance alone, because the sidecar uses its own values for dogs.
Finally, run Administration → Job Queues → Face Detection → Refresh. This takes a long time on a large library.
Full instructions: sidecar/README.md
What to expect
- Dogs with plenty of photos get one large person, plus a few strays to merge
- Dogs with only a handful of photos may not group at all
- Similar-looking dogs, especially from the same home, sometimes get merged
- About one cat photo in nine is detected as a person
The full numbers are in the model card.
How
The sidecar runs as a docker container next to Immich's other containers. It intercepts Immich's face detection requests and finds the dogs in each photo, computing an identity embedding for each one. It then calls the real Immich ML container for the human faces and returns both sets of results together.
Assets
The model files are attached here and baked into the published image. Verify them with sha256sum -c SHA256SUMS.
Licence
AGPL-3.0, because the detector is fine-tuned from Ultralytics YOLO11. Some of the training data is licensed for non-commercial research and personal use only; see NOTICE.

