Evaluation
What is semantic deduplication?
Semantic deduplication drops near-duplicate items by embedding similarity, not only by exact hash equality.
Updated · Part of What is an embedding?
How it works
- Embed the items.
- Find neighbors closer than a threshold.
- Keep one row from each cluster.
What it is not
It is not a UNIQUE constraint on a filename, and it is not an MD5 of the file.
semantic deduplication: this, and the thing it is confused with
| This | Not this | |
|---|---|---|
| Catches | Near duplicates | Only identical bytes |
| Uses | Embeddings | A hash |
| Keeps | One representative | Every file |
Where Pixeltable fits
Pixeltable finds those neighbors with an embedding index, then a query or an aggregate picks the row to keep.
Questions
- How does semantic deduplication work?
- Embed the items. Find neighbors closer than a threshold. Keep one row from each cluster.
- What is semantic deduplication often confused with?
- It is not a UNIQUE constraint on a filename, and it is not an MD5 of the file.