Evaluation

What is semantic deduplication?

Semantic deduplication drops near-duplicate items by embedding similarity, not only by exact hash equality.

Updated · Part of What is an embedding?

How it works

  • Embed the items.
  • Find neighbors closer than a threshold.
  • Keep one row from each cluster.

What it is not

It is not a UNIQUE constraint on a filename, and it is not an MD5 of the file.

semantic deduplication: this, and the thing it is confused with

semantic deduplication: this, and the thing it is confused with
ThisNot this
CatchesNear duplicatesOnly identical bytes
UsesEmbeddingsA hash
KeepsOne representativeEvery file

Where Pixeltable fits

Pixeltable finds those neighbors with an embedding index, then a query or an aggregate picks the row to keep.

Questions

How does semantic deduplication work?
Embed the items. Find neighbors closer than a threshold. Keep one row from each cluster.
What is semantic deduplication often confused with?
It is not a UNIQUE constraint on a filename, and it is not an MD5 of the file.