---
title: "What is semantic deduplication?"
description: "Semantic deduplication drops near-duplicate items by embedding similarity, not only by exact hash equality."
url: "https://pixeltable.com/learn/what-is-semantic-deduplication"
updated: "2026-09-29"
vertical: "Evaluation"
doc: "https://docs.pixeltable.com/datastore/embedding-index"
---

# What is semantic deduplication?

Semantic deduplication drops near-duplicate items by embedding similarity, not only by exact hash equality.

Updated: 2026-09-29
Part of [What is an embedding?](https://pixeltable.com/learn/what-is-an-embedding).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-semantic-deduplication#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-semantic-deduplication#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-semantic-deduplication#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-semantic-deduplication#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-semantic-deduplication#questions)

## How it works {#how-it-works}


- Embed the items.
- Find neighbors closer than a threshold.
- Keep one row from each cluster.

## What it is not {#what-it-is-not}

It is not a UNIQUE constraint on a filename, and it is not an MD5 of the file.

## semantic deduplication: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Catches | Near duplicates | Only identical bytes |
| Uses | Embeddings | A hash |
| Keeps | One representative | Every file |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable finds those neighbors with an embedding index, then a query or an aggregate picks the row to keep.

## Questions {#questions}

### How does semantic deduplication work? {#faq-1}

Embed the items. Find neighbors closer than a threshold. Keep one row from each cluster.

### What is semantic deduplication often confused with? {#faq-2}

It is not a UNIQUE constraint on a filename, and it is not an MD5 of the file.

## Related

- [Documentation](https://docs.pixeltable.com/datastore/embedding-index)
- [Stop Labeling Duplicates: Semantic Deduplication for Multimodal Datasets](https://pixeltable.com/blog/semantic-deduplication-curation)
- [Computer vision pipeline](https://pixeltable.com/use-cases/computer-vision-pipeline-optimization)
