---
title: "What is dataset curation?"
description: "Dataset curation selects, cleans, and versions the rows that train or evaluate a model."
url: "https://pixeltable.com/learn/what-is-dataset-curation"
updated: "2026-09-29"
vertical: "Evaluation"
doc: "https://docs.pixeltable.com/overview/quick-start"
---

# What is dataset curation?

Dataset curation selects, cleans, and versions the rows that train or evaluate a model.

Updated: 2026-09-29
Part of [What is active learning?](https://pixeltable.com/learn/what-is-active-learning).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-dataset-curation#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-dataset-curation#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-dataset-curation#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-dataset-curation#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-dataset-curation#questions)

## How it works {#how-it-works}


- Gather candidate rows.
- Drop exact and near duplicates, fix labels, and hold out a test slice.
- Pin the version the run used.

## What it is not {#what-it-is-not}

It is not an unversioned scrape, and it is not only storing weights.

## dataset curation: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Output | A versioned set of rows | A dump |
| Includes | Labels and media | Weights |
| Tools | Filters, embeddings, people | A spreadsheet with paths |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable curation is queries and columns on the media table: dedup, labels, and a version you can name.

## Questions {#questions}

### How does dataset curation work? {#faq-1}

Gather candidate rows. Drop exact and near duplicates, fix labels, and hold out a test slice. Pin the version the run used.

### What is dataset curation often confused with? {#faq-2}

It is not an unversioned scrape, and it is not only storing weights.

## In the blog

- [Stop Rebuilding Training Datasets: How Training Engineers Cut Model Development Time by 90%](https://pixeltable.com/blog/training-engineer-model-development-pytorch-integration)
- [Stop Labeling Duplicates: Semantic Deduplication for Multimodal Datasets](https://pixeltable.com/blog/semantic-deduplication-curation)
- [Beyond Pandas: Why Pixeltable Is the Ultimate Tool for Multimodal Data Wrangling](https://pixeltable.com/blog/pixeltable-vs-pandas-multimodal-data-wrangling)
- [From Data Chaos to Dataset Mastery: How ML Engineers Are Transforming Autonomous Vehicle Workflows](https://pixeltable.com/blog/ml-engineer-dataset-chaos-autonomous-vehicle-workflows)
- [Stop Juggling Tools: Why Modern AI Teams Are Moving Beyond Traditional Data Infrastructure](https://pixeltable.com/blog/traditional-tables-to-multimodal-ai)
- [Unified Multimodal AI Infrastructure: Escape Data Plumbing with Pixeltable](https://pixeltable.com/blog/unified-multimodal-ai-infrastructure-pixeltable)

## Related

- [Documentation](https://docs.pixeltable.com/overview/quick-start)
- [Stop Labeling Duplicates: Semantic Deduplication for Multimodal Datasets](https://pixeltable.com/blog/semantic-deduplication-curation)
- [From Excel to PyTorch: The Complete Guide to Converting Spreadsheets into Training Data](https://pixeltable.com/blog/excel-pandas-pytorch-training-data-guide)
- [From Data Chaos to Dataset Mastery: How ML Engineers Are Transforming Autonomous Vehicle Workflows](https://pixeltable.com/blog/ml-engineer-dataset-chaos-autonomous-vehicle-workflows)
- [Pixeltable vs Voxel51](https://pixeltable.com/compare/pixeltable-vs-voxel51)
