---
title: "What is ML evaluation?"
description: "ML evaluation scores model or pipeline outputs against labels or a judge, so you can compare versions."
url: "https://pixeltable.com/learn/what-is-ml-evaluation"
updated: "2026-09-29"
vertical: "Evaluation"
doc: "https://docs.pixeltable.com/overview/quick-start"
---

# What is ML evaluation?

ML evaluation scores model or pipeline outputs against labels or a judge, so you can compare versions.

Updated: 2026-09-29


## On this page

- [How it works](https://pixeltable.com/learn/what-is-ml-evaluation#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-ml-evaluation#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-ml-evaluation#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-ml-evaluation#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-ml-evaluation#questions)

## How it works {#how-it-works}


- Fix the inputs at a table version.
- Score the output columns.
- Compare scores across versions.

## What it is not {#what-it-is-not}

It is not the database that stores those outputs. An eval framework is not a multimodal table.

## ML evaluation: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Scores | Outputs | The storage engine |
| Needs | Labels or a judge | Only a metric name |
| Lives beside | The table of outputs | Instead of the table |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable stores the outputs you score. It is not an eval product. A score can still be a column or an external judge reading the table.

## Questions {#questions}

### How does ML evaluation work? {#faq-1}

Fix the inputs at a table version. Score the output columns. Compare scores across versions.

### What is ML evaluation often confused with? {#faq-2}

It is not the database that stores those outputs. An eval framework is not a multimodal table.

## In the blog

- [You Don't Need an Eval Framework: You Need Data Infrastructure](https://pixeltable.com/blog/pixeltable-not-eval-framework)
- [Beyond AVG(): Building Custom Aggregation Functions for AI Workflows with Pixeltable UDA](https://pixeltable.com/blog/beyond-avg-custom-aggregations-uda)
- [DIY Scripts vs. Declarative Pipelines: Choosing the Right Framework for Your AI/ML Project](https://pixeltable.com/blog/declarative-vs-imperative-ai-pipelines)
- [Why Pixeltable is the Ultimate Agent Harness](https://pixeltable.com/blog/pixeltable-agent-harness)
- [What Is Jev? TypeSafe's System One Model — and Why the Decision Belongs in the Table](https://pixeltable.com/blog/jev-system-one-model)
- [From Excel to PyTorch: The Complete Guide to Converting Spreadsheets into Training Data](https://pixeltable.com/blog/excel-pandas-pytorch-training-data-guide)

## Related

- [Documentation](https://docs.pixeltable.com/overview/quick-start)
- [You Don't Need an Eval Framework: You Need Data Infrastructure](https://pixeltable.com/blog/pixeltable-not-eval-framework)
- [What ML Infrastructure Engineers Actually Want: Design Principles for Modern AI Data Platforms](https://pixeltable.com/blog/ml-infrastructure-design-principles-evaluation)
- [What Is Jev? TypeSafe's System One Model — and Why the Decision Belongs in the Table](https://pixeltable.com/blog/jev-system-one-model)
- [score](https://pixeltable.com/score)
