---
title: "Closing the Loop: Automating Active Learning Pipelines from Production to Training"
date: "2025-01-30"
author: "Marcel Kornacker"
tags:
  - Active Learning
  - MLOps
  - Computer Vision
  - Automation
description: "How to build automated feedback loops that capture edge cases in production and feed them back into training, without manual plumbing."
url: "https://pixeltable.com/blog/closing-the-loop-active-learning"
---

# Closing the Loop: Automating Active Learning Pipelines from Production to Training

The hardest part of ML isn't training the model. It's knowing *what* to train it on. Most teams treat training and production as separate worlds. Here's how to connect them into a continuous, automated active learning loop.
 

 
## The Production Gap

 
 

 You train a model on a curated dataset. It gets 95% accuracy. You deploy it to a drone or a security camera. Suddenly, it fails on "red trucks at night" or "people holding umbrellas."
 

 

 **The Manual Fix:**
 

 

 - Engineers notice the failure in logs.

 - Someone manually downloads the "bad" videos.

 - They upload them to S3.

 - They send links to a labeling team (Label Studio/Scale).

 - They wait for labels.

 - They manually merge the new labels into the training set.

 - They re-train.

 

 

 This cycle takes weeks. It should take minutes.
 

 
## Automating the Feedback Loop

 

 With Pixeltable, you can treat production inference logs as just another data source. Because Pixeltable handles both media storage and metadata, you can query for "hard examples" directly.
 

 
### Step 1: Capture Production Data

 
 

 Instead of just logging text, log the actual image/frame references to a Pixeltable table.
 

 
```python
# Production table
prod_logs = pxt.create_table('prod_logs', {
 'image': pxt.Image,
 'inference': pxt.Json,
 'confidence': pxt.Float
})
```

 
### Step 2: Identify Edge Cases

 
 

 We want to find images where the model was unsure (low confidence).
 

 
```python
# Create a view of "hard examples"
hard_examples = pxt.create_view(
 'hard_examples',
 prod_logs,
 filter=prod_logs.confidence < 0.6
)
```

 
### Step 3: Integrate with Labeling

 
 

 Pixeltable integrates directly with Label Studio. You can sync your "hard examples" view to a labeling project automatically.
 

 
```python
from pixeltable.integrations import label_studio

# Sync the view to a Label Studio project
ls_project = label_studio.create_project(
 title='Active Learning - Red Trucks',
 view=hard_examples
)
```

 
### Step 4: Merge and Retrain

 
 

 Once labeled, the annotations flow back into Pixeltable. You can create a unified training dataset that combines your original gold set with these new, high-value edge cases.
 

 
```python
# Combine original train set + new labeled examples
training_set = pxt.create_view(
 'training_set',
 original_data
)

# Add the new labels
training_set.insert(ls_project.annotations)

# Export for PyTorch
pxt.io.export_pytorch(
 training_set,
 format='coco'
)
```

 
## The Flywheel Effect

 

 By closing this loop, you turn your production environment into a data mining engine. Every failure becomes a training signal. The model gets smarter automatically, focusing exactly on the data it finds most difficult.
 

 

 This is how companies like Tesla and Waymo build data moats. With Pixeltable, you don't need a 50-person infrastructure team to build it.