---
title: "From Excel to PyTorch: The Complete Guide to Converting Spreadsheets into Training Data"
date: "2025-01-23"
author: "Pixeltable Team"
tags:
  - PyTorch
  - Excel
  - Training Data
  - Data Import
  - ML Training
  - Pandas
  - Dataset Preparation
  - Machine Learning
  - Data Engineering
  - Model Training
description: "Learn how to import data from Excel files and convert them into PyTorch-ready training datasets. Complete tutorial covering Excel/CSV import, data validation, preprocessing, and PyTorch Dataset creation with Pixeltable for production-ready ML workflows."
url: "https://pixeltable.com/blog/excel-pandas-pytorch-training-data-guide"
---

# From Excel to PyTorch: The Complete Guide to Converting Spreadsheets into Training Data

## The Excel to PyTorch Challenge: Bridging Traditional Data and Modern ML

 
Many AI/ML projects start with data in familiar formats: Excel spreadsheets, CSV files, or database exports. But getting this data into a format PyTorch can use for training often requires writing tedious boilerplate code, manual data validation, and custom preprocessing pipelines. This tutorial shows you the complete path from Excel sheets to production-ready PyTorch datasets.

 
 
We'll cover three approaches: traditional pandas/PyTorch, advanced Pixeltable workflows, and best practices for [production training pipelines](/blog/training-engineer-model-development-pytorch-integration).

 
## Traditional Approach: Pandas + PyTorch Manual Pipeline

 
Let's start with the conventional method to understand the challenges:

 
### Step 1: Import Excel Data with Pandas

 
```python

import pandas as pd
import torch
from torch.utils.data import Dataset, DataLoader
import numpy as np

# Import Excel file (supports .xlsx, .xls, .csv)
df = pd.read_excel('training_data.xlsx')

# Or for CSV
df = pd.read_csv('training_data.csv')

# Inspect the data
print(f"Shape: {df.shape}")
print(f"Columns: {df.columns.tolist()}")
print(df.head())

# Example output:
# image_path category label quality_score
# 0 images/img1.jpg cat 0 0.95
# 1 images/img2.jpg dog 1 0.87
# 2 images/img3.jpg cat 0 0.92
 
```

 
### Step 2: Data Validation and Cleaning

 
```python

# Manual data validation
def validate_training_data(df):
 """Validate imported data for ML training"""
 issues = []
 
 # Check for missing values
 missing_counts = df.isnull().sum()
 if missing_counts.any():
 issues.append(f"Missing values: {missing_counts[missing_counts > 0].to_dict()}")
 
 # Check for duplicate entries
 duplicates = df.duplicated().sum()
 if duplicates > 0:
 issues.append(f"Duplicate rows: {duplicates}")
 
 # Validate file paths exist
 import os
 if 'image_path' in df.columns:
 missing_files = []
 for path in df['image_path']:
 if not os.path.exists(path):
 missing_files.append(path)
 if missing_files:
 issues.append(f"Missing files: {len(missing_files)}")
 
 # Check label distribution
 if 'label' in df.columns:
 label_counts = df['label'].value_counts()
 print(f"Label distribution:
{label_counts}")
 
 # Warn about class imbalance
 if label_counts.max() / label_counts.min() > 10:
 issues.append("Severe class imbalance detected")
 
 return issues

validation_issues = validate_training_data(df)
if validation_issues:
 print("⚠️ Validation Issues:")
 for issue in validation_issues:
 print(f" - {issue}")

# Clean the data
df_clean = df.dropna() # Remove rows with missing values
df_clean = df_clean.drop_duplicates() # Remove duplicates
 
```

 
### Step 3: Create PyTorch Dataset

 
```python

from torch.utils.data import Dataset
from PIL import Image
import torchvision.transforms as transforms

class ExcelToTorchDataset(Dataset):
 """Convert Excel/CSV data to PyTorch Dataset"""
 
 def __init__(self, dataframe, image_column='image_path', label_column='label', transform=None):
 """
 Args:
 dataframe: Pandas DataFrame with training data
 image_column: Column name containing image file paths
 label_column: Column name containing labels
 transform: Optional torchvision transforms
 """
 self.df = dataframe.reset_index(drop=True)
 self.image_column = image_column
 self.label_column = label_column
 self.transform = transform
 
 # Validate columns exist
 if image_column not in self.df.columns:
 raise ValueError(f"Column '{image_column}' not found")
 if label_column not in self.df.columns:
 raise ValueError(f"Column '{label_column}' not found")
 
 def __len__(self):
 return len(self.df)
 
 def __getitem__(self, idx):
 # Get row data
 row = self.df.iloc[idx]
 
 # Load image
 image_path = row[self.image_column]
 try:
 image = Image.open(image_path).convert('RGB')
 except Exception as e:
 print(f"Error loading {image_path}: {e}")
 # Return dummy data on error (not ideal!)
 image = Image.new('RGB', (224, 224), color='black')
 
 # Get label
 label = row[self.label_column]
 
 # Apply transforms
 if self.transform:
 image = self.transform(image)
 
 # Convert to dict for flexibility
 sample = {
 'image': image,
 'label': torch.tensor(label, dtype=torch.long),
 'image_path': image_path,
 }
 
 # Add optional metadata columns
 for col in self.df.columns:
 if col not in [self.image_column, self.label_column]:
 sample[col] = row[col]
 
 return sample

# Define transforms
train_transform = transforms.Compose([
 transforms.Resize((224, 224)),
 transforms.RandomHorizontalFlip(),
 transforms.ColorJitter(brightness=0.2, contrast=0.2),
 transforms.ToTensor(),
 transforms.Normalize(mean=[0.485, 0.456, 0.406], 
 std=[0.229, 0.224, 0.225])
])

# Create dataset
dataset = ExcelToTorchDataset(
 dataframe=df_clean,
 image_column='image_path',
 label_column='label',
 transform=train_transform
)

# Create DataLoader
dataloader = DataLoader(
 dataset,
 batch_size=32,
 shuffle=True,
 num_workers=4
)

# Test it
print(f"Dataset size: {len(dataset)}")
batch = next(iter(dataloader))
print(f"Batch image shape: {batch['image'].shape}")
print(f"Batch labels shape: {batch['label'].shape}")
 
```

 
## Advanced Preprocessing: Handling Complex Excel Data

 
### Working with Multiple Feature Columns

 
```python

# Excel with multiple feature columns
# image_path, feature1, feature2, feature3, label, metadata

# Convert multiple columns to feature tensor
class MultiColumnDataset(Dataset):
 def __init__(self, dataframe, image_col, feature_cols, label_col, transform=None):
 self.df = dataframe
 self.image_col = image_col
 self.feature_cols = feature_cols
 self.label_col = label_col
 self.transform = transform
 
 def __getitem__(self, idx):
 row = self.df.iloc[idx]
 
 # Load image
 image = Image.open(row[self.image_col]).convert('RGB')
 if self.transform:
 image = self.transform(image)
 
 # Combine numerical features
 features = torch.tensor(
 [row[col] for col in self.feature_cols],
 dtype=torch.float32
 )
 
 # Get label
 label = torch.tensor(row[self.label_col], dtype=torch.long)
 
 return {
 'image': image,
 'features': features,
 'label': label
 }

# Usage
dataset = MultiColumnDataset(
 dataframe=df,
 image_col='image_path',
 feature_cols=['brightness', 'contrast', 'sharpness'],
 label_col='category_id',
 transform=train_transform
)
 
```

 
### Handling Categorical Variables

 
```python

# Convert categorical columns to numeric
from sklearn.preprocessing import LabelEncoder

# Encode string labels to integers
label_encoder = LabelEncoder()
df['label_encoded'] = label_encoder.fit_transform(df['category'])

# Save encoder for later use
import joblib
joblib.dump(label_encoder, 'label_encoder.pkl')

# One-hot encoding for multi-class features
df_encoded = pd.get_dummies(df, columns=['weather_condition', 'time_of_day'])

print(f"Original columns: {df.columns.tolist()}")
print(f"Encoded columns: {df_encoded.columns.tolist()}")
 
```

 
## The Pixeltable Approach: Declarative Data Import and Training

 
Pixeltable transforms the Excel-to-PyTorch workflow from manual scripting to [declarative data management](/blog/declarative-multimodal-incremental):

 
### Step 1: Import Excel to Pixeltable

 
```python

import pixeltable as pxt
import pandas as pd

# Read Excel with pandas first (or use Pixeltable's import)
df = pd.read_excel('training_data.xlsx')

# Create Pixeltable table with proper schema
training_data = pxt.create_table('ml_training.dataset', {
 'image': pxt.Image,
 'category': pxt.String,
 'label': pxt.Int,
 'quality_score': pxt.Float,
 'metadata': pxt.Json
})

# Bulk insert from DataFrame
records = []
for _, row in df.iterrows():
 records.append({
 'image': row['image_path'],
 'category': row['category'],
 'label': row['label'],
 'quality_score': row['quality_score'],
 'metadata': {
 'source': 'excel_import',
 'original_index': int(row.name)
 }
 })

training_data.insert(records)

print(f"✓ Imported {training_data.count()} records to Pixeltable")
 
```

 
### Step 2: Automatic Validation with Computed Columns

 
```python

# Add validation as computed columns
@pxt.udf
def validate_image_quality(image: pxt.Image, quality_score: float) -> dict:
 """Validate image meets training requirements"""
 if image is None:
 return {'valid': False, 'reason': 'image_missing'}
 
 width, height = image.size
 
 # Check minimum dimensions
 if width 3.0 or aspect_ratio pxt.Image:
 """Preprocess images for training"""
 from PIL import Image
 
 # Resize
 resized = image.resize(target_size, Image.Resampling.LANCZOS)
 
 # Convert to RGB if needed
 if resized.mode != 'RGB':
 resized = resized.convert('RGB')
 
 return resized

valid_data.add_computed_column(
 preprocessed_image=preprocess_for_training(valid_data.image)
)

# Calculate training statistics
@pxt.udf
def calculate_normalization_stats(images: list) -> dict:
 """Calculate mean and std for normalization"""
 import numpy as np
 from PIL import Image
 
 all_pixels = []
 for img in images[:1000]: # Sample for efficiency
 img_array = np.array(img) / 255.0
 all_pixels.append(img_array)
 
 all_pixels = np.stack(all_pixels)
 
 return {
 'mean': all_pixels.mean(axis=(0, 1, 2)).tolist(),
 'std': all_pixels.std(axis=(0, 1, 2)).tolist()
 }

# Query for normalization stats
images_sample = valid_data.select(valid_data.preprocessed_image).limit(1000).collect()
norm_stats = calculate_normalization_stats([r['preprocessed_image'] for r in images_sample])

print(f"Dataset normalization stats:")
print(f" Mean: {norm_stats['mean']}")
print(f" Std: {norm_stats['std']}")
 
```

 
### Step 4: Export to PyTorch Dataset

 
```python

# Export Pixeltable table to PyTorch Dataset
from torchvision import transforms

# Define transforms using calculated statistics
training_transforms = transforms.Compose([
 transforms.ToTensor(),
 transforms.Normalize(
 mean=norm_stats['mean'],
 std=norm_stats['std']
 ),
 transforms.RandomHorizontalFlip(p=0.5),
 transforms.RandomRotation(degrees=15)
])

# Pixeltable to PyTorch Dataset
pytorch_dataset = valid_data.to_pytorch_dataset(
 image_column='preprocessed_image',
 label_column='label',
 transform=training_transforms
)

# Create DataLoader
train_loader = DataLoader(
 pytorch_dataset,
 batch_size=32,
 shuffle=True,
 num_workers=4,
 pin_memory=True
)

# Verify
print(f"PyTorch Dataset created: {len(pytorch_dataset)} samples")
batch = next(iter(train_loader))
print(f"Batch shapes - Image: {batch['image'].shape}, Label: {batch['label'].shape}")
 
```

 
## Handling Complex Scenarios

 
### Multiple Excel Sheets

 
```python

# Import multiple sheets from Excel
excel_file = 'multi_sheet_data.xlsx'

# Read all sheets
sheet_dict = pd.read_excel(excel_file, sheet_name=None)

print(f"Found sheets: {list(sheet_dict.keys())}")

# Combine sheets with metadata
combined_records = []
for sheet_name, sheet_df in sheet_dict.items():
 for _, row in sheet_df.iterrows():
 record = {
 'image': row['image_path'],
 'label': row['label'],
 'source_sheet': sheet_name, # Track origin
 'metadata': {
 'sheet': sheet_name,
 'row_number': int(row.name)
 }
 }
 combined_records.append(record)

# Import to Pixeltable
training_data.insert(combined_records)
 
```

 
### Text and Image Multimodal Data

 
```python

# Excel with both text and images
# Columns: image_path, caption, category, label

multimodal_df = pd.read_excel('multimodal_training.xlsx')

# Create multimodal Pixeltable table
multimodal_data = pxt.create_table('ml_training.multimodal', {
 'image': pxt.Image,
 'caption': pxt.String,
 'category': pxt.String,
 'label': pxt.Int
})

# Import
multimodal_records = []
for _, row in multimodal_df.iterrows():
 multimodal_records.append({
 'image': row['image_path'],
 'caption': row['caption'],
 'category': row['category'],
 'label': row['label']
 })

multimodal_data.insert(multimodal_records)

# Add text embeddings with Pixeltable
from pixeltable.functions import openai

multimodal_data.add_computed_column(
 caption_embedding=openai.embeddings(
 multimodal_data.caption,
 model='text-embedding-3-small'
 )
)

# Custom PyTorch Dataset for multimodal training
class MultimodalDataset(Dataset):
 def __init__(self, pixeltable_data, transform=None):
 self.data = pixeltable_data.select(
 pixeltable_data.image,
 pixeltable_data.caption_embedding,
 pixeltable_data.label
 ).collect()
 self.transform = transform
 
 def __len__(self):
 return len(self.data)
 
 def __getitem__(self, idx):
 row = self.data[idx]
 
 # Process image
 image = row['image']
 if self.transform:
 image = self.transform(image)
 
 # Get text embedding (already computed by Pixeltable)
 text_embedding = torch.tensor(row['caption_embedding'], dtype=torch.float32)
 
 # Get label
 label = torch.tensor(row['label'], dtype=torch.long)
 
 return {
 'image': image,
 'text_embedding': text_embedding,
 'label': label
 }
 
```

 
## Data Augmentation Strategies

 
### Defining Augmentation Rules in Excel

 
```python

# Excel with augmentation parameters
# image_path, label, augment_rotate, augment_flip, augment_brightness

aug_df = pd.read_excel('training_with_augmentation.xlsx')

# Create dynamic transforms based on Excel specifications
class DynamicAugmentationDataset(Dataset):
 def __init__(self, dataframe):
 self.df = dataframe
 
 def __getitem__(self, idx):
 row = self.df.iloc[idx]
 
 # Load image
 image = Image.open(row['image_path']).convert('RGB')
 
 # Build transform chain from Excel specifications
 transform_list = [transforms.ToTensor()]
 
 if row.get('augment_rotate', False):
 transform_list.append(
 transforms.RandomRotation(degrees=row.get('rotation_degrees', 15))
 )
 
 if row.get('augment_flip', False):
 transform_list.append(
 transforms.RandomHorizontalFlip(p=1.0)
 )
 
 if row.get('augment_brightness', False):
 transform_list.append(
 transforms.ColorJitter(brightness=row.get('brightness_factor', 0.2))
 )
 
 transform_list.append(
 transforms.Normalize(mean=[0.485, 0.456, 0.406], 
 std=[0.229, 0.224, 0.225])
 )
 
 # Apply transforms
 transform = transforms.Compose(transform_list)
 image = transform(image)
 
 return {
 'image': image,
 'label': torch.tensor(row['label'], dtype=torch.long)
 }
 
```

 
## Production-Ready Pipeline with Pixeltable

 
For [production training workflows](/blog/training-engineer-model-development-pytorch-integration), Pixeltable provides comprehensive data management:

 
### Complete Excel → Pixeltable → PyTorch Workflow

 
```python

import pixeltable as pxt
import pandas as pd
from datetime import datetime

# 1. Import Excel data to Pixeltable with full tracking
df = pd.read_excel('production_training_data.xlsx')

production_dataset = pxt.create_table('production.training_v1', {
 'image': pxt.Image,
 'label': pxt.Int,
 'metadata': pxt.Json,
 'import_date': pxt.Timestamp
})

# Import with metadata
import_records = []
for _, row in df.iterrows():
 import_records.append({
 'image': row['image_path'],
 'label': row['label'],
 'metadata': {
 'excel_source': 'production_training_data.xlsx',
 'original_row': int(row.name),
 'quality_score': float(row.get('quality_score', 1.0))
 },
 'import_date': datetime.now()
 })

production_dataset.insert(import_records)

# 2. Add validation and preprocessing
@pxt.udf
def comprehensive_validation(image: pxt.Image, label: int, metadata: dict) -> dict:
 """Production-grade validation"""
 import numpy as np
 
 if image is None:
 return {'valid': False, 'reason': 'missing_image', 'severity': 'critical'}
 
 width, height = image.size
 validations = {
 'valid': True,
 'checks': []
 }
 
 # Dimension check
 if width 100: # Assuming max 100 classes
 validations['valid'] = False
 validations['checks'].append('invalid_label')
 
 # Image content check (detect corrupt images)
 try:
 img_array = np.array(image)
 if img_array.std() dict:
 """Determine augmentation based on class and image properties"""
 width, height = image.size
 aspect_ratio = width / height
 
 # Different augmentation for different scenarios
 if aspect_ratio > 2.0: # Wide images (panoramas)
 return {
 'horizontal_flip': False, # Don't flip panoramas
 'rotation': 5, # Minimal rotation
 'crop': True
 }
 else:
 return {
 'horizontal_flip': True,
 'rotation': 15,
 'crop': True,
 'color_jitter': True
 }

valid_training_data.add_computed_column(
 augmentation_strategy=determine_augmentation_strategy(
 valid_training_data.label,
 valid_training_data.image
 )
)

# 5. Create versioned snapshot for reproducibility
training_snapshot = pxt.create_snapshot(
 'production_training_v1.0_2025_01_23',
 valid_training_data
)

print(f"✓ Created immutable training snapshot")
print(f" Valid samples: {training_snapshot.count()}")
print(f" Snapshot URI: pxt://production_training_v1.0_2025_01_23")

# 6. Export to PyTorch with full lineage
pytorch_dataset = training_snapshot.to_pytorch_dataset(
 image_column='image',
 label_column='label'
)

print(f"✓ PyTorch dataset ready: {len(pytorch_dataset)} samples")
 
```

 
## Handling Common Import Errors

 
### Missing Image Files

 
```python

# Identify and handle missing files
missing_files = training_data.select(
 training_data.image,
 training_data.validation_result
).where(
 training_data.validation_result['reason'] == 'image_missing'
).collect()

if missing_files:
 print(f"⚠️ Found {len(missing_files)} missing image files")
 
 # Export list for correction
 missing_df = pd.DataFrame([
 {'image_path': r['image'], 'status': 'missing'}
 for r in missing_files
 ])
 missing_df.to_csv('missing_files_report.csv', index=False)
 
 print(" Exported missing files report to missing_files_report.csv")
 
```

 
### Label Inconsistencies

 
```python

# Detect label distribution issues
label_distribution = training_data.select(
 training_data.label,
 count=pxt.functions.count()
).group_by(
 training_data.label
).collect()

print("Label Distribution:")
for item in label_distribution:
 print(f" Label {item['label']}: {item['count']} samples")

# Identify under-represented classes
min_samples = min(item['count'] for item in label_distribution)
max_samples = max(item['count'] for item in label_distribution)

imbalance_ratio = max_samples / min_samples
if imbalance_ratio > 5:
 print(f"⚠️ Class imbalance detected: {imbalance_ratio:.1f}x")
 print(" Consider: stratified sampling or weighted loss function")
 
```

 
## Creating Train/Validation Splits

 
### Stratified Split with Pixeltable

 
```python

# Create stratified train/validation split
@pxt.udf
def assign_split(label: int, random_seed: int = 42) -> str:
 """Assign train/val split maintaining class balance"""
 import random
 random.seed(random_seed + label) # Consistent per-class split
 
 return 'train' if random.random() dict:
 """Extract features for product classification"""
 from PIL import ImageStat
 import numpy as np
 
 # Color histogram
 hist = image.histogram()
 
 # Basic stats
 stat = ImageStat.Stat(image)
 
 return {
 'avg_brightness': stat.mean[0] if len(stat.mean) > 0 else 0,
 'color_variance': stat.stddev[0] if len(stat.stddev) > 0 else 0,
 'dominant_color': 'rgb' # Simplified
 }

products.add_computed_column(
 visual_features=extract_visual_features(products.product_image)
)

# 5. Create balanced training set
# Under-sample majority classes or over-sample minority classes
category_counts = products.select(
 products.category_id,
 count=pxt.functions.count()
).group_by(products.category_id).collect()

min_count = min(item['count'] for item in category_counts)
print(f"Smallest class has {min_count} samples")

# Sample balanced dataset
@pxt.udf
def sample_balanced(category_id: int, row_number: int, samples_per_class: int = min_count) -> bool:
 """Select balanced samples from each class"""
 import random
 random.seed(category_id + row_number)
 return random.random() 0.88
).order_by(
 experiments.results['val_accuracy'], asc=False
).limit(1).collect()[0]

# Get exact dataset used
original_snapshot = pxt.get_snapshot(best_experiment['dataset_snapshot'])
print(f"Best model used {original_snapshot.count()} training samples")
print(f"Excel source: {best_experiment['excel_source']}")

# Reproduce training
reproduced_dataset = original_snapshot.to_pytorch_dataset()
 
```

 
## Best Practices for Excel to PyTorch Conversion

 
### Data Quality Checklist

 

 - ✅ **Validate file paths** before import

 - ✅ **Check label consistency** (no out-of-range values)

 - ✅ **Handle missing values** explicitly

 - ✅ **Verify image accessibility** and format compatibility

 - ✅ **Document data sources** and collection methods

 - ✅ **Version your Excel files** (or better: use Pixeltable snapshots)

 

 
### Performance Optimization Tips

 

 - **Use num_workers:** Set DataLoader num_workers=4 for parallel loading

 - **Enable pin_memory:** Use pin_memory=True for GPU training

 - **Prefetch batches:** Use prefetch_factor to prepare batches ahead

 - **Cache preprocessed images:** Let Pixeltable handle caching automatically

 - **Monitor memory:** Watch for memory leaks in custom Dataset classes

 

 
## Common Pitfalls and Solutions

 
### File Path Issues

 
```python

# Problem: Relative paths in Excel don't work after moving files
# Solution: Normalize paths during import

import os

def normalize_image_path(excel_path: str, excel_dir: str) -> str:
 """Convert relative paths to absolute paths"""
 if os.path.isabs(excel_path):
 return excel_path
 
 # Combine with Excel directory
 absolute_path = os.path.join(excel_dir, excel_path)
 
 if not os.path.exists(absolute_path):
 # Try common variations
 alternatives = [
 os.path.join(excel_dir, 'images', os.path.basename(excel_path)),
 os.path.join(excel_dir, '..', excel_path),
 ]
 for alt_path in alternatives:
 if os.path.exists(alt_path):
 return alt_path
 
 return absolute_path

# Apply during import
excel_dir = os.path.dirname(os.path.abspath('product_catalog.xlsx'))
df['normalized_path'] = df['image_path'].apply(
 lambda p: normalize_image_path(p, excel_dir)
)
 
```

 
### Character Encoding Issues

 
```python

# Handle Excel encoding issues
df = pd.read_excel(
 'international_data.xlsx',
 encoding='utf-8' # Specify encoding
)

# Clean string columns
df['product_name'] = df['product_name'].str.strip() # Remove whitespace
df['category'] = df['category'].str.lower() # Normalize case
 
```

 
## Conclusion: From Spreadsheets to Production Training

 
Converting Excel data to PyTorch training datasets doesn't have to be painful. Whether you use the traditional pandas approach for simple projects or Pixeltable's [declarative infrastructure](/blog/declarative-multimodal-incremental) for production workflows, the key is having a systematic approach that handles validation, preprocessing, and versioning properly.

 
 
For teams building serious ML systems, Pixeltable provides critical advantages: automatic validation, built-in versioning, complete lineage tracking, and seamless integration with training frameworks. This transforms ad-hoc Excel imports into reproducible, auditable data pipelines that scale from prototype to production.

 
## Learn More About ML Data Pipelines

 

 - **[Training Engineer Guide](/blog/training-engineer-model-development-pytorch-integration)** - Production training workflows

 - **[ML Engineer Dataset Management](/blog/ml-engineer-dataset-chaos-autonomous-vehicle-workflows)** - Managing complex datasets

 - **[Pixeltable vs Pandas](/blog/pixeltable-vs-pandas-multimodal-data-wrangling)** - When to use what

 - **[Your First Pixeltable Project](/blog/your-first-pixeltable-project)** - Getting started tutorial

 - **[Data Versioning Guide](/blog/pixeltable-versioning-time-travel)** - Reproducible training

 - **[PyTorch Data Loading Tutorial](https://pytorch.org/tutorials/beginner/data_loading_tutorial.html)** - Official PyTorch guide

 - **[Pixeltable on GitHub](https://github.com/pixeltable/pixeltable)** - Complete examples

 - **[Join our Discord](https://discord.gg/QPyqFYx2UN)** - Get help with data import

 

 
*Transform your spreadsheets into production-ready training data with proper validation, versioning, and lineage tracking.* 🚀