---
title: "Build Production Document RAG Pipelines with Pixeltable"
date: "2025-12-09"
author: "Pixeltable Team"
tags:
  - Document Processing
  - PDF
  - RAG
  - Retrieval
  - Pixeltable
  - LLM
  - Embeddings
  - OCR
  - Text Extraction
  - Knowledge Base
description: "Learn how to build end-to-end document processing pipelines that extract, chunk, embed, and search PDFs and documents using Pixeltable's declarative approach, with no external orchestration needed."
url: "https://pixeltable.com/blog/document-pdf-processing-rag-pixeltable"
---

# Build Production Document RAG Pipelines with Pixeltable

## The Document Processing Challenge

 
Building a document RAG (Retrieval-Augmented Generation) system sounds simple: extract text from PDFs, chunk it, embed it, and search. But production reality is messier:

 

 - PDFs have tables, images, headers, and footers that need special handling

 - Documents get updated and need re-processing

 - You need to track which chunks came from which documents

 - Embeddings need to stay in sync when documents change

 - The pipeline needs to scale to thousands of documents

 

 
Pixeltable solves all of this with a declarative approach that handles extraction, chunking, embedding, and search in a single unified system.

 
## Basic Document Pipeline

 
 
Let's start with a simple but complete document processing pipeline:

 
```python

import pixeltable as pxt
from pixeltable.functions.document import document_splitter
from pixeltable.functions.huggingface import sentence_transformer

# Create a document library
pxt.drop_dir('docs', force=True)
pxt.create_dir('docs')

documents = pxt.create_table('docs.library', {
 'document': pxt.Document, # Native document type
 'filename': pxt.String,
 'uploaded_at': pxt.Timestamp
})

# Step 1: Create chunks using document_splitter
chunks_view = pxt.create_view(
 'docs.chunks',
 documents,
 iterator=document_splitter(
 documents.document,
 separators='token_limit',
 limit=512,
 overlap=50,
 )
)

# Step 2: Create a vector index (embeddings computed on demand)
embed_fn = sentence_transformer.using(model_id='sentence-transformers/all-MiniLM-L6-v2')
chunks_view.add_embedding_index('chunk_idx', string_col='text', string_embed=embed_fn)

# Insert documents - everything else happens automatically!
documents.insert([
 {'document': 'report_q1.pdf', 'filename': 'Q1 Report', 'uploaded_at': '2025-01-15'},
 {'document': 'handbook.pdf', 'filename': 'Employee Handbook', 'uploaded_at': '2025-01-10'},
])

# Search your documents
def search_docs(query: str, top_k: int = 5):
 sim = chunks_view.text.similarity(string=query)
 return chunks_view.order_by(sim, asc=False).limit(top_k).select(
 chunks_view.filename,
 chunks_view.text,
 chunks_view.pos # Chunk position in document
 ).collect()

# Find relevant passages
results = search_docs("What is the vacation policy?")
 
```

 
## Advanced: Handling Complex PDFs

 
 
Real PDFs have tables, images, and complex layouts. Here's how to handle them:

 
```python

import pixeltable as pxt
from pixeltable.functions.huggingface import sentence_transformer
from pixeltable.functions import openai

# Create table for complex documents
docs = pxt.create_table('docs.complex', {
 'document': pxt.Document,
 'doc_type': pxt.String, # 'report', 'contract', 'manual', etc.
})

# Extract text with different strategies based on document type
docs.add_computed_column(
 raw_text=docs.document.extract_text()
)

# For documents with tables, use structured extraction
docs.add_computed_column(
 tables=docs.document.extract_tables() # Returns list of table data
)

# For documents with images, extract them separately
docs.add_computed_column(
 images=docs.document.extract_images() # Returns list of images
)

# Use LLM to clean and structure messy OCR output
docs.add_computed_column(
 cleaned_text=openai.chat_completions(
 model='gpt-4o-mini',
 messages=[{
 'role': 'system',
 'content': 'Clean and structure the following extracted document text. Fix OCR errors, remove headers/footers, and format properly.'
 }, {
 'role': 'user', 
 'content': docs.raw_text
 }]
 )['choices'][0]['message']['content']
)
 
```

 
## Smart Chunking Strategies

 
 
Different documents need different chunking strategies:

 
```python

import pixeltable as pxt

docs = pxt.create_table('docs.smart_chunks', {
 'document': pxt.Document,
 'doc_type': pxt.String
})

import pixeltable as pxt
from pixeltable.functions.document import document_splitter

docs = pxt.create_table('docs.strategies', {
 'document': pxt.Document,
 'doc_type': pxt.String
})

# Strategy 1: Fixed-size chunks with overlap (good for general text)
fixed_chunks = pxt.create_view(
 'docs.fixed_chunks',
 docs.where(docs.doc_type == 'general'),
 iterator=document_splitter(
 docs.document,
 separators='token_limit',
 limit=500,
 overlap=50,
 )
)

# Strategy 2: Paragraph-based chunks (good for reports)
paragraph_chunks = pxt.create_view(
 'docs.paragraph_chunks', 
 docs.where(docs.doc_type == 'report'),
 iterator=document_splitter(
 docs.document,
 separators='paragraph',
 limit=1000,
 )
)

# Strategy 3: Sentence-based chunks (good for Q&A)
sentence_chunks = pxt.create_view(
 'docs.sentence_chunks',
 docs.where(docs.doc_type == 'faq'),
 iterator=document_splitter(
 docs.document,
 separators='sentence',
 limit=200,
 )
)
 
```

 
## Enriching Chunks with Metadata

 
 
Add useful metadata to each chunk for better retrieval:

 
```python

import pixeltable as pxt
from pixeltable.functions import openai

docs = pxt.create_table('docs.enriched', {
 'document': pxt.Document,
 'filename': pxt.String,
 'category': pxt.String
})

import pixeltable as pxt
from pixeltable.functions.document import document_splitter
from pixeltable.functions import openai

docs = pxt.create_table('docs.enriched', {
 'document': pxt.Document,
 'filename': pxt.String,
 'category': pxt.String
})

# Create chunks
chunks = pxt.create_view(
 'docs.enriched_chunks',
 docs,
 iterator=document_splitter(
 docs.document,
 separators='token_limit',
 limit=500,
 overlap=50,
 )
)

# Add the parent document info to each chunk
# (automatically inherited from parent table!)

# Generate a summary for each chunk
chunks.add_computed_column(
 summary=openai.chat_completions(
 model='gpt-4o-mini',
 messages=[{
 'role': 'user',
 'content': pxt.functions.string.format(
 'Summarize this text in 1-2 sentences: {text}',
 text=chunks.text
 )
 }]
 )['choices'][0]['message']['content']
)

# Extract key entities from each chunk
chunks.add_computed_column(
 entities=openai.chat_completions(
 model='gpt-4o-mini',
 messages=[{
 'role': 'user',
 'content': pxt.functions.string.format(
 'Extract key entities (people, organizations, dates, numbers) from this text as JSON: {text}',
 text=chunks.text
 )
 }],
 response_format={'type': 'json_object'}
 )['choices'][0]['message']['content']
)

# Generate embeddings for both text and summary
from pixeltable.functions.huggingface import sentence_transformer

chunks.add_computed_column(
 text_embedding=sentence_transformer.embed(chunks.text)
)
chunks.add_computed_column(
 summary_embedding=sentence_transformer.embed(chunks.summary)
)

# Index both for different search strategies
chunks.add_embedding_index('text_idx', column=chunks.text_embedding)
chunks.add_embedding_index('summary_idx', column=chunks.summary_embedding)
 
```

 
## Complete RAG Pipeline with LLM

 
 
Here's a full RAG system that retrieves relevant chunks and generates answers:

 
```python

import pixeltable as pxt
from pixeltable.functions import openai
from pixeltable.functions.huggingface import sentence_transformer

# Setup document library (from previous examples)
# ... docs table and chunks view with embeddings ...

# Create a queries table to track all questions
queries = pxt.create_table('docs.queries', {
 'question': pxt.String,
 'asked_at': pxt.Timestamp
})

# Step 1: Embed the question
queries.add_computed_column(
 question_embedding=sentence_transformer.embed(queries.question)
)

# Step 2: Retrieve top-k relevant chunks
# This uses a retrieval UDF for semantic search
@pxt.udf
def retrieve_context(question_embedding: pxt.Array, top_k: int = 5) -> str:
 """Retrieve relevant document chunks for a question."""
 results = chunks.order_by(
 chunks.embedding.similarity(question_embedding, metric='cosine'),
 asc=False
 ).limit(top_k).select(chunks.text, chunks.filename).collect()
 
 context_parts = []
 for r in results:
 context_parts.append(f"[Source: {r['filename']}]\n{r['text']}")
 
 return "\n\n---\n\n".join(context_parts)

queries.add_computed_column(
 context=retrieve_context(queries.question_embedding, top_k=5)
)

# Step 3: Generate answer using retrieved context
queries.add_computed_column(
 answer=openai.chat_completions(
 model='gpt-4o',
 messages=[{
 'role': 'system',
 'content': '''You are a helpful assistant that answers questions based on the provided context.
 Only use information from the context. If the answer isn't in the context, say so.
 Always cite your sources by mentioning the document name.'''
 }, {
 'role': 'user',
 'content': pxt.functions.string.format(
 'Context:\n{context}\n\nQuestion: {question}',
 context=queries.context,
 question=queries.question
 )
 }]
 )['choices'][0]['message']['content']
)

# Ask questions - retrieval and generation happen automatically!
queries.insert([
 {'question': 'What is the company vacation policy?', 'asked_at': '2025-01-15 10:00:00'},
 {'question': 'How do I submit an expense report?', 'asked_at': '2025-01-15 10:05:00'},
])

# View questions with their answers
queries.select(queries.question, queries.answer).collect()
 
```

 
## Handling Document Updates

 
 
When documents change, Pixeltable's [incremental processing](/blog/incremental-embedding-indexes) ensures only affected chunks are recomputed:

 
```python

# Update a document - only its chunks get re-embedded!
documents.update(
 where=documents.filename == 'handbook.pdf',
 values={'document': 'handbook_v2.pdf'} # New version
)

# Pixeltable automatically:
# 1. Re-extracts text from the updated document
# 2. Re-generates chunks
# 3. Re-computes embeddings ONLY for changed chunks
# 4. Updates the vector index
# 5. Maintains lineage so you can track versions

# Delete a document - chunks and embeddings are cleaned up
documents.delete(where=documents.filename == 'old_report.pdf')
 
```

 
## Scaling to Large Document Libraries

 
 
Tips for handling thousands of documents:

 
```python

import pixeltable as pxt

# 1. Use batch inserts for efficiency
documents.insert([
 {'document': f'doc_{i}.pdf', 'filename': f'Document {i}'}
 for i in range(1000)
])

# 2. Partition by document type or date for faster queries
docs_2025 = documents.where(documents.uploaded_at >= '2025-01-01')

# 3. Use appropriate embedding models for your scale
# - all-MiniLM-L6-v2: Fast, good for <10K docs
# - all-mpnet-base-v2: Better quality, slower
# - OpenAI ada-002: Best quality, API costs

# 4. Monitor your index size
index_info = chunks_view.describe() # Shows index statistics
 
```

 
## Why Pixeltable vs Traditional RAG Stacks

 
 
| Task | Traditional Stack | Pixeltable |
| --- | --- | --- |
| Text extraction | PyPDF2, pdfplumber, unstructured | Built-in extract_text() |
| Chunking | LangChain, custom code | Built-in chunks() iterator |
| Embeddings | Separate embedding service | Computed column with caching |
| Vector store | Pinecone, Weaviate, pgvector | Built-in embedding index |
| Orchestration | Airflow, Prefect, custom | Automatic with dependencies |
| Updates | Manual re-indexing | Automatic incremental updates |
| Lineage | External tracking | Built-in versioning |

 
## Conclusion

 
 
Building production document RAG systems doesn't have to be complex. With Pixeltable's declarative approach:

 

 - Define your pipeline as computed columns

 - Let Pixeltable handle extraction, chunking, and embedding

 - Get automatic incremental updates when documents change

 - Scale from prototypes to production without changing code

 

 
The best part? Your entire document pipeline (from raw PDFs to searchable embeddings) lives in one place, fully versioned and queryable.

 
The class-based `TableModel` receipt for the same job is [DocuVision](/blog/docuvision-pdf-chart-qa).

 
## Resources

 

 - [Pixeltable Documentation](https://docs.pixeltable.com)

 - [Production RAG: A Data-Centric Approach](/blog/production-rag-data-centric)

 - [Embedding Management Guide](/blog/embedding-management-guide)

 - [Incremental Embedding Indexes](/blog/incremental-embedding-indexes)

 - [Pixeltable on GitHub](https://github.com/pixeltable/pixeltable)