---
title: "What is local inference?"
description: "Local inference runs a model on hardware you control, instead of calling a hosted model API."
url: "https://pixeltable.com/learn/what-is-local-inference"
updated: "2026-09-29"
vertical: "Orchestration"
doc: "https://docs.pixeltable.com/overview/quick-start"
---

# What is local inference?

Local inference runs a model on hardware you control, instead of calling a hosted model API.

Updated: 2026-09-29
Part of [What is a model provider?](https://pixeltable.com/learn/what-is-a-model-provider).

## On this page

- [How it works](https://pixeltable.com/learn/what-is-local-inference#how-it-works)
- [What it is not](https://pixeltable.com/learn/what-is-local-inference#what-it-is-not)
- [Comparison](https://pixeltable.com/learn/what-is-local-inference#comparison)
- [Where Pixeltable fits](https://pixeltable.com/learn/what-is-local-inference#where-pixeltable-fits)
- [Questions](https://pixeltable.com/learn/what-is-local-inference#questions)

## How it works {#how-it-works}


- A runtime such as Ollama, llama.cpp, or vLLM serves the model.
- A computed column calls that local endpoint.
- The row still caches the output.

## What it is not {#what-it-is-not}

It is not free in electricity or GPUs, and it is not required for the open-source database engine.

## local inference: this, and the thing it is confused with {#comparison}

|  | This | Not this |
| --- | --- | --- |
| Bill | Your machine | Per token to a vendor |
| Weights | On disk | Behind an API |
| Schema | Unchanged aside from the function | A different pipeline |

## Where Pixeltable fits {#where-pixeltable-fits}

Pixeltable columns can call Ollama, llama.cpp, or vLLM the same way they call a hosted provider.

## Questions {#questions}

### How does local inference work? {#faq-1}

A runtime such as Ollama, llama.cpp, or vLLM serves the model. A computed column calls that local endpoint. The row still caches the output.

### What is local inference often confused with? {#faq-2}

It is not free in electricity or GPUs, and it is not required for the open-source database engine.

## In the blog

- [llama.cpp: High-Performance Local LLM Inference with Quantized Models in Pixeltable](https://pixeltable.com/blog/llama-cpp-local-inference-pixeltable)
- [vLLM: High-Throughput Local LLM Inference in Pixeltable](https://pixeltable.com/blog/vllm-high-throughput-local-inference-pixeltable)
- [Why the Local-Cloud Loop Matters](https://pixeltable.com/blog/why-local-cloud-loop-matters)
- [Ollama: Run Local LLMs with Llama, Qwen, and More in Pixeltable](https://pixeltable.com/blog/ollama-local-llm-pixeltable)
- [Together AI: Fast Open-Source LLM Inference with Llama 3.3 and FLUX in Pixeltable](https://pixeltable.com/blog/together-ai-llm-inference-pixeltable)
- [Pixeltable May/June 2026 Release Highlights](https://pixeltable.com/blog/pixeltable-may-june-2026-release-highlights)

## Related

- [Documentation](https://docs.pixeltable.com/overview/quick-start)
- [Ollama: Run Local LLMs with Llama, Qwen, and More in Pixeltable](https://pixeltable.com/blog/ollama-local-llm-pixeltable)
- [llama.cpp: High-Performance Local LLM Inference with Quantized Models in Pixeltable](https://pixeltable.com/blog/llama-cpp-local-inference-pixeltable)
- [vLLM: High-Throughput Local LLM Inference in Pixeltable](https://pixeltable.com/blog/vllm-high-throughput-local-inference-pixeltable)
