Skip to content
DigitalTatva
How-To Guides

How to Run an LLM Locally With Ollama

Run an LLM locally with Ollama in minutes: install it on Mac, Windows or Linux, pick a model that fits your RAM, fix the 4,096 token context default and use the local API.

How to Run an LLM Locally With Ollama
On this page
  1. Key takeaways
  2. Why run an LLM locally?
  3. What you need before you start
  4. How to install Ollama and run your first model
  5. Use Ollama from your own apps with the local API
  6. Common mistakes when running LLMs locally
  7. Local LLM vs cloud API: is it actually cheaper?
  8. Advanced tips for local LLMs
  9. Frequently asked questions
  10. Next steps

To run an LLM locally with Ollama, install Ollama from ollama.com, open a terminal and type ollama run gemma4 (or another model name). Ollama downloads the model once, then you chat with it on your own computer, offline and free, with no prompts sent to the cloud. A small model such as Llama 3.2 3B is a 2.0GB download, so even an ordinary laptop can run one; bigger models need much more memory.

This guide walks you through choosing hardware and a model, installing Ollama on Mac, Windows or Linux, chatting in the terminal, using the local API from your own code, and avoiding the mistakes that make local models feel slow or forgetful.

Key takeaways

  • Ollama is free and open source for local use. On macOS and Linux install it with one curl command; on Windows use the PowerShell installer or the download page.
  • Run a model with ollama run MODEL. Verified download sizes range from 523MB (qwen3:0.6b) to 2.0GB (llama3.2:3b), 14GB (gpt-oss:20b) and 142GB (qwen3:235b).
  • Ollama’s default context window is 4,096 tokens. Raise it with OLLAMA_CONTEXT_LENGTH if the model forgets long documents.
  • Ollama serves a local API at localhost port 11434, including an OpenAI-compatible endpoint, so many existing apps work with a one-line change.
  • Local AI wins on privacy and offline use. On cost alone, cheap cloud models such as GPT-6 Luna are often only a few dollars a month.
Bar chart of Ollama model download sizes from qwen3 0.6b at 523MB to gpt-oss 120b at 65GB
Ollama model download sizes

Why run an LLM locally?

A large language model (LLM) is the engine behind chatbots like ChatGPT and Claude. Running one locally means the model file sits on your disk and the computing happens on your own CPU or GPU. Ollama is a free tool that handles the hard parts: downloading models, loading them into memory and giving you a chat window and an API.

There are four good reasons to do it:

  • Privacy. Ollama’s FAQ states: “Ollama runs locally. We don’t see your prompts or data when you run locally.” That matters for client files, contracts, health or financial data, and anything covered by your company’s AI usage policy.
  • Offline use. Once a model is downloaded, it works on a plane or a weak connection.
  • No per-token bills. You pay only for hardware and electricity, which helps with high-volume, repetitive jobs.
  • Control. You choose the exact model and version, set a custom system prompt, and nothing changes under you overnight.

The trade-off is quality and speed. Small open models are capable for summaries, drafts, classification and coding help, but they do not match frontier cloud models such as Claude Opus 5.5 or GPT-6 Astra on hard reasoning. If you mainly want free access to a strong assistant, our list of free ChatGPT Plus alternatives may suit you better.

What you need before you start

Hardware

The key rule: the model has to fit in memory with room to spare. A model needs at least as much free RAM (or GPU memory) as its download size, plus extra for the conversation and your other apps. Ollama’s own guidance for gpt-oss:20b, a 14GB download, is that it can “run on systems with as little as 16GB memory”.

Ollama uses your GPU when it can. Its GPU documentation lists support for NVIDIA cards with compute capability 5.0 or newer (driver 550 or later), AMD Radeon cards through ROCm v7 (broadest on Linux), Apple Silicon Macs through Metal, and extra Intel and AMD support on Windows and Linux through Vulkan. Without a supported GPU, Ollama runs on the CPU, which works but is slower.

Choosing your first model

These sizes come from the official Ollama library pages. Context is the maximum the model supports, which is different from Ollama’s default setting explained later.

Model tag Download size Max context Good for
qwen3:0.6b 523MB 40K Testing on very old or low-memory machines
llama3.2:1b 1.3GB 128K Fast simple tasks, summaries
llama3.2:3b 2.0GB 128K Best first model for most laptops
qwen3:8b 5.2GB 40K Better reasoning and writing on 16GB machines
gemma4:12b About 7.7 to 8GB 256K Text and image input, long documents
qwen3:14b 9.3GB 40K Stronger general model for capable desktops
gpt-oss:20b 14GB 128K OpenAI’s open-weight reasoning model, 16GB+ memory
gemma4:31b About 19 to 20GB 256K High quality on workstations with lots of memory

The library also lists coding models such as qwen2.5-coder and qwen3-coder, DeepSeek’s deepseek-r1 reasoning models, Microsoft’s phi4 and Mistral’s mistral 7B. Very large models such as qwen3:235b (142GB) and gpt-oss:120b (65GB, designed to “fit on a single 80GB GPU”) are for servers, not laptops.

Tip: “B” in a model name means billions of parameters, a rough measure of size. More parameters usually means better answers but more memory and slower replies. Start small, then step up only if the answers are not good enough.

How to install Ollama and run your first model

  1. Install Ollama. On macOS or Linux, open Terminal and run the official install script shown below. On Windows, open PowerShell and run the Windows command, or download the installer from the Ollama download page. On macOS and Windows you can also use the desktop app, which updates itself automatically.
  2. Check it works. Type ollama in a new terminal window. You should see a list of commands. If you get “command not found”, close and reopen the terminal.
  3. Download and run a model. Type ollama run llama3.2. The first run downloads the model (about 2.0GB), then shows a prompt where you can type. Replies stream back word by word.
  4. Chat. Ask a question, paste text to summarize, or request a draft. Type /bye to leave the chat. The model stays loaded in memory for 5 minutes by default, so the next run starts faster.
  5. Try a second model. Run ollama pull qwen3:8b to download without chatting, then ollama run qwen3:8b. Compare answers to the same prompt to see which suits your work.
  6. Manage your models. Use ollama list to see what is installed, ollama ps to see what is loaded (the Processor column shows whether it is on CPU or GPU), ollama stop MODEL to unload it and ollama rm MODEL to delete it and free disk space.

The install commands from Ollama’s official GitHub README:

macOS and Linux:
curl -fsSL https://ollama.com/install.sh | sh

Windows (PowerShell):
irm https://ollama.com/install.ps1 | iex

Docker:
docker pull ollama/ollama

Models are stored in ~/.ollama/models on macOS, /usr/share/ollama/.ollama/models on Linux and C:\Users\%username%\.ollama\models on Windows. Check your free disk space before pulling large models.

Use Ollama from your own apps with the local API

While Ollama is running, it serves an API on your machine. By default it binds to 127.0.0.1 port 11434, which means only your own computer can reach it. Here is the chat request from Ollama’s README:

curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{"role": "user", "content": "Why is the sky blue?"}],
  "stream": false
}'

OpenAI-compatible endpoint

Ollama also offers an OpenAI-compatible API at http://localhost:11434/v1/, with chat completions, completions, models, embeddings and the Responses API. Code written for OpenAI’s Python library works by changing the base URL. The API key is required by the library but ignored by Ollama:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama")

reply = client.chat.completions.create(
    model="gpt-oss:20b",
    messages=[{"role": "user", "content": "Write a 3 line product summary."}],
)
print(reply.choices[0].message.content)

This is how you connect local models to coding tools, note apps, automation tools and chat interfaces. It is also a cheap way to prototype before you pay for a cloud API. If you later move to the cloud, our guide on how to estimate AI API costs before you build helps you budget, and our guide to adding an AI chatbot to your website covers the hosted route. For Python projects there is also an official library: pip install ollama.

Create a custom model with a Modelfile

A Modelfile lets you save a base model plus your own settings and system prompt under a new name. This is ideal for a reusable “brand voice writer” or “invoice classifier”:

FROM llama3.2
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
SYSTEM You are a concise marketing copywriter for an Indian D2C brand. Use simple English.

Save it as a file named Modelfile, then run ollama create brand-writer -f ./Modelfile and ollama run brand-writer. Good system prompts matter even more for small models; our prompt engineering guide has techniques that transfer directly.

Common mistakes when running LLMs locally

1. Ignoring the 4,096 token default context

Ollama’s default context window is 4,096 tokens, even if the model supports 128K or more. Paste a long report and the model quietly loses the beginning. Raise it with an environment variable, for example OLLAMA_CONTEXT_LENGTH=8192 ollama serve, or set num_ctx in a Modelfile. Larger contexts use more memory, so increase in steps. Our explainers on context windows and what tokens are cover the basics.

2. Picking a model that does not fit in memory

If replies crawl at a word every few seconds, the model is probably too big for your RAM or GPU and is spilling over. Run ollama ps and check the Processor column. Drop to a smaller tag such as llama3.2:3b or qwen3:8b.

3. Exposing the API to the network by accident

Ollama listens only on your own machine by default. Changing OLLAMA_HOST to listen on all network interfaces makes your model reachable by other devices, with no password. Only do it on a trusted network or behind a firewall or reverse proxy with authentication.

4. Expecting frontier-model quality

Local models are good assistants, not replacements for the best cloud models on complex research, tricky code or high-stakes writing. Always check facts, because small models invent details more often.

5. Forgetting that cloud models are not local

The Ollama library includes cloud variants (tags such as gpt-oss:20b-cloud and gemma4:cloud). These run on Ollama’s servers, not your computer. Ollama says it processes those prompts to provide the service but does not store or log the content. If privacy is your reason for going local, stick to non-cloud tags.

Local LLM vs cloud API: is it actually cheaper?

Local inference has no per-token fee, but it is not automatically the cheapest option. Here is a worked example using official API prices from our facts.

Say a small business processes 10 million input tokens and 2 million output tokens a month (thousands of support emails summarized and tagged):

  • GPT-6 Luna at $0.10 per 1M input and $0.50 per 1M output: 10 x $0.10 = $1.00, plus 2 x $0.50 = $1.00, total about $2 a month.
  • Claude Haiku 4.5 at $1 input and $5 output: 10 x $1 = $10, plus 2 x $5 = $10, total about $20 a month.
  • Ollama on hardware you already own: $0 in API fees, plus electricity and slower processing.

So for light workloads, the main reason to go local is privacy, offline access or control, not savings. Local starts to pay off when volumes are very high, data cannot leave your machine, or you already own a capable GPU. For a wider view of model costs, see how to pick the cheapest AI model for each task and our OpenAI API pricing breakdown.

Option Price Where it runs Best for
Ollama local Free Your computer Private and offline work
Ollama Cloud Free $0, starter credits Ollama servers Trying bigger models, 1 concurrent request
Ollama Pro $20/month or $200/year Ollama servers $60 usage credits a month, 3 concurrent requests
Ollama Max $100/month Ollama servers $300 usage credits a month, 10 concurrent requests

Advanced tips for local LLMs

  • Use a coding model in your editor. Point an editor extension at the OpenAI-compatible endpoint and use qwen2.5-coder or qwen3-coder for private code help. For full agentic coding, cloud tools are still stronger; compare them in our list of best AI coding assistants.
  • Build private document search. Pair an embedding model such as nomic-embed-text or mxbai-embed-large with a chat model to answer questions over your own files without uploading them.
  • Keep models warm. Models unload after 5 minutes idle. Set OLLAMA_KEEP_ALIVE or the keep_alive parameter if you call the API in bursts and want to avoid reload delays.
  • Use vision models. Gemma 4 accepts images as well as text, useful for reading screenshots or product photos locally.
  • Update regularly. The Mac and Windows apps update automatically. On Linux, re-run the install script.
  • Prefer a graphical app? Desktop apps such as LM Studio also run open models locally with a point-and-click interface, but Ollama’s command line and API make it the easier choice for automation.

Frequently asked questions

Is Ollama free to use?

Yes. Running models locally with Ollama is free and open source, and you only pay for your own hardware and electricity. Ollama also sells optional cloud plans for running larger models on its servers: a Free cloud tier with starter credits, Pro at $20 a month or $200 a year with $60 of monthly usage credits, and Max at $100 a month.

How much RAM do I need to run an LLM locally?

You need more free memory than the model’s download size. Small models like llama3.2:3b (2.0GB) run on most modern laptops, qwen3:8b (5.2GB) is a sensible choice on a 16GB machine, and Ollama says gpt-oss:20b (14GB) can run with as little as 16GB of memory. A supported GPU or Apple Silicon Mac makes replies much faster.

Does Ollama send my data to the internet?

Not for local models. Ollama’s FAQ says “Ollama runs locally. We don’t see your prompts or data when you run locally.” The internet is used to download models and updates. Cloud model tags are different: they run on Ollama’s servers, which process prompts to provide the service but, according to Ollama, do not store or log that content.

What is the best model to run locally with Ollama?

For most laptops, start with llama3.2:3b because it is a small 2.0GB download and fast. On a 16GB machine, try qwen3:8b or gemma4:12b for better writing and reasoning. For coding, try qwen2.5-coder. For OpenAI-style reasoning on stronger hardware, gpt-oss:20b is a good step up. Test the same prompts across two or three models.

Can I use Ollama with apps built for the OpenAI API?

Usually yes. Ollama exposes an OpenAI-compatible API at http://localhost:11434/v1/ that supports chat completions, completions, embeddings, model listing and the Responses API. In most apps you change the base URL, enter any placeholder API key such as “ollama”, and select a local model name like gpt-oss:20b.

Next steps

Install Ollama, run llama3.2:3b, and spend an hour testing it on your real tasks: summarizing documents, drafting emails or tagging data. If quality is good enough, increase the context window, save a custom Modelfile for your most common job, and connect it to your tools through the local API. If it is not, keep sensitive work local and use a cloud model for the hard problems; our DeepSeek vs ChatGPT comparison and roundup of the best free AI tools are good next reads. Ollama’s cloud plans and model library change often, so confirm current details on the official Ollama site before you rely on them.

Pricing and features are checked at the time of writing and can change. Some links may be affiliate links, which never affect our verdicts.

Written by

Ketan Parmar

Ketan Parmar has spent more than 15 years in digital marketing, helping brands grow through SEO, Google Ads, Meta Ads, content strategy and social media. Today he focuses on AI search visibility: how businesses get found and recommended in ChatGPT, Gemini, Perplexity and Google's AI answers.

Get the weekly AI tools brief

New tools, price changes and money-saving deals. One email a week, no spam.