One of Japan's largest directories x find the right AI in as little as a minute

▶︎ For those who want to list their service

  1. AI BEST SEARCH
  2. Ollama
Ollama

Ollama
: How to use it, features, and the business problems it solves

Add bookmark

What is Ollama?

An open-model runtime from US-based Ollama Inc. that lets you download and run open models such as Gemma, Qwen, DeepSeek and gpt-oss on your own computer (Mac, Windows, Linux). A single command, "ollama run", gets you started, and it also works as an OpenAI-compatible API server. Running locally is free and unlimited. Paid plans for running large models in the cloud are Pro at $20/month and Max at $100/month. The interface is in English.

Business problems it solves

About "Ollama"

What is Ollama?

Ollama is a tool for downloading open AI models such as Gemma, Qwen, DeepSeek and gpt-oss and running them on your own computer. It supports Mac, Windows and Linux, and once installed, a single command naming a model, such as ollama run gemma4, takes care of everything from the download to starting a chat. The core is open source (MIT License) and has about 180,000 stars on GitHub (as of October 10, 2026). The official website states more than 9 million installs a month and more than 1 billion model downloads.

When it starts, an API server runs on your computer (http://localhost:11434). Because it also supports an OpenAI-compatible API, you can point apps written for the ChatGPT API, or coding tools such as Cline and Kilo Code, at the models on your machine. While you run models locally, the text you enter is not sent to Ollama.

Running models locally is free and unlimited. There are also paid plans (Pro, Max, Team and Enterprise) for running large models in Ollama's cloud when your own computer can't handle them.

Provider
Ollama Inc. (USA; its Terms of Service are governed by California law)
Category
A runtime for running open models locally + an inference service for running models in the cloud
Supported OS
macOS (Sonoma 14 or later), Windows (10 22H2 or later), Linux
Models
245 models in the official library (232 local, 17 cloud, at the time of checking)
Japanese
The website and documentation are in English. How well it responds in Japanese depends on the model you choose
Pricing
Local use is free. Cloud: Free (with starter credits), Pro $20/month, Max $100/month, Team $500/month, Enterprise custom quote. Billed in USD
Eligibility
You must be at least 18 to use the services (per the Terms of Service)

How to use

  1. Install it Get the macOS or Windows installer from "Download" on the official website. From a terminal, you can also install it with curl -fsSL https://ollama.com/install.sh | sh on macOS and Linux, or irm https://ollama.com/install.ps1 | iex in PowerShell on Windows.
  2. Run a model Open the app, or run ollama in a terminal and follow the prompts to choose a model. If you name a model, as in ollama run gemma4, it is downloaded the first time and the chat starts right away. Models range from a few GB to hundreds of GB, so you need free disk space.
  3. Manage models Use ollama pull (download), ollama ls (list), ollama ps (running models) and ollama rm (remove). By writing a system prompt and other settings in a Modelfile, you can also create a customized model for yourself.
  4. Use it from apps and coding agents ollama launch claude, ollama launch codex and ollama launch opencode start coding agents such as Claude Code with Ollama models. From your own apps, call the API at http://localhost:11434 (including the OpenAI-compatible /v1/chat/completions).
  5. Use cloud models (optional) After signing in with ollama signin, you can call cloud models with names such as gemma4:cloud. With an API key, you can send requests directly to https://ollama.com/v1 even from environments where Ollama isn't installed.

Key features

01Running locally

  • One command — Just name a model and it is downloaded, loaded and ready to chat
  • Automatic GPU use — Runs on Apple M-series GPUs, NVIDIA GPUs and AMD Radeon GPUs (Intel Macs use the CPU only)
  • Model library — Filter models by capability, such as text, image understanding (Vision), tool calling, reasoning (Thinking) and embeddings
  • Modelfile — Create your own model by adding a system prompt and settings to a base model

02API and integrations

  • Local API server — In addition to its own API on `localhost:11434`, it supports OpenAI-compatible and Anthropic-compatible APIs (each covering a subset of features)
  • ollama launch — Starts Claude Code, Codex, OpenCode and others already connected to Ollama models
  • Desktop app connections — On macOS, you can set up Claude Desktop or the ChatGPT desktop app (in Codex mode) to use Ollama models

03Cloud

  • Cloud models — Run large models that are hard to run locally in Ollama's cloud, without downloading them
  • Data handling — Ollama officially states that cloud prompts and responses are not stored or logged and are never used for training. Servers are primarily in the United States, and requests may be routed to Europe for extra capacity
  • Disabling cloud features — You can restrict Ollama to local models only, for example with the environment variable `OLLAMA_NO_CLOUD=1`

Pricing

Running models on your computer is free, with no usage limits. You only pay when you run models in the cloud. Prices are in USD and are based on the official pricing page as of October 10, 2026.

PlanPriceMonthly creditsConcurrent requestsMain features
Free$0Starter amount (starter models only)1Local use, starter cloud models. Buying credits unlocks all models
Pro$20/month ($200/year billed annually)$603Larger models, running multiple models concurrently
Max$100/month$30010Early access to the newest models
Team (early access)$500/month$1,000 (shared across the team)10Unlimited users, centralized billing and administration, priority support
EnterpriseCustom quote——Model access controls, cost budgets per user and API key, dedicated support

How credits work: Cloud usage is calculated at each model's rate (per million tokens). Credits included in your plan are used first, and anything beyond that is drawn from purchased credits. Included credits reset on your monthly renewal date and do not roll over. Purchased credits expire one year after they are added (per the Terms of Service). On paid plans, Ollama emails you when you reach 90% of your included usage.

Example cloud model rates (per million tokens, from the official pricing page)

ModelInputOutput
gpt-oss:20b$0.07$0.30
gemma4$0.14$0.40
deepseek-v4.1-flash$0.30$1.20
mistral-large-4$0.68$2.09
kimi-k3$3.00$15.00

Some models (such as DeepSeek) have cheaper off-peak rates that apply outside 12:00–18:00 UTC on weekdays and all day on weekends. Requests beyond your concurrency limit are queued, and if the queue is full they are rejected.

Payment
Payments are processed by Stripe
Renewal
Subscriptions renew automatically unless cancelled before the renewal date
Cancellation
Cancel from your account settings or by contacting Ollama
Refunds
The Terms of Service and pricing page say nothing about refunds (at the time of checking)

System requirements

  • macOS — Sonoma (14) or later. Apple M-series Macs use both the CPU and GPU; Intel (x86) Macs run on the CPU only.
  • Windows — Windows 10 22H2 or later (Home or Pro). NVIDIA GPUs require driver 551.61 or later. AMD Radeon GPUs are supported via ROCm or Vulkan. The app itself needs at least 4GB of free space.
  • Linux — Besides the install script, there are manual installs and packages for AMD GPUs and ARM64.
  • Model size — A single model can take tens or even hundreds of GB. Larger models need more memory (GPU memory).

Japanese support

  • Interface and documentation — The official website and documentation are in English.
  • Conversations in Japanese — Whether you can use Japanese depends on the model. Choose a multilingual model such as Qwen or Gemma to chat in Japanese.
  • Pricing — Prices are in USD, with no yen pricing shown.

Pros and cons

Pros

  • Running locally is free and unlimited, and what you enter is not sent outside your machine
  • One command runs a model, so getting started with local LLMs takes little effort
  • The OpenAI-compatible API lets you point existing apps and coding tools at local models
  • Large models that won't run locally can be run in the cloud the same way

Cons

  • Running models comfortably on your machine requires a powerful GPU or plenty of memory
  • The interface and documentation are in English
  • The default context length (how much text it handles at once) is 4,096 tokens, so long documents require changing the setting
  • Paid cloud plans are billed in USD, and refund terms are not stated officially

Reputation and reviews

Ollama has about 180,000 GitHub stars, and its website states more than 9 million installs a month, making it a widely used way to run LLMs locally. On this site, the pages for Cline, Kilo Code, Vanna AI and others mention Ollama as a way to connect local models. It is a common choice when you don't want data to leave your organization or want to try models without paying API fees. On the other hand, depending on your computer's performance, large models may not run or may be slow, and the paid cloud plans exist to fill that gap.

Frequently asked questions (FAQ)

Q. Is Ollama free? A. Yes. Running models on your own computer is free, with no usage limits. You pay only when running models in Ollama's cloud, through a plan such as Pro ($20/month) or purchased credits.

Q. Is what I enter sent anywhere? A. While you use local models, your prompts and data are not sent to Ollama. When you use cloud models, they are processed to generate responses, but Ollama officially states that they are not stored or logged and are never used for training.

Q. Can I use it in Japanese? A. The interface and documentation are in English. For conversations, choose a model that supports Japanese, such as Qwen or Gemma.

Q. What kind of computer do I need? A. It depends on the size of the model. Macs with Apple M-series chips and PCs with NVIDIA or AMD GPUs can run models quickly on the GPU. Larger models need more memory, so it's best to start with a small model.

Q. Can I use it with apps built for the OpenAI API? A. Yes. Change the endpoint to http://localhost:11434/v1/ and you can call local models from the OpenAI client (an API key value is required but ignored locally). Only a subset of the OpenAI API is supported.

Q. Will I get a refund if I cancel a paid plan? A. You can cancel from your account settings, but the Terms of Service and pricing page say nothing about refunds (as of October 10, 2026).

How Ollama differs from other AI tools

ServiceTypeBest for
OllamaRuntime for running open models locally + cloud inferenceRunning models on your own machine, pointing apps at a local LLM
Hugging FaceSharing platform for models and datasets + inference APIFinding and downloading many kinds of models, publishing your own
QwenAlibaba's LLM (chat, API, open weights)Using a published model itself
ClineCoding agent that runs in VS CodeWriting code with models such as those run by Ollama

Qwen and Mistral are "models," Hugging Face is the "place" where models are distributed and shared, and Ollama is the "runtime" that runs those models on your machine. Tools such as Cline can use models running in Ollama as their backend.

The information on this page is as of October 10, 2026, and is based on Ollama's official website (home page, pricing page, model library, Terms of Service and Privacy Policy) and official documentation (quickstart, CLI, OpenAI compatibility, cloud, macOS/Windows/Linux and FAQ). Pricing and features may change, so please check the official website for the latest information.

Try Ollama

Frequently compared services