Skip to content

Qwen

Qwen is Alibaba's model line, and the part this guide recommends is the open-weight family: the models you download and serve on your own hardware. One clarification matters up front: the closed-weight flagship is API-only while a Max-class model has also been released with open weights under a custom licence, and the open-weight line is what this page describes.

A paid hosted API also exists, running on Alibaba Cloud, and for some organisations that carries a jurisdiction question. Self-hosting the open weights removes it entirely, because data governance stays in your hands rather than with any provider.

The ceiling is worth naming plainly. On ambiguous specs and sustained multi-agent work the closed frontier models still win, so the working pattern is not Qwen instead of a frontier model but Qwen as the daily driver, with a retained frontier seat for the hardest problems. No tool is best for everyone; this one is best when you want to own the model layer.

01FACTS
Cost
Freemium (Free tier + paid plans)
Ease
Model
Runs privately (self-hostable)
Checked
August 2026

Prices, plans and model versions change fast: this is a mid-2026 snapshot; check the tool's official site for the latest.

02FIT

Best for

  • The current default for new self-hosted coding work
  • Apache-2.0 open weights, free to download and run
  • Coding agents on infrastructure you control
  • Local prototyping with Ollama, production with vLLM
  • An actively maintained open-weight coding line

Less suited to

Anyone who wants coding AI without owning infrastructure has a hosted route through Alibaba's own coding subscription, and may still prefer a packaged assistant. GitHub Copilot is the broadest and cheapest entry, working in-editor with a free tier; Claude Code goes deeper when the specification is ambiguous and the work genuinely hard; and Kimi Code offers an open-source coding agent you adopt whole, with no serving stack to build first.

The standing obligation is operational. The weights cost nothing to download, but the spend shifts to GPUs and the engineering time to serve them, and that commitment runs for as long as the deployment does. Treat the decision as buying infrastructure rather than software: capacity to plan, a serving stack to operate, and a team that owns it when it breaks.

03EVIDENCE

Costs & data, in short

The open-weight line is free to download and self-host, with spend shifting to GPUs and the engineering time to serve them; a paid hosted API exists alongside; the flagship is closed-weight and API-only, though a Max-class model has also been released with open weights under a custom licence.

Qwen's hosted API runs on Alibaba Cloud, which carries a jurisdiction question for some organisations. Self-hosting the open weights removes it: the model runs on your own infrastructure, so data governance stays entirely in your hands.

Plans

Published plans and prices
FreeFree

Prices as of August 2026. Prices and plans change regularly. Check with the provider before you buy.

04IN PRACTICE

In practice

How Qwen is used, area by area.

Jobs

Operations
See all Operations tools →

For operations

For operations, Qwen changes the shape of the AI estate. Inference runs on hardware you govern, so the data-residency answer is structural rather than contractual: nothing leaves your infrastructure because nothing needs to. The hosted API on Alibaba Cloud exists for teams that want it, but the question it raises for some organisations simply does not arise when you serve the weights yourself. Cost becomes an infrastructure line, GPUs and the engineering time to run them, instead of a dependency on an external provider, and the same vLLM stack will serve DeepSeek as a second open-weight option if model diversity matters. Operations leaders who need a defensible data-governance story alongside predictable infrastructure planning gain the most.

Example tasks

  • Serve an internal model with no data leaving your estate
  • Meet data-residency requirements with in-house inference
  • Standardise one vLLM serving stack across teams
  • Plan AI spend as GPU capacity and engineering time
  • Add or swap open models without changing the stack

Limits

Without GPU capacity or serving expertise, the operational load lands squarely on your team. If the organisation is satisfied with a hosted provider's data posture, running your own inference is effort you do not need to spend.

Compares

vsPick Qwen whenPick the other when
DeepSeekFull comparison →Qwen is the current default for a new self-hosted deployment, with DeepSeek the natural second open-weight option on the same serving stackits free hosted chat and low-cost API fit your governance requirements and you would rather not run infrastructure at all
Llama (Meta)Full comparison →Qwen makes the data-residency answer structural rather than contractual, since nothing leaves your infrastructure because nothing needs to, though without GPU capacity or serving expertise the operational load lands squarely on your teamweights that do not change underneath you matter more, on tooling that stays broad
Software development
See all Software development tools →

Qwen earns its place in a development team's stack by being the open-weight line built for exactly this work: agentic coding, under an Apache-2.0 licence, kept current rather than frozen at an old release. The path from experiment to production is well worn, Ollama on a developer's laptop and vLLM once the whole team depends on it, so adopting it is an engineering project with known shapes rather than a research effort. It is built to carry the everyday volume, with the retained frontier seat picking up whatever exceeds it. Development teams that want their everyday coding model under their own control gain the most.

Example tasks

  • Run an agentic coding model on infrastructure you control
  • Generate and refactor code without source leaving the building
  • Power an internal coding assistant behind the firewall
  • Prototype coding workflows locally through Ollama
  • Serve the whole team's coding traffic through vLLM

Limits

On ambiguous specs and sustained multi-agent work the closed frontier models still do better, so keep a frontier seat for the hardest problems. If nobody on the team wants to own model infrastructure, a hosted coding tool is the simpler choice.

Compares

vsPick Qwen whenPick the other when
Claude CodeQwen gives you open weights to run and govern yourself where Claude Code is a token-metered frontier agentthe work is ambiguous and genuinely hard, the territory where the closed frontier models still win
Kimi CodeFull comparison →Qwen makes adoption an engineering project with known shapes rather than a research effort: Ollama on a laptop, vLLM once the team depends on it, and source that never leaves the buildingyou would rather adopt a finished client that speaks ACP and reuses the MCP servers you already have
Llama (Meta)Full comparison →Qwen is the open-weight line built for agentic coding specifically and kept current rather than frozen at an old release, with a frontier seat still worth keeping for ambiguous specsthe broader tooling legacy and an estate that already works are the reassurance that matters
GLM (Z.ai)Full comparison →Qwen asks somebody on the team to own model infrastructure, and if nobody wants to, a hosted coding tool is the simpler choice rather than a compromisethe running cost should be restructured without changing how anyone works
TabnineFull comparison →Qwen powers an internal coding assistant behind the firewall from weights the team holds, and is built to carry the everyday volume rather than the exceptional casethe decision belongs to enterprise procurement rather than to the developer who would run it
Devin DesktopFull comparison →Qwen puts a development team's everyday coding model under the team's own control rather than inside a vendor's editorthe editor should stay close to a normal one while quietly doing abnormal amounts of the work
CursorFull comparison →most of a day's coding does not need the strongest model available, so the everyday volume runs on weights the team serves while a frontier seat is kept for the ambiguous problems that still beat themthe editor is where the work happens and the strongest AI should be woven through it
Devin CloudFull comparison →the code never leaves the building: generation and refactoring happen on infrastructure the team runs, which is the whole answer for anyone who cannot send a repository into somebody else's environmentdelegation with review fits the workflow better than another pair of hands
OpenAI CodexFull comparison →the open-weight line keeps moving rather than being frozen at whatever was last published, so running the model yourself does not mean settling for an older onethe backlog holds more well-scoped tasks than hands and delegation would genuinely clear it
DeepSeekFull comparison →the line is tuned for agentic coding rather than for general work, so what it is strongest at is the thing a development team actually runs it foryou want strong models cheap at volume, or fully under your control

Tasks

Coding & software development
See all Coding & software development tools →

Within coding AI

Within coding AI, Qwen's distinct posture is that the model layer is the centre of the offer, with a first-party coding agent and a hosted coding subscription alongside it. You take the Apache-2.0 weights, serve them, and connect them to whatever agentic workflow you run, which suits organisations that treat coding tooling as owned infrastructure rather than a per-seat purchase. The line is tuned for agentic coding, and heavy everyday workloads run without per-token metering because the weights themselves are free; the spend is GPUs and the engineers who serve them. Qwen therefore slots in as the daily driver beside a retained frontier seat rather than as a replacement for one. Engineering organisations running high-volume agentic coding on owned infrastructure gain the most.

Example tasks

  • Anchor a coding stack you own end to end
  • Run heavy agentic workloads without per-token charges
  • Pair a self-hosted daily driver with a frontier seat
  • Wire open weights into your existing coding agents
  • Scale from solo Ollama use to team-wide vLLM serving

Limits

Developers who want completions working in minutes are better served by an in-editor product; Qwen's own coding agent ships IDE plugins and is configured with an API key rather than a deployment. And the hardest, most ambiguous engineering problems still belong to the closed frontier models.

Compares

vsPick Qwen whenPick the other when
Kimi CodeFull comparison →both vendors ship a first-party open-source coding agent, and Qwen additionally publishes the open weights beneath ityou want a working coding agent from about $19 a month rather than infrastructure to build
GLM (Z.ai)Full comparison →Qwen runs heavy agentic workloads without per-token charges once the weights are yours to serve, and its own coding agent ships IDE plugins configured with an API keya flat Coding Plan is the better cost shape
TabnineFull comparison →Qwen anchors a coding stack you own end to end, from the weights upward rather than from a vendor's deployment options downwardthe security review needs documented retention guarantees before any assistant is permitted
Devin DesktopFull comparison →Qwen scales from solo Ollama use to team-wide vLLM serving, so the growth path is infrastructure rather than seatsthe daily driver should be an agentic editor with a local agent in it
Sourcegraph CodyFull comparison →Qwen wires open weights into the coding agents you already run, so the model changes without the tooling changinga codebase has outgrown the single-repo tools
DeepSeekFull comparison →Qwen pairs a self-hosted daily driver with a retained frontier seat, so the split is planned rather than discoveredcoding AI spend matters, or you want strong open weights you control
OpenAI CodexFull comparison →Qwen ships Apache-2.0 weights you take and serve yourself, so the bill becomes GPUs and the engineers who run them rather than a subscriptionwhat is wanted is a full agent rather than the model layer somebody else builds one from
CursorFull comparison →Qwen suits organisations that treat coding tooling as owned infrastructure rather than a per-seat purchaseyou live in an editor all day and want the strongest AI woven into it
Devin CloudFull comparison →Qwen's line is tuned for agentic coding but arrives as weights rather than as an agent, so the workflow around it stays yours to buildyou have a backlog of well-defined tasks and would rather review PRs than write them
Private, local & self-hosted
See all Private, local & self-hosted tools →

This category is Qwen's home ground

This category is Qwen's home ground. The open-weight line is Apache-2.0, free to download and run, and actively maintained, which matters because a self-hosted deployment is a long-term commitment and an unmaintained line ages badly. That freshness is the substance behind its position as the current default for new self-hosted coding work. The tooling path is well established: Ollama on a laptop to prove the idea, vLLM when it becomes infrastructure, and DeepSeek available as the second open-weight option on the same stack, so choosing Qwen never locks you into a single model. Teams standardising on self-hosted open weights, and wanting the line most likely to stay current, gain the most.

Example tasks

  • Self-host current open weights for coding work
  • Keep prompts, code and outputs entirely on your hardware
  • Prove the idea on a laptop before buying GPUs
  • Serve concurrent users at production throughput with vLLM
  • Keep DeepSeek as a second option on the same stack

Limits

The closed-weight flagship is API-only, so that part of the range cannot be self-hosted, though a Max-class model has also been released with open weights under a custom licence. Production serving is also real infrastructure work; at personal scale, stay on Ollama and defer the rest.

Compares

vsPick Qwen whenPick the other when
LlamaFull comparison →Qwen is the actively maintained line, while Meta has shipped no major new open-weight family since April 2025an existing deployment and its wide tooling support outweigh having the freshest weights
GLM (Z.ai)Full comparison →Qwen's closed-weight flagship is API-only, though a Max-class model has also been released with open weights under a custom licence, and production serving is real infrastructure work rather than a downloadits earlier-generation published weights are the better fit for the self-hosted stack
TabnineFull comparison →Qwen leaves the model choice open, keeping DeepSeek available as a second open-weight option on the same stack so the deployment never locks you to one vendor's modelsthe requirement is to serve regulated teams that cloud assistants exclude
DeepSeekFull comparison →Qwen serves concurrent users at production throughput with vLLM, which is the step from one person's deployment to a team'syou want top-tier open weights on your own hardware, not the hosted app
06FAQ

Common questions

Is Qwen free?

There's a free tier to start; paid plans add capacity and features.

Where does Qwen fit best?

Qwen fits best in Software development and Operations; see its practice notes for how.

Before sharing confidential or personal data, check this tool's data-governance and training policies. They differ between providers and can change.

Last checked: August 2026

Keep reading