Model Params Context Licence Capabilities Sovereignty
Qwen3.6-35B-A3B ★ Latest Qwen. 35B MoE with 3B active — small footprint, agentic reasoning.
35B / 3B active (MoE)
128K
Apache 2.0
Text generationReasoning
Sovereign · AU
Qwen3.5 Feb 2026 flagship. Multimodal agents, faster + cheaper than U.S. rivals.
397B / 17B active (MoE)
128K
Apache 2.0
Text generationReasoning
Sovereign · AU
Qwen3-VL-235B-A22B Vision-language flagship. Sharper vision, deeper reasoning, broader action.
235B / 22B active (MoE)
128K
Apache 2.0
VisionText generation
Sovereign · AU
Qwen3-Next-80B-A3B Efficient hybrid. Best tokens-per-watt in the Qwen 3 line.
80B / 3B active (MoE)
128K
Apache 2.0
Text generationReasoning
Sovereign · AU
Qwen3-Coder-480B-A35B ★ Alibaba's most powerful coding model. Repo-scale agentic workflows.
480B / 35B active (MoE)
256K
Apache 2.0
CodeText generation
Sovereign · AU
Qwen3-235B-A22B ★ Qwen 3 flagship. Hybrid thinking + non-thinking modes. 36T training tokens.
235B / 22B active (MoE)
128K
Apache 2.0
Text generationReasoning
Sovereign · AU
Qwen3-32B Dense Qwen 3. Strong single-GPU deployment target.
32B dense
128K
Apache 2.0
Text generationReasoning
Sovereign · AU
Qwen3-30B-A3B Cost-efficient MoE. Runs on modest hardware.
30B / 3B active (MoE)
128K
Apache 2.0
Text generation
Sovereign · AU
Qwen2.5-VL-32B-Instruct Vision-language Qwen. Surpasses Qwen2.5-VL-72B and GPT-4o mini on benchmarks.
32B
128K
Apache 2.0
VisionText generation
Sovereign · AU
QwQ-32B Compact reasoning Qwen. Matches DeepSeek-R1 with far smaller compute.
32B dense
32K
Apache 2.0
ReasoningText generation
Sovereign · AU
Qwen2.5-VL-72B-Instruct Vision-language flagship. Long-video understanding + high-res OCR.
72B
128K
Qwen
VisionText generation
Sovereign · AU
DeepSeek-V4-Pro ★ V4 preview flagship. 1.6T parameter MoE, 1M context. Adopted by Huawei + Cambricon.
1.6T (MoE)
1M
MIT
Text generationReasoning
Sovereign · AU
DeepSeek-V4-Flash V4 low-latency variant. Most of the smarts at a fraction of the compute.
284B (MoE)
1M
MIT
Text generationCode
Sovereign · AU
DeepSeek-V3.2 Uses DeepSeek Sparse Attention. More efficient long-context inference.
671B / 37B active (MoE)
128K
MIT
Text generationReasoning
Sovereign · AU
DeepSeek-V3.1 Hybrid thinking + non-thinking modes. +40% on SWE-Bench vs V3.
671B / 37B active (MoE)
128K
MIT
Text generationReasoning
Sovereign · AU
DeepSeek-R1-0528 ★ R1 refresh. Reasoning model with visible chain-of-thought traces.
671B / 37B active (MoE)
128K
MIT
ReasoningText generation
Sovereign · AU
DeepSeek-V3-0324 V3 refresh. Workhorse chat + code model, MIT-licensed.
671B / 37B active (MoE)
128K
MIT
Text generationCode
Sovereign · AU
Gemma 3 27B ★ Largest open Gemma 3. Multimodal, 140+ languages, GQA + SigLIP vision.
27B dense
128K
Gemma
Text generationReasoning
Sovereign · AU
Gemma 3 12B Mid-tier Gemma 3. Multimodal, best-in-class 12B open model.
12B dense
128K
Gemma
Text generationVision
Sovereign · AU
Mistral Small 3.2 Latest Mistral Small. Multimodal, function-calling, edge-deployable.
24B dense
128K
Apache 2.0
Text generationVision
Sovereign · AU
Phi-4 Reasoning Small-model reasoning leader. MIT-licensed, 14B fits on one GPU.
14B dense
32K
MIT
ReasoningText generation
Sovereign · AU
Kimi K2.5 ★ Multimodal upgrade to K2 — adds 400M-param MoonViT vision encoder.
1T / 32B active (MoE)
256K
Modified MIT
Text generationVision
Sovereign · AU
Kimi K2 Thinking ★ Reasoning + agentic tool-calling. INT4 native, 200-300 sequential tool calls autonomously. Beats GPT-5 + Claude Sonnet 4.5 on HLE (44.9%) + SWE-Bench Verified (71.3%).
1T / 32B active (MoE)
256K
Modified MIT
ReasoningText generation
Sovereign · AU
Kimi-K2-Instruct-0905 K2 refresh. Doubled context to 256K, improved agentic coding.
1T / 32B active (MoE)
256K
Modified MIT
Text generationCode
Sovereign · AU
Kimi K2 Moonshot's flagship. 1T total / 32B active, trained on 15.5T tokens. Beats GPT-4o + Claude on coding at a fraction of the price.
1T / 32B active (MoE)
128K
Modified MIT
Text generationCode
Sovereign · AU
GLM-5.2 ★ 1M context window. Tops GPT-5.5 on key benchmarks.
~355B / 32B active (MoE)
1M
MIT
Text generationReasoning
Sovereign · AU
GLM-5.1 Long-horizon tasks. Coding agents can run autonomously for hours.
~355B / 32B active (MoE)
200K
MIT
Text generationReasoning
Sovereign · AU
GLM-5 From vibe coding to agentic engineering. Z.ai's first major post-IPO release.
~355B / 32B active (MoE)
128K
MIT
Text generationReasoning
Sovereign · AU
GLM-4.7 Coding specialist. Surpasses Gemini 3.0 Pro on some coding tests.
~355B / 32B active (MoE)
200K
MIT
CodeText generation
Sovereign · AU
GLM-4.6 First FP8 + Int4 integration on Cambricon chips. Native FP8 on Moore Threads GPUs.
~355B / 32B active (MoE)
200K
MIT
Text generationReasoning
Sovereign · AU
GLM-4.5V Vision-language flagship. Called the best-performing 100B-class VLM globally.
106B (MoE)
128K
MIT
VisionText generation
Sovereign · AU
GLM-4.5 ★ Z.ai's first MIT-licensed flagship. Runs on 8× NVIDIA H20.
~355B / 32B active (MoE)
128K
MIT
Text generationReasoning
Sovereign · AU
GLM-4.5-Air Lightweight GLM-4.5. Cost-efficient inference for high-throughput workloads.
~106B (MoE)
128K
MIT
Text generationCode
Sovereign · AU