Meta's Llama 3.2 goes small with 1B and 3B models.
ollama run llama3.2
策展 + Ollama Library 全量 + HF GGUF 热门 合并目录:参数/显存、许可、 Hugging Face / GGUF / Ollama / ModelScope 下载,并跳转 本地算力计算器 。HF 全库十万+——本频道做可跑清单、深链与算力账,不镜像权重。 同步于 2026-08-03T07:33:50Z。
要「全部开源」请用下列权威源检索;本站策展负责选型 + 部署路径 + 算力账。
匹配 325 · 展示 40 · 点「设为 A/B」后底部一键对比 · 本站不托管权重
Meta's Llama 3.2 goes small with 1B and 3B models.
ollama run llama3.2
本地助手与 coding 入门标准;8–12GB Q4 舒适。
ollama run llama3.1:8b
超大 MoE;机房多卡 / 云推理,不适合单消费卡。
中文本地日用首选之一;8GB 卡 Q4 可跑。
ollama run qwen2.5:7b
R1 蒸馏小模型;消费级可跑推理风格。
ollama run deepseek-r1:8b
Gemma 4 models are designed to deliver frontier-level performance at each size. They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding.
ollama run gemma4
Meta Llama 3: The most capable openly available LLM to date
ollama run llama3
The latest series of Code-Specific Qwen models, with significant improvements in code generation, code reasoning, and code fixing.
ollama run qwen2.5-coder
Qwen3 单卡旗舰向;强推理与中文。
ollama run qwen3:32b
蒸馏 32B;24GB 卡推理能力明显强于 8B。
ollama run deepseek-r1:32b
4090 24GB 本地强 coding / 中文综合甜点。
ollama run qwen2.5:32b
Qwen3 小参数;思考/非思考模式按发行版。
ollama run qwen3:8b
开源旗舰对话质量;单卡 48GB+ Q4 或双 24GB TP。
ollama run llama3.3:70b
接近商用中文质量;48GB+ 或双卡。
ollama run qwen2.5:72b
OpenAI 开放权重线;本地/云双通道。
16GB 卡甜点;中文写作与中等 coding。
ollama run qwen2.5:14b
代码专用线;本地 Agent/IDE 后端常见。
ollama run qwen3-coder
多语 Embedding 标准件;CPU/小 GPU 即可。
A large language model that can use text prompts to generate and discuss code.
ollama run codellama
Gemma is a family of lightweight, state-of-the-art open models built by Google DeepMind. Updated to version 1.1
ollama run gemma
GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture.
ollama run glm-ocr
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.
ollama run gpt-oss
Llama 4 多模态 MoE 线;激活参数可控,机房/大显存。
Llama 2 is a collection of foundation language models ranging from 7B to 70B parameters.
ollama run llama2
A series of multimodal LLMs (MLLMs) designed for vision-language understanding.
ollama run minicpm-v
A state-of-the-art 12B model with 128k context length, built by Mistral AI in collaboration with NVIDIA.
ollama run mistral-nemo
State-of-the-art large embedding model from mixedbread.ai
ollama run mxbai-embed-large
Phi-3 is a family of lightweight 3B (Mini) and 14B (Medium) state-of-the-art open models by Microsoft.
ollama run phi3
Qwen 1.5 is a series of large language models by Alibaba Cloud spanning from 0.5B to 110B parameters
ollama run qwen
Qwen2 is a new series of large language models from Alibaba group
ollama run qwen2
The most powerful vision-language model in the Qwen model family to date.
ollama run qwen3-vl
Qwen 3.5 is a family of open-source multimodal models that delivers exceptional utility and performance.
ollama run qwen3.5
Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models.
ollama run qwen3.6
The TinyLlama project is an open endeavor to train a compact 1.1B Llama model on 3 trillion tokens.
ollama run tinyllama
Gemma 3 中档;多语言与指令跟随。
ollama run gemma3:12b
Gemma 3 大档本地旗舰向。
ollama run gemma3:27b
更大 MoE;多卡/机房推理为主。
经典 7B;生态与教程最多。
ollama run mistral:7b
本地截图/文档视觉问答。
ollama run qwen2-vl
中型开源;24GB 可 Q4/Q5。
ollama run mistral-small
ollama run,不二次托管权重。Clavue CLI / imux IDE / clavue-2.1
CLI / GUI Agent 运行时、设备登录、会员额度与 OpenAI 兼容 API(api.clavue.com)。开发脚本、CI 与本地工具一条链路。
打开 Clavue ↗Clavue 平台的原生 IDE:多 Agent 分屏、Ghostty 级终端、Agent Chat、浏览器自动化与 Supervisor——不是又一个 Electron 壳。
了解 imux ↗旗舰模型 clavue-2.1(及 fast / pro / rev):适合复杂推理与长程 Agent 循环。Chat、imux、API 共用会员额度。
查看模型 ↗