Built-in AI Models
Clearminutes ships with 9 local AI models for generating meeting summaries. They run entirely on your device — no internet connection, no API key, no per-summary cost. This is the default and recommended way to use Clearminutes.
Models are downloaded on demand from Clearminutes' servers and stored locally. Once downloaded, they work completely offline.
Clearminutes tests your hardware at first launch and recommends the best model your machine can run. You'll see the suggested model highlighted in the grid — just download it and you're ready to go.
How to Choose
Settings → Summary → Provider: Built-in AI — then pick a model from the grid.
The app will block models that exceed your machine's available RAM. Models are grouped by how much RAM they need:
All 9 Models
Qwen3 1.7B
Best for: Low-memory machines, quick summaries
- Size: ~1 GB download
- RAM required: 4 GB minimum
- Context window: 32,000 tokens
- Languages: Multilingual
- Alibaba's Qwen3 series — punches well above its size. A solid choice on older hardware or machines with limited RAM.
Phi-4 Mini Instruct
Best for: Low-memory machines needing a large context window
- Size: ~2 GB download
- RAM required: 4 GB minimum
- Context window: 128,000 tokens
- Languages: Multilingual
- Microsoft's Phi-4 Mini is fast and accurate for its size, with an exceptionally large 128k context window. Can process very long meetings without chunking.
Gemma 4 E2B
Best for: Most users — great balance of quality and speed
- Size: ~3 GB download
- RAM required: 6 GB minimum
- Context window: 128,000 tokens
- Languages: Multilingual
- Google's next-generation Gemma 4 in its efficient 2B variant. Excellent summary quality with a 128k context window. Frequently recommended for mid-range hardware.
Gemma 4 E4B
Best for: Higher-quality summaries on 6 GB+ machines
- Size: ~4 GB download
- RAM required: 6 GB minimum
- Context window: 128,000 tokens
- Languages: Multilingual
- The larger Gemma 4 efficient variant. Noticeably better summaries than E2B, still fits in 6 GB RAM. A strong upgrade if your machine can handle it.
Qwen3 4B
Best for: Balanced quality and speed on 6 GB machines
- Size: ~3 GB download
- RAM required: 6 GB minimum
- Context window: 32,000 tokens
- Languages: Multilingual
- Comparable quality to models twice its size. The 32k context window is sufficient for most meetings under 2 hours.
Ministral 3 8B Instruct
Best for: Strong reasoning on technical or complex meetings
- Size: ~4 GB download
- RAM required: 8 GB minimum
- Context window: 128,000 tokens
- Languages: Multilingual
- Mistral's Ministral 3 8B delivers strong reasoning and structured output with a 128k context window. A great choice for engineering or product discussions.
Llama 3 8B Instruct
Best for: Excellent instruction following, detailed summaries
- Size: ~5 GB download
- RAM required: 10 GB minimum
- Context window: 8,000 tokens
- Languages: Multilingual
- Meta's Llama 3 8B follows complex instructions reliably. Note the smaller 8k context window — very long meetings will be summarised in chunks.
Qwen3 8B
Best for: High-quality summaries on 10 GB machines
- Size: ~5 GB download
- RAM required: 10 GB minimum
- Context window: 32,000 tokens
- Languages: Multilingual
- Excellent summary quality. Like Qwen3 4B but with more parameters — noticeably better on nuanced or technical content.
Gemma 3 12B
Best for: Near-frontier quality on well-equipped machines
- Size: ~7 GB download
- RAM required: 12 GB minimum
- Context window: 128,000 tokens
- Languages: Multilingual
- Google's Gemma 3 12B delivers near-frontier quality locally. If you have the RAM, this produces the best summaries of any built-in model. Ideal for executives, sales, or any meeting where summary quality is critical.
GPU Acceleration
All built-in models benefit from GPU acceleration — summaries complete in seconds rather than minutes on capable hardware:
- macOS (Apple Silicon): Metal is enabled automatically. M-series chips with unified memory are excellent for local AI.
- Windows (NVIDIA): CUDA acceleration via the CUDA build
- Windows (AMD/Intel): Vulkan acceleration
Without GPU acceleration, models still work but take longer to generate summaries. See GPU Acceleration for setup details.
Storage Locations
Downloaded models are stored at:
Models can be deleted from within the app (Settings → Summary → [model] → Delete) or by removing the .gguf file directly.
Compared to Cloud AI
For most users, built-in AI is the right choice. Consider Cloud AI only if you need frontier-level quality and are comfortable sharing meeting content with a third party.
Related
- Cloud AI for Summaries — Claude, Groq, OpenAI, OpenRouter
- Ollama Setup — use any Ollama model locally
- GPU Acceleration — make built-in models run faster
- Generate Summaries — how summaries are triggered and customised