Choose the Right Model
Clearminutes uses Whisper for transcription. Different models trade off speed, accuracy, file size, and memory usage. This guide helps you pick the right one.
What is Real-Time Factor (RTF)?
Before comparing models, understand RTF — how fast a model transcribes relative to audio length:
- RTF = 0.5: Transcribes twice as fast as real-time (a 60-min meeting transcribes in 30 min)
- RTF = 1.0: Real-time transcription (a 60-min meeting transcribes in 60 min)
- RTF > 1.0: Slower than real-time (a 60-min meeting transcribes in longer than 60 min)
Lower RTF = faster results, which matters for important meetings.
Whisper Model Comparison
All Whisper models are trained on 680,000 hours of multilingual audio. Larger models are more accurate but slower and require more memory.
Notes:
- RTF estimates are for typical hardware (RTX 3070 for GPU, Ryzen 5 for CPU)
- Actual performance varies by device, background apps, and audio quality
- GPU acceleration (Metal/CUDA/Vulkan) is 5–10× faster than CPU
- Accuracy % reflects performance on OpenAI's Whisper test set, but all models are very accurate for most meetings
Recommended Models by Hardware
macOS
Windows
Model Details
large-v3-turbo (RECOMMENDED)
Best for: Most users. Exceptional balance of speed, accuracy, and resource use.
- Size: 800 MB (moderate download)
- RTF: 0.16 (GPU), 3.0 (CPU)
- Accuracy: 93% on test set (competitive with large-v3)
- Memory: 6 GB RAM, 4 GB VRAM (GPU)
- Download time: ~2 minutes on a 50 Mbps connection (varies with provider)
Ideal scenario:
- 45-minute meeting on Apple Silicon Mac with GPU
- Transcription starts immediately after recording stops
- Results available in ~7 minutes (45 min × 0.16 RTF)
- High accuracy for all languages
Pros:
- Fast transcription without sacrificing accuracy
- Works on most hardware (even older laptops)
- Supports 99+ languages
- Small enough to fit on mobile hardware
Cons:
- Slightly lower accuracy than full large-v3 (~2%)
- Still requires a few minutes for longer meetings
Start here if: You're unsure what to pick, or this is your first Clearminutes recording.
large-v3
Best for: Maximum accuracy when speed isn't critical.
- Size: 3.0 GB
- RTF: 0.35 (GPU), 8.0 (CPU)
- Accuracy: 95% on test set (highest)
- Memory: 10 GB RAM, 6 GB VRAM (GPU)
- Download time: ~8 minutes on a 50 Mbps connection (varies with provider)
Ideal scenario:
- High-stakes meeting (board meeting, legal deposition, customer call)
- 45-minute meeting on an M3 Mac with 32 GB RAM: RTF 0.35 means roughly 16 minutes of processing on the GPU
- Exceptional accuracy, especially for technical terms and accents
Pros:
- Highest accuracy (95%) — catches edge cases other models miss
- Better at technical jargon, uncommon words, accents
- Excellent for non-English languages
Cons:
- Slow (RTF 0.35) — a 45-minute meeting takes ~16 minutes on GPU
- Requires significant disk space (3 GB)
- High VRAM requirement (6 GB) — uses half the VRAM on typical gaming laptops
- CPU-only transcription is impractical (RTF 8.0 = 6+ hours per 45-min meeting)
Use if: You need maximum accuracy and can wait 10–30 minutes for results.
medium
Best for: Older or resource-constrained hardware (older laptops, budget desktops).
- Size: 1.5 GB
- RTF: 0.25 (GPU), 4.0 (CPU)
- Accuracy: 86% on test set
- Memory: 5 GB RAM, 2.5 GB VRAM (GPU)
- Download time: 30–45 minutes
Ideal scenario:
- Intel i5 laptop with 8 GB RAM and no GPU
- 30-minute meeting on CPU takes roughly 2 hours (RTF 4.0); use GPU acceleration if you have it
- Reliable accuracy for clear speech in quiet environments
Pros:
- Works on older hardware with limited RAM
- Fast enough on CPU without GPU (RTF 4.0)
- Small VRAM requirement (2.5 GB) — fits on integrated graphics
Cons:
- Lower accuracy (86%) — more transcription errors, especially with:
- Overlapping speech
- Heavy accents
- Technical terminology
- Background noise
- Still moderately large (1.5 GB)
Use if: Your hardware has <8 GB RAM or no GPU, or you're on a budget machine.
small
Best for: Specific cases where you have good hardware but want faster results than medium.
- Size: 460 MB
- RTF: 0.15 (GPU), 2.5 (CPU)
- Accuracy: 82% on test set
- Memory: 2 GB RAM, 1 GB VRAM (GPU)
- Download time: 10–15 minutes
Ideal scenario:
- MacBook Air with 16 GB RAM, need quick transcription
- 30-minute meeting transcribes in ~4 minutes with GPU
- Good accuracy for general meetings with minimal background noise
Pros:
- Very fast (RTF 0.15)
- Small download and disk footprint (460 MB)
- Minimal VRAM needed (1 GB)
- Good accuracy for clear, focused meetings
Cons:
- Lower accuracy than medium (82%)
- Struggles with accents and technical jargon more
- Not ideal for noisy environments
Use if: You want speed and have good hardware, but are willing to accept slightly lower accuracy.
tiny & base
Best for: Testing, demonstrations, or extreme resource constraints.
tiny:
- Size: 40 MB | RTF: 0.05 (GPU) | Accuracy: 60% | Memory: 1 GB RAM / 200 MB VRAM
base:
- Size: 140 MB | RTF: 0.08 (GPU) | Accuracy: 74% | Memory: 1 GB RAM / 500 MB VRAM
Use cases:
- Quick testing: Download in <1 minute, try Clearminutes without committing to large models
- Old hardware: Machines from 2010–2014 with <2 GB RAM
- Embedded systems: Edge devices, Raspberry Pi, etc.
Limitations:
- Accuracy below 75% — expect frequent errors
- Not recommended for important meetings
- Struggles with multiple speakers, noise, and accents
Skip these unless: You're testing or have extremely limited hardware.
Language Support
All Whisper models support 99+ languages. When you record:
- Clearminutes auto-detects the language from the first 30 seconds of audio
- Transcription continues in that language
- To manually set language: Settings → Transcription → Language
Accuracy for non-English languages:
- large-v3 / large-v3-turbo: Near-native accuracy
- medium: Good accuracy for common languages, fair for less common
- small/tiny: Acceptable for clear speech in common languages, poor for rare languages
Switching Models
Download a New Model
- Open Settings (⌘, on macOS)
- Go to Transcription section
- Click the Model dropdown
- Select a new model (e.g., "large-v3")
- Click Download
- A progress bar shows download speed and time remaining
- [SCREENSHOT: Model download progress]
- Keep Clearminutes open — don't close during download
- When complete, the model is active immediately
Switch Between Models
- Go to Settings → Transcription → Model
- Click the dropdown
- Select from downloaded models
- Click Switch or confirm
- Future recordings use the new model
In-progress recordings: If you're recording when you switch models, the current recording finishes with the old model. The new model is used for the next recording.
Delete a Model
To free up disk space:
- Go to Settings → Transcription
- Find the model you want to delete
- Click Delete (trash icon)
- Confirm the deletion
- The model is uninstalled (you can download it again anytime)
Disk space freed:
- tiny: 40 MB
- base: 140 MB
- small: 460 MB
- medium: 1.5 GB
- large-v3: 3.0 GB
- large-v3-turbo: 800 MB
Summary Models
After transcription, Clearminutes can generate a summary. For built-in AI model details and recommendations, see Built-in AI Models.
Performance Troubleshooting
Model Is Too Slow
- Switch to a smaller model: large-v3-turbo → medium → small
- Check GPU: Is it being used? See GPU Acceleration
- Close other apps: Free up RAM and GPU memory
- Check CPU temperature: If your machine is throttling due to heat, performance drops
Out of Memory Error
The model is too large for your available RAM/VRAM.
- Close other apps
- Switch to a smaller model (large-v3-turbo → medium)
- Check available RAM: Activity Monitor (macOS) or Task Manager (Windows)
- If still failing, upgrade RAM or use a smaller model permanently
Model Won't Load
- Check disk space — ensure 2× the model size is available free
- Delete other models to free space
- Restart Clearminutes
- Try downloading the model again (file may be corrupted)
See Model Download Errors for more troubleshooting.
Changing Models: Step-by-Step Walkthrough
Scenario: You've been using medium but want to try large-v3-turbo for better accuracy.
- Click Settings (⌘, or gear icon)
- Navigate to Transcription section
- Find the Model dropdown
- Click to open the list of available models
- Select large-v3-turbo from the list
- If not downloaded: click Download, wait 15–20 min
- If already downloaded: click Switch to this model
- Confirm: "Switch to large-v3-turbo?"
- Click Confirm
- The next recording uses large-v3-turbo
- (Optional) Delete medium: go back to Transcription, find medium, click Delete to free 1.5 GB