The Curriculum / Reader / 05 — Model Routing
LEVEL 1 · ESSENTIALS · INDIVIDUAL TRACK

05 — Model Routing

This page compiles 3 files from the repository, verbatim, in reading order. The living version: this folder on GitHub.

level-1-essentials/individual/05-model-routing/README.md

05 — Model Routing

Which model to use when. The single biggest power-user skill after prompting.

Files in this folder

  1. model-routing-cheat-sheet.md — the one-page you print and keep near your monitor
  2. provider-strengths-matrix.md — the deeper table with reasoning

The principle

Every model has a shape. Use the right shape for the job.

Definition of done

level-1-essentials/individual/05-model-routing/model-routing-cheat-sheet.md

Model Routing Cheat Sheet

Print this. Tape it above your monitor.

Route by task

Task First choice Second choice Notes
Long-doc reasoning (contracts, deep analysis) Claude Opus Gemini 2.x Pro Claude wins at nuance; Gemini wins on context length
Coding — small changes Claude Sonnet in Cursor GPT-5 in Cursor Sonnet is the daily driver
Coding — large refactor Claude Code (agent) Cursor Composer Agent mode for multi-file
Web research with citations Perplexity Sonar/Pro Perplexity Deep Research Search-grounded by default
Real-time facts (news, prices, sports) Perplexity Grok (X data) Never trust a non-grounded model on live facts
Google Workspace tasks Gemini in Workspace Copilot 365 (Microsoft) Native integration wins
Microsoft Office tasks Copilot 365 Gemini in Workspace Native integration wins
Image generation — realistic GPT (Sora/Images), Midjourney Ideogram Midjourney for hero art
Image generation — with text in image Ideogram GPT Images Ideogram nails typography
Video generation Sora, Veo 3 Higgsfield, Runway Model-specific look; test both
Voice / TTS ElevenLabs OpenAI TTS ElevenLabs for characters
Transcription Whisper (OpenAI API) AssemblyAI, Otter Whisper is baseline; Otter for meetings
Structured extraction (JSON) GPT-5 JSON mode Claude with structured output Both strong
Quick chat / drafting ChatGPT (any) Claude Sonnet Whichever tab is open
Cheap batch inference Groq (LLaMA, Mixtral) OpenRouter with cheap models Groq for speed
Uncensored generation Grok, local models For legit adult / edgy use cases
Agent orchestration Perplexity Computer, Claude Code Custom (LangGraph) This product; Claude Code for local
Deep research (multi-source) Perplexity Deep Research Gemini Deep Research Both are strong; different styles
Long context (>200k tokens) Gemini 2.x Pro (2M), Claude (1M) GPT-5 (400k) Gemini has the longest
Reasoning-heavy math/logic o-series (OpenAI), Claude thinking DeepSeek Turn on extended thinking
Multimodal (image + text in) GPT-5, Gemini Claude Gemini strong on video input

Route by data class

Data class Consumer tools OK? API OK? Notes
Public / already-shipped ✅ any ✅ any Anything goes
Personal notes / drafts ✅ any ✅ any
Internal work docs (redacted) ⚠️ enterprise only ✅ with DPA Not free tiers
Client PII ✅ enterprise API + DPA Zero-retention required
Legal-privileged ❌ except with counsel Consider local models
Passwords / keys / SSN Never

Fallback flow

If your first-choice model is rate-limited, down, or the output is bad:

  1. Try the second-choice model with the same prompt
  2. If both fail, simplify the prompt (fewer variables, smaller scope)
  3. If it's a research question, try Perplexity for the base facts, then hand to Claude for synthesis
  4. If it's a code question, add the exact error message to the prompt and try Claude Opus

Cost mental model

Approximate cost per million tokens output (as of August 2026 — verify current pricing):

Tier Models Cost / MTok output
Frontier Claude Opus, GPT-5, Gemini 2.x Pro $30–$75
Workhorse Claude Sonnet, GPT-5 mini, Gemini Flash $3–$15
Cheap Claude Haiku, GPT nano, Gemini Nano $0.5–$2
Open-weight on Groq LLaMA, Mixtral, Qwen $0.2–$1

Rule: Start with workhorse. Escalate to frontier only when workhorse can't do it.

level-1-essentials/individual/05-model-routing/provider-strengths-matrix.md

Provider Strengths Matrix

The deep version. Read once, refer back quarterly.

OpenAI (ChatGPT + API)

Signature strengths - Widest ecosystem: Custom GPTs, marketplace, integrations - Best multimodal breadth: image gen (DALL·E / Sora Images), video (Sora), voice (advanced voice mode), Whisper - Structured output: JSON mode is the most reliable - o-series (o1, o3, o4, o5-preview) for reasoning-heavy tasks - Sora for video generation - Massive third-party integration surface

Signature weaknesses - Less nuanced than Claude for long, subtle prose - Custom GPT memory is opaque compared to Claude Projects - Rate limits on peak models can bite

Best for - Multimodal work (image + text) - Custom GPTs shared with others - Batch structured extraction (JSON mode) - Voice interactions - When you need "just works" across every modality

Anthropic Claude

Signature strengths - Best long-form reasoning and prose - Longest sustained coherence over 100k+ token inputs - Artifacts (inline runnable code / documents) - Projects (persistent knowledge + custom instructions) - Claude Code (best agentic coding tool) - Skills (reusable capability packages) - Extended thinking mode surfaces reasoning

Signature weaknesses - No native image generation (as of Aug 2026) - Smaller integration ecosystem than OpenAI - Fewer built-in tools

Best for - Long documents (contracts, reports, research) - Complex coding (via Claude Code) - Writing that needs voice preservation - Anything requiring subtle judgment - Skills-based workflows

Google Gemini

Signature strengths - Longest context (2M tokens on Pro tiers) - Native Google Workspace integration (Docs, Sheets, Gmail, Meet) - Deep Research mode - Veo 3 for video - Multimodal — best at video input - Free tier is generous

Signature weaknesses - Product surface changes often (feature churn) - Consumer UX inconsistent - Voice mode less polished than OpenAI's

Best for - Workspace-native workflows (Google users) - Very long context needs (2M tokens) - Video understanding (feed a video, ask questions) - Research via Deep Research

Perplexity

Signature strengths - Search-grounded by default — every answer cites sources - Real-time facts - Comet browser + Computer (this product) — agentic web workflows - Spaces (persistent contexts with connectors) - Deep Research - API returns citations natively

Signature weaknesses - Less "raw" LLM power — it's a search-tuned system - Long-doc reasoning is not its strength - Some models better for pure generation tasks

Best for - Any question about current events, prices, people - Research briefs with citations - Agent workflows requiring web browsing - Fact-checking

Microsoft Copilot

Signature strengths - Native in Word, Excel, Outlook, Teams, Windows - Best-in-class for Office-native workflows - SharePoint / OneDrive grounding - Enterprise controls and compliance

Signature weaknesses - Feels bolted-on outside Office - Less flexible for creative or open-ended work - Locked to Microsoft's model choices

Best for - Excel formulas, PowerPoint decks, Word docs, Outlook triage - Enterprises already on M365 - Compliance-sensitive environments

xAI Grok

Signature strengths - Real-time X (Twitter) data - More permissive on edgy topics - Aurora image gen without heavy filters

Signature weaknesses - Smaller ecosystem - Fewer integrations - Less careful reasoning than Claude/GPT

Best for - Real-time X monitoring - Uncensored image gen (with judgment) - Contrarian second opinion

Groq (inference provider)

Signature strengths - Fastest inference on the market - Open-weight models (LLaMA, Mixtral, Qwen) - Cheap - Great free tier

Signature weaknesses - Only serves open models — no Claude, GPT, Gemini - Model catalog rotates

Best for - Batch processing where speed matters - Cost-sensitive workloads - When latency is critical (real-time apps)

OpenRouter

Signature strengths - One API key → 300+ models across all providers - Easy A/B testing - Automatic fallback across providers - Great for experimentation

Signature weaknesses - Adds a small margin over provider prices - Rate limits inherited from underlying provider

Best for - Experimentation - Multi-model apps - Fallback resilience

Local models (via Ollama, LM Studio, MLX)

Signature strengths - Zero data leakage - No per-token cost - Works offline - Available: Llama 3.x, Qwen 3, Mistral, DeepSeek, Phi

Signature weaknesses - Much weaker than frontier models - Requires decent hardware (M-series Mac or good GPU) - No multimodal parity

Best for - Highly sensitive data - Offline scenarios - Learning how models work under the hood

← 04 — Prompt Library 06 — Memory Hygiene →