170+ Models

Every Model, One CLI

Use any LLM from any provider. Switch models mid-session without losing context. All via litellm.

Anthropic
OpenAI
Google
xAI
DeepSeek
Mistral
Alibaba
Meta
Cohere
Groq
Microsoft
Cerebras
Fireworks
Together
Hugging
Perplexity
NVIDIA
Cloudflare
Amazon
OpenRouter
Local

Switch Models Mid-Session

bash
/model gpt-4o

Switch in-session

bash
/models

List all models

bash
apex --model gpt-4o "prompt"

Specify on launch

claude-sonnet-4.6
anthropic/claude-sonnet-4-6

Claude Sonnet 4.6 — 1M context, extended thinking

Best for: General coding, debugging, refactoring

claude-opus-4.7
anthropic/claude-opus-4-7

Claude Opus 4.7 — 1M context, most capable

Best for: Complex architecture, multi-file refactoring

claude-sonnet-4.5
anthropic/claude-sonnet-4-5

Claude Sonnet 4.5 — latest reasoning

Best for: Advanced coding, complex reasoning

claude-opus-4.5
anthropic/claude-opus-4-5

Claude Opus 4.5 — premium reasoning

Best for: Complex tasks, deep analysis

claude-haiku-4.5
anthropic/claude-haiku-4-5

Claude Haiku 4.5 — fast & affordable

Best for: Quick edits, simple queries, chat

claude-3.7-sonnet
anthropic/claude-3-7-sonnet-20250219

Claude 3.7 Sonnet — enhanced coding

Best for: Advanced coding, complex reasoning

claude-3.5-haiku
anthropic/claude-3-5-haiku-20241022

Claude 3.5 Haiku — fastest Claude

Best for: Quick edits, simple queries, cost-effective

Cost Comparison

Approximate pricing per 1M tokens. Prices vary by provider and may change.

ModelInput/1MOutput/1MSpeedBest For
GPT-5$1.25$10.00FastGeneral coding
GPT-5 Mini$0.25$2.00Very FastBudget coding
GPT-5 Nano$0.05$0.40Very FastUltra-budget
GPT-4o$2.50$10.00FastGeneral coding
GPT-4o Mini$0.15$0.60Very FastBudget coding
Claude Sonnet 4.6$3.00$15.00FastCode quality
Claude Opus 4.7$5.00$25.00MediumComplex tasks
Claude Haiku 4.5$1.00$5.00Very FastQuick edits
Gemini 2.5 Pro$1.25$10.00FastLong context
Gemini 2.5 Flash$0.30$2.50Very FastFast coding
DeepSeek V4 Flash$0.14$0.28FastBest value
DeepSeek V4 Pro$1.74$3.48MediumReasoning
Grok 4 Fast$0.20$0.50Very FastFast reasoning
Mistral Large 3$0.50$1.50FastBalanced coding
Qwen3 Coder Plus$1.00$5.00FastCode specialist
Command A$2.50$10.00FastEnterprise RAG
Ollama (Local)FreeFreeVariesPrivacy/offline

Performance Optimization

Use Fast Models for Edits

Switch to gpt-4o-mini or claude-3.5-haiku for quick edits and simple queries. Save expensive models for complex tasks.

Clear Context Regularly

Use /clear to free up context window. Long conversations consume tokens even for small requests.

Use Streaming Mode

Enable streaming with apex --stream for real-time output. Reduces perceived latency significantly.

Leverage Local Models

Use Ollama for privacy-sensitive code or when offline. Zero cost, zero data leaving your machine.

Monitor Costs

Use /cost to check token usage and estimated costs in real-time. Set budget alerts in your config.

Choose the Right Model

Reasoning tasks → o1/deepseek-r1. Code gen → gpt-4o/claude-sonnet. Budget → gpt-4o-mini/deepseek-chat.