Provider Caching Rules
OpenAI
- Min 1,024 tokens to cache
- Cached reads: 50% off input price
- No write surcharge
- Auto-caching on all requests
Anthropic
- Min 1,024 tokens (Haiku: 2,048)
- Cached reads: 90% off input price
- Cache write: 25% surcharge
- 5-min TTL (extended on hit)
Google
- Min 32,768 tokens
- Cached reads: 75% off input price
- Storage: $1.00/hr per 1M tokens
- Min 1-hour TTL
DeepSeek
- Min 1,024 tokens
- Cached reads: 90% off input price
- No write surcharge
- Auto prefix matching
Your Setup
Tokens in your system prompt / context that stay constant
User message + dynamic context that changes each call
How often your cached prefix is reused (80%+ is typical for consistent system prompts)