Free-Tier Model Allocation &
Rate-Limit Capacity Matrix
Our intelligent ModelRouter maps conversational, visual, and memory retrieval tasks to the optimal free-tier model on Google AI Studio, maintaining zero cloud expenditure under strict rate limits.
| Model Identifier | Assigned Role | Daily Limit (RPD) | Rate Limit (RPM) | Context Window (TPM) | Inference Latency | Usage Strategy |
|---|---|---|---|---|---|---|
|
Gemma 4 26B
Primary Chat
|
High-frequency DMs, urban Hinglish banter, personality dialogue | 14,400 / day | 30 RPM | 16,000 TPM | ~750ms | $0.00 Free • Workhorse |
|
Gemma 4 31B
Reasoning
|
Complex social graph reflection & high-tier friend banter fallback | 14,400 / day | 30 RPM | 16,000 TPM | ~920ms | $0.00 Free • Secondary |
|
Gemini 3.5 Flash Lite
Vision & Video
|
Instagram Reel video QA, photo aesthetic scoring, structured JSON parsing | 500 / day | 15 RPM | 250,000 TPM | ~1,100ms | $0.00 Free • High-Context |
|
Gemini 3.1 Flash Lite
|
Multimodal reserve & structured Pydantic memory schema extraction | 500 / day | 15 RPM | 250,000 TPM | ~980ms | $0.00 Free • Hot Reserve |
|
Gemini 3.8 Flash
Flagship
|
High-stakes aesthetic evaluation & deep cultural context verification | 20 / day | 5 RPM | 250,000 TPM | ~1,450ms | $0.00 Free • Reserved |
|
Gemini Embedding 2
Vector Memory
|
3,072-dimensional embeddings for episodic semantic memory search | 1,000 / day | 100 RPM | 30,000 TPM | ~65ms | $0.00 Free • Vector Store |
Intelligent Task Allocation & Failover
Incoming requests are classified by media type and required context length before reaching model dispatch:
Conversational Inbound (Text Only)
Direct messages under 4,000 tokens dispatch immediately to Gemma 4 26B. If Gemma 4 approaches 90% quota, automatic failover routes to Gemma 4 31B.
Shared Reel / Video (Multimodal)
Video clips require massive context capacity. The request dispatches to Gemini 3.5 Flash Lite (250,000 TPM) with extracted keyframes and local OCR overlays.
Episodic Memory Recall (Vectors)
Query terms are embedded into 3,072 dimensions using Gemini Embedding 2 at 100 RPM and compared against cosine distance metrics in the vector database.