Navigation
Overview Research Publications Safety & Policy Architecture Engineering News Try Ishita (Public Access)
Home Overview Research Publications Alignment Science Rate-Limit Economics Empirical Benchmarks Newsroom & Blog Announcements & Releases Engineering Stories Free-Tier Model Matrix 14-Step Execution Pipeline Character-State Architecture Pros & Cons Tradeoffs Safety & Ethics Demo Sandbox Launch Control Center →
Multi-Model Cognitive Layer

Free-Tier Model Allocation &
Rate-Limit Capacity Matrix

Our intelligent ModelRouter maps conversational, visual, and memory retrieval tasks to the optimal free-tier model on Google AI Studio, maintaining zero cloud expenditure under strict rate limits.

6
Supported Model Endpoints
14,400
Daily Gemma 4 RPD
250K
Tokens / Min Video Context
3,072-dim
Vector Memory Embeddings
Model Identifier Assigned Role Daily Limit (RPD) Rate Limit (RPM) Context Window (TPM) Inference Latency Usage Strategy
Gemma 4 26B Primary Chat
High-frequency DMs, urban Hinglish banter, personality dialogue 14,400 / day 30 RPM 16,000 TPM ~750ms $0.00 Free • Workhorse
Gemma 4 31B Reasoning
Complex social graph reflection & high-tier friend banter fallback 14,400 / day 30 RPM 16,000 TPM ~920ms $0.00 Free • Secondary
Gemini 3.5 Flash Lite Vision & Video
Instagram Reel video QA, photo aesthetic scoring, structured JSON parsing 500 / day 15 RPM 250,000 TPM ~1,100ms $0.00 Free • High-Context
Gemini 3.1 Flash Lite
Multimodal reserve & structured Pydantic memory schema extraction 500 / day 15 RPM 250,000 TPM ~980ms $0.00 Free • Hot Reserve
Gemini 3.8 Flash Flagship
High-stakes aesthetic evaluation & deep cultural context verification 20 / day 5 RPM 250,000 TPM ~1,450ms $0.00 Free • Reserved
Gemini Embedding 2 Vector Memory
3,072-dimensional embeddings for episodic semantic memory search 1,000 / day 100 RPM 30,000 TPM ~65ms $0.00 Free • Vector Store
Execution Dispatch Logic

Intelligent Task Allocation & Failover

Incoming requests are classified by media type and required context length before reaching model dispatch:

Conversational Inbound (Text Only)

Direct messages under 4,000 tokens dispatch immediately to Gemma 4 26B. If Gemma 4 approaches 90% quota, automatic failover routes to Gemma 4 31B.

Route: Inbound → Gemma 4 26B → Sanitizer

Shared Reel / Video (Multimodal)

Video clips require massive context capacity. The request dispatches to Gemini 3.5 Flash Lite (250,000 TPM) with extracted keyframes and local OCR overlays.

Route: Video → Keyframes → Gemini 3.5 Flash Lite

Episodic Memory Recall (Vectors)

Query terms are embedded into 3,072 dimensions using Gemini Embedding 2 at 100 RPM and compared against cosine distance metrics in the vector database.

Route: Query → Embedding 2 → Cosine Rerank