Google Releases Gemini 1.5

3 minute read

Published:

Google announced Gemini 1.5 Pro on February 15, 2024 in a research preview that made it available to select developers through Google AI Studio and the Gemini API. The headline capability was a context window of up to one million tokens — roughly 750,000 words, equivalent to approximately 30,000 lines of code, 11 hours of audio, 1 hour of video, or several complete novels in a single request. For comparison, GPT-4 Turbo (announced November 6, 2023) offered a 128,000-token context, and Claude 2.1 (November 2023) offered 200,000 tokens; Gemini 1.5’s one-million-token window was five to eight times larger than any commercially available alternative at the time. Google demonstrated the capability with tasks including analyzing an entire 402-page technical manual to answer detailed questions, understanding a full-length movie provided as video frames, and performing few-shot in-context learning from hundreds of examples rather than the dozens that fit in previous context windows. The model also supported multimodal inputs — text, images, audio, and video — within the same context window, enabling queries like “at what timestamp in this hour-long lecture does the professor make an error in this equation?”

Gemini 1.5 Pro used a Mixture of Experts (MoE) architecture, departing from the dense Transformer architecture of Gemini 1.0. In an MoE model, the total number of parameters is divided among many “expert” subnetworks, and a routing mechanism selects a small subset of experts to process each token — meaning that while the model’s total parameter count is large, only a fraction are activated for any given forward pass. This substantially reduces the compute cost per token relative to a dense model of equivalent total capacity, while allowing the model to maintain specialized pathways for different input types and domains. Google did not disclose the specific number of parameters or experts in Gemini 1.5 Pro. To evaluate long-context performance rigorously, Google used “needle in a haystack” benchmarks — inserting a specific fact into a very long document and testing whether the model could retrieve it — and reported near-perfect recall across the full one-million-token context length, contrasting with earlier long-context models that showed degraded retrieval accuracy in the middle sections of very long inputs. A 2M-token preview was available to a smaller set of developers simultaneously with the 1M-token version.

Google made Gemini 1.5 Pro generally available in June 2024 with pricing of $3.50 per million input tokens and $10.50 per million output tokens (for inputs under 128K; higher rates applied above that threshold). In May 2024, Google also announced Gemini 1.5 Flash — a smaller, faster, cheaper model in the 1.5 family optimized for high-throughput use cases — priced at $0.35 per million input tokens and $1.05 per million output tokens, making long-context processing economically viable for production applications. For developers, the long context window changed several architectural patterns. Retrieval-Augmented Generation (RAG) systems — which fetched relevant chunks from a vector database and inserted them into a shorter context — could be partially replaced by simply loading entire codebases or document corpora into context directly, at least for smaller repositories. Agentic systems could maintain longer action histories without truncation. Multi-modal applications could analyze full conversations including images and voice recordings in a single pass. The context window race that Gemini 1.5 accelerated drove rapid iteration from competitors: Anthropic extended Claude’s context to one million tokens in the subsequent model generation, and OpenAI expanded GPT-4o’s effective context handling, making long-context capability a standard expectation for frontier models within approximately one year of Gemini 1.5’s announcement.