Google Releases Gemini 1.5
Published:
Google announced Gemini 1.5 Pro on February 15, 2024 in a research preview that made it available to select developers through Google AI Studio and the Gemini API. The headline capability was a context window of up to one million tokens — roughly 750,000 words, equivalent to approximately 30,000 lines of code, 11 hours of audio, 1 hour of video, or several complete novels in a single request. For comparison, GPT-4 Turbo (announced November 6, 2023) offered a 128,000-token context, and Claude 2.1 (November 2023) offered 200,000 tokens; Gemini 1.5’s one-million-token window was five to eight times larger than any commercially available alternative at the time. Google demonstrated the capability with tasks including analyzing an entire 402-page technical manual to answer detailed questions, understanding a full-length movie provided as video frames, and performing few-shot in-context learning from hundreds of examples rather than the dozens that fit in previous context windows. The model also supported multimodal inputs — text, images, audio, and video — within the same context window, enabling queries like “at what timestamp in this hour-long lecture does the professor make an error in this equation?” Read more
