Updated 28 Jul 2026

The AI Model Timeline

Every major AI model in the order it shipped — from AlexNet and Google's Transformer paper to trillion-parameter open weights landing weekly. Expand any release to see what actually changed, in plain English.

161
releases tracked
40
labs
83
open weights
New here? How AI grew up, in six steps

The whole timeline below is one story: a system that learned to see, then to read, then to be helpful, then to think, and finally to do the work itself. Pick a step to see only the releases from that era. Swipe the cards sideways.

2026

52 releases
·AnthropicLandmarkNew

Claude Opus 5

Anthropic's default for complex agentic coding, at half the price of Fable 5 with a May 2026 knowledge cutoff.

What changed in this release
  • Anthropic's default for complex agentic coding and enterprise work.
  • Half the price of Fable 5 with a 1 million token context.
  • Knowledge current to May 2026 — the most recent cutoff of any Claude model.
What it means

The current recommended workhorse: near the top of the range, at a price that makes sustained use practical.

ProprietaryReasoning1M tokens
·Sakana AINew

Fugu-Ultra v1.1

Sakana's evolutionary model-merging approach, scaled up.

What changed in this release
  • Scaled up Sakana's approach of evolving and merging existing models rather than training from scratch.
What it means

A genuinely different recipe: breed models from existing ones instead of building each from nothing.

ProprietaryReasoning
·Black Forest LabsLandmarkNew

FLUX 3

Black Forest Labs' first multimodal frontier model — images plus 20-second video with audio, phased rollout starting with video in early access.

What changed in this release
  • Black Forest Labs' first multimodal model: images plus 20-second video with audio.
  • Phased rollout, with only the video variant in early access initially.
What it means

The leading open image lab moved to video — and, notably, did not release these weights openly.

ProprietaryVideo
·xAINew

Grok STT 1.0

xAI's first dedicated speech-to-text model.

What changed in this release
  • xAI's first dedicated speech-to-text model.
What it means

Every major lab now wants to own the full voice pipeline, not just the text in the middle.

ProprietaryAudio
·Ant GroupNew

Ling-3.0-flash

Efficiency-focused MoE, closing out the busiest release week on record.

What changed in this release
  • An efficiency-focused mixture-of-experts, closing out the busiest release week on record.
What it means

Seven significant models in seven days — an unprecedented concentration.

Open weightsReasoning
·Google DeepMindNew

Gemini 3.5 Flash-Lite + Flash Cyber

Cheapest 3.5 tier plus a security-specialised variant.

What changed in this release
  • The cheapest tier in the 3.5 family, plus a variant specialised for security work.
What it means

Specialised variants for specific industries are becoming standard.

ProprietaryMultimodal
·Google DeepMindNew

Gemini 3.6 Flash

Announced together with 3.5 Flash-Lite and 3.5 Flash Cyber — three models in one drop.

What changed in this release
  • Announced together with two sibling models in a single drop.
What it means

Labs now ship families rather than models, so there is something for every price point at once.

ProprietaryMultimodal
·poolsideNew

Laguna S 2.1

Open-weight coding model from a lab that had been entirely closed until now.

What changed in this release
  • A previously entirely closed lab released open coding weights.
What it means

Open weights used as a distribution strategy, not an ideology.

Open weightsCode118B (8B active)OpenMDW 1.1
·AlibabaNew

Qwen-Image-3.0

Completed Alibaba's three-model, three-modality week.

What changed in this release
  • Completed a text, audio and image release inside a single week.
What it means

One lab covering every modality in three days — the scale of the current arms race.

ProprietaryImage
·AlibabaNew

Qwen-Audio-3.0-TTS

Flash and Plus speech tiers, shipped a day after the Max preview.

What changed in this release
  • Dedicated speech generation in two speed and quality tiers.
What it means

Voice is now its own product line rather than a feature of the main model.

ProprietaryAudio
·AlibabaNew

Qwen3.8-Max-Preview

First of three Qwen releases inside a 72-hour window.

What changed in this release
  • A 2.4 trillion parameter preview, first of three Qwen releases in 72 hours.
What it means

Release cadence compressed to days, which is genuinely difficult to keep up with.

ProprietaryReasoning2.4T
·Moonshot AINew

Kimi K3

Moonshot's largest model, and the point at which the Kimi flagship went closed.

What changed in this release
  • 2.8 trillion parameters, Moonshot's largest.
  • The flagship went closed after four open generations.
What it means

The same pattern again: build an audience with open weights, then close the very top.

ProprietaryReasoning2.8T
·Thinking Machines LabLandmarkNew

Inkling

Thinking Machines' debut — near-trillion-scale weights released under Apache 2.0.

What changed in this release
  • A new lab's debut at 975 billion parameters, released under a fully permissive licence.
What it means

A brand-new company opened with near-trillion-scale open weights — the barrier to entry has moved, but not as far as it looks.

Open weightsReasoning975B (41B active)Apache 2.0
·OpenAILandmarkNew

GPT-5.6 (Sol, Terra, Luna)

Three tiers after a staged government-safety review from late June. Sol is tuned for biology, chemistry and cybersecurity, and reported ~54% more token-efficient on coding.

What changed in this release
  • Three tiers: Sol, Terra and Luna, from strongest to cheapest.
  • Held back for a staged government safety review before public release.
  • Sol is reported to be about 54% more token-efficient on coding, and is tuned for biology, chemistry and cybersecurity.
What it means

The safety review is the story. Governments now sit between a frontier model being finished and being released.

ProprietaryReasoning
·Meta AINew

Muse Spark 1.1

First iteration on Meta's closed flagship line.

What changed in this release
  • First iteration on Meta's closed flagship.
What it means

Meta's closed strategy settling into a regular cadence.

ProprietaryMultimodal
·xAINew

Grok 4.5

xAI's agentic flagship, built around long-running autonomous tasks.

What changed in this release
  • Built around long-running autonomous tasks rather than conversation.
What it means

Agents, not chatbots, are now what frontier models are designed for.

ProprietaryReasoning
·Cognition AINew

SWE-1.7

The model behind the Devin coding agent, trained on full engineering trajectories.

What changed in this release
  • Trained on complete engineering work sessions rather than isolated code snippets.
What it means

Learns how software actually gets built — including the debugging and backtracking.

ProprietaryCode
·TencentNew

Hy3

Tencent's Hunyuan line moved to a fully permissive licence.

What changed in this release
  • Tencent's model family moved to a fully permissive licence.
What it means

Another major platform choosing openness to attract developers.

Open weightsMultimodal295B (21B active)Apache 2.0
·AnthropicNew

Claude Sonnet 5

Roughly Opus 4.8 capability at Sonnet pricing, with introductory rates through August 2026.

What changed in this release
  • Roughly the capability of the previous flagship at mid-tier pricing.
  • Introductory rates lower still through August 2026.
What it means

Last quarter's frontier at half the price — the clearest illustration of how fast costs fall.

ProprietaryReasoning1M tokens
·MeituanNew

LongCat-2.0

A food-delivery company shipping one of the largest models of the year.

What changed in this release
  • A 1.6 trillion parameter model from a food delivery company.
What it means

Frontier-scale training is now within reach of any large technology business.

Open weightsReasoning1.6T (48B active)
·Z.ai

GLM-5.2

Z.ai held to a roughly two-month open-weight cadence through 2026.

What changed in this release
  • Continued Z.ai's roughly two-month open-weight cadence.
What it means

Predictable, frequent open releases became a competitive strategy in itself.

Open weightsReasoning744B (40B active)MIT
·Moonshot AI

Kimi K2.7 Code

A code-specialised trillion-parameter model with open weights.

What changed in this release
  • A trillion-parameter model specialised for code, with open weights.
What it means

Free weights aimed squarely at the most commercially valuable task.

Open weightsCode1T (32B active)Modified MIT
·AnthropicLandmark

Claude Fable 5 + Mythos 5

Anthropic's most capable widely released model, with always-on adaptive thinking. Mythos 5 stayed invitation-only under Project Glasswing.

What changed in this release
  • Anthropic's most capable widely available model, built for long-running agents.
  • Thinking is always on and cannot be switched off.
  • A sibling, Mythos 5, stayed invitation-only for defensive cybersecurity work.
What it means

Notable for what came with it: the most capable variant was restricted rather than sold, a first for a mainstream lab.

ProprietaryReasoning1M tokens
·NVIDIA

Nemotron 3 Ultra

NVIDIA's largest open release, tuned for its own inference stack.

What changed in this release
  • NVIDIA's largest open release, tuned for its own inference software.
What it means

The chip maker gives away models to make its hardware the obvious place to run them.

Open weightsReasoning550B (55B active)OpenMDW 1.1
·MiniMax

MiniMax-M3

Agentic MoE aimed squarely at long-horizon tool use.

What changed in this release
  • An efficient mixture-of-experts built for long-horizon tool use.
What it means

Purpose-built for agents that run for hours, not chat sessions that run for minutes.

Open weightsReasoning428B (23B active)MiniMax Community
·StepFun

Step 3.7 Flash

Multimodal refresh of the sparse Flash line.

What changed in this release
  • A multimodal refresh of the extremely sparse Flash line.
What it means

Cheap multimodal AI keeps getting cheaper.

Open weightsMultimodal198B (11B active)Apache 2.0
·Anthropic

Claude Opus 4.8

Effort defaulted to high across every surface, including Claude Code.

What changed in this release
  • Maximum thinking effort became the default everywhere, including in Claude Code.
What it means

Quality over speed as the out-of-the-box choice, reflecting cheap enough inference to afford it.

ProprietaryReasoning1M tokens
·Alibaba

Qwen3.7-Max

The closed Max line kept pace with the open Qwen releases.

What changed in this release
  • Alibaba's closed flagship kept pace with its own open releases.
What it means

A dual strategy: give away the strong models, sell the strongest.

ProprietaryReasoning
·Google DeepMind

Gemini 3.5 Flash

Frontier sustained performance shipped in the Flash tier rather than Pro.

What changed in this release
  • Sustained frontier performance shipped in the cheap, fast tier rather than the premium one.
What it means

The premium tier stopped being where the interesting capability lives.

ProprietaryReasoning
·Cursor

Composer 2.5

An IDE company training its own trillion-parameter agentic coding model.

What changed in this release
  • A code editor company trained its own trillion-parameter model.
What it means

Application companies started building their own models rather than renting someone else's.

ProprietaryCode1T
·Mistral AI

Mistral Medium 3.5

Mistral's dense flagship, still open at a time when most flagships were not.

What changed in this release
  • A dense 128B flagship — no mixture-of-experts — kept open.
What it means

Europe's leading lab held the line on open weights as US labs closed up.

Open weightsReasoning128BModified MIT
·Xiaomi

MiMo-V2.5-Pro

Tuned specifically for agentic coding, alongside an omni-modal 310B sibling.

What changed in this release
  • Tuned specifically for agentic coding, with an omni-modal sibling.
What it means

Specialisation is now how labs differentiate, since raw capability has largely converged.

ProprietaryReasoning1.02T
·DeepSeekLandmark

DeepSeek-V4 (Flash + Pro)

Split into a cheap Flash tier and a 1.6T Pro tier, both under MIT.

What changed in this release
  • Split into a cheap Flash tier and a 1.6 trillion parameter Pro tier.
  • Both released under a fully permissive licence.
What it means

A 1.6T model given away free — an unthinkable proposition eighteen months earlier.

Open weightsReasoning284B / 1.6TMIT
·OpenAI

GPT-5.5

Shipped alongside a Pro tier for the hardest reasoning workloads.

What changed in this release
  • A general capability step, with a Pro tier for the hardest reasoning.
What it means

Paying more for extra thinking time became a standard product tier.

ProprietaryReasoning
·Moonshot AI

Kimi K2.6

Kept trillion-parameter open weights on a roughly quarterly cadence.

What changed in this release
  • Kept trillion-parameter open weights on a roughly quarterly release cycle.
What it means

Frontier-scale open releases became routine rather than exceptional.

Open weightsReasoning1T (32B active)Modified MIT
·Anthropic

Claude Opus 4.7

Introduced a new tokenizer, so token counts shifted roughly 30% against earlier models.

What changed in this release
  • A new tokenizer — the way text is chopped into pieces for the model.
  • The same text now counts as roughly 30% more tokens than on older models.
What it means

Worth knowing if you pay per token: identical inputs bill differently across this boundary.

ProprietaryReasoning1M tokens
·Alibaba

Qwen3.6

3B active parameters at near-frontier quality — the efficiency frontier of its moment.

What changed in this release
  • Near-frontier quality with only 3B parameters active per word.
What it means

Strong reasoning cheap enough to run at consumer scale.

Open weightsReasoning35B (3B active)Apache 2.0
·Meta AILandmark

Muse Spark

Meta's pivot away from open Llama releases toward a closed flagship line.

What changed in this release
  • Meta's new flagship line, kept closed.
  • A clear break from the open Llama strategy it had pursued since 2023.
What it means

The company that made open weights mainstream stepped back from it — a real loss for the open ecosystem.

ProprietaryMultimodal
·Z.ai

GLM-5.1

Refinement of GLM-5 focused on agentic tool use.

What changed in this release
  • Refined GLM-5 for tool use and multi-step agent work.
What it means

Open models tuned specifically for doing things, not just answering.

Open weightsReasoning744B (40B active)MIT
·Google DeepMind

Gemma 4

Gemma moved to a true Apache licence, dropping the custom terms.

What changed in this release
  • Google's open family moved to a genuinely standard permissive licence.
What it means

Removed the custom terms that had made Gemma awkward for commercial use.

Open weightsMultimodalup to 31BApache 2.0
·Xiaomi

MiMo-V2-Pro

A consumer-electronics company shipping a trillion-parameter frontier model.

What changed in this release
  • A consumer electronics manufacturer shipped a trillion-parameter frontier model.
What it means

Frontier AI stopped being confined to AI companies.

ProprietaryReasoning1T (43B active)
·Mistral AI

Mistral Small 4

6B active parameters — "small" now means active cost, not total size.

What changed in this release
  • 119B total parameters, only 6B active per word.
What it means

'Small' now describes running cost, not knowledge. These two numbers have fully decoupled.

Open weightsReasoning119B (6B active)Apache 2.0
·OpenAI

GPT-5.4

Mid-cycle frontier refresh, four months after GPT-5.2.

What changed in this release
  • A mid-cycle capability refresh across the GPT-5 line.
What it means

Frontier releases moved to a roughly two-month rhythm, from roughly annual two years earlier.

ProprietaryReasoning
·Sarvam AI

Sarvam-105B

India's first independently trained large model, built for Indic languages.

What changed in this release
  • India's first large model trained independently end to end, built around Indic languages.
What it means

Sovereign AI became real. Countries increasingly want models trained on their own languages and values.

Open weightsReasoning105BApache 2.0
·Anthropic

Claude Sonnet 4.6

Sonnet pricing with a 1M window and adaptive thinking.

What changed in this release
  • Mid-tier pricing with the 1M context and adaptive thinking from the top tier.
What it means

Long-context work stopped being a premium feature.

ProprietaryReasoning1M tokens
·Alibaba

Qwen3.5

Apache-licensed frontier-class weights, still free to use commercially.

What changed in this release
  • Frontier-class capability at 397B total, 17B active, fully permissive.
What it means

The strongest weights anyone can download and use commercially without asking permission.

Open weightsReasoning397B (17B active)Apache 2.0
·Z.ai

GLM-5

MIT-licensed weights at a scale that had been strictly proprietary a year earlier.

What changed in this release
  • 744 billion parameters released under a fully permissive licence.
What it means

Weights that would have been the world's most valuable secret two years earlier, given away.

Open weightsReasoning744B (40B active)MIT
·StepFun

Step 3.5 Flash

Extreme sparsity — 11B active out of 196B — aimed at cheap multimodal serving.

What changed in this release
  • 196B total parameters but only 11B active — unusually aggressive sparsity.
What it means

Multimodal AI at a price point that suits high-volume consumer products.

Open weightsMultimodal196B (11B active)Apache 2.0
·Anthropic

Claude Opus 4.6

Opus to a 1M-token window, with adaptive thinking replacing the extended-thinking toggle.

What changed in this release
  • Top tier moved to a 1 million token context.
  • Thinking became adaptive by default rather than a switch you set.
What it means

Whole codebases fit in one conversation, and you no longer configure how hard the model thinks.

ProprietaryReasoning1M tokens
·OpenAI

GPT-5.3-Codex

A frontier release shipped code-first rather than chat-first.

What changed in this release
  • A frontier release that led with coding rather than chat.
What it means

Software work is now the highest-value use of these models, and the release schedule reflects it.

ProprietaryCode
·Moonshot AILandmark

Kimi K2.5

A trillion-parameter open MoE — the largest open weights anyone had published.

What changed in this release
  • A trillion parameters with open weights — the largest ever published at the time.
  • Only 32B run per word, so it stays affordable to serve.
What it means

Trillion-parameter models stopped being the exclusive property of the largest US labs.

Open weightsReasoning1T (32B active)Modified MIT
·Alibaba

Qwen3-Max-Thinking

Alibaba's first closed flagship, reasoning with tool use built in.

What changed in this release
  • Alibaba's first flagship held back from open release, with tool use built into its reasoning.
What it means

Even the most open-friendly labs started keeping their very best work proprietary.

ProprietaryReasoning

2025

31 releases
·Z.ai

GLM-4.7

Closed most of the remaining gap to closed frontier models on agentic benchmarks.

What changed in this release
  • Closed most of the remaining gap to closed models on agent-style tasks.
What it means

By the end of 2025 the best free weights were within touching distance of the best paid APIs.

Open weightsReasoning355B (32B active)MIT
·OpenAI

GPT-5.2

Shipped with a Codex variant tuned specifically for agentic software work.

What changed in this release
  • Better reasoning, plus a Codex variant tuned for autonomous software work.
What it means

Coding became important enough to justify its own dedicated frontier model.

ProprietaryReasoning
·DeepSeek

DeepSeek-V3.2

Sparse attention cut long-context inference cost by roughly half.

What changed in this release
  • A sparse attention mechanism roughly halved the cost of long-context work.
What it means

Long documents got cheaper to process, not just possible.

Open weightsReasoning685BMIT
·Anthropic

Claude Opus 4.5

Opus pricing cut to a third, with an effort parameter to trade cost against depth.

What changed in this release
  • Top-tier pricing cut to roughly a third.
  • An effort setting to trade cost against depth of thinking.
What it means

Frontier capability moved within reach of small teams and individuals.

ProprietaryReasoning200k tokens
·Allen Institute for AI

Olmo 3

Fully open reasoning models with every intermediate checkpoint published.

What changed in this release
  • Fully open reasoning models with every training checkpoint published.
What it means

The only reasoning models researchers can inspect end to end to understand how the ability forms.

Open weightsReasoning7B / 32BApache 2.0
·Google DeepMindLandmark

Gemini 3 Pro

Shipped with Deep Think and a generative UI surface. Retook the frontier.

What changed in this release
  • A Deep Think mode for the hardest problems.
  • Can generate a working interface on the fly rather than replying in prose.
What it means

Google retook the lead, and answers started arriving as interactive tools rather than paragraphs.

ProprietaryReasoning1M tokens
·OpenAI

GPT-5.1

Instant and Thinking variants with adaptive reasoning effort, plus warmer default tone.

What changed in this release
  • Adjusts how long it thinks based on how hard the question looks.
  • Warmer default tone after complaints that GPT-5 felt cold.
What it means

Faster on easy questions, more careful on hard ones, without you having to say which is which.

ProprietaryReasoning
·Anthropic

Claude Haiku 4.5

Near-frontier quality at the fastest tier — still the current Haiku as of this dataset.

What changed in this release
  • The cheapest tier reached roughly the capability of the flagship from five months earlier.
What it means

A good illustration of the pace: today's budget option is last spring's state of the art.

ProprietaryReasoning200k tokens
·Z.ai

GLM-4.6

200k context and a large jump in real-world coding evaluations.

What changed in this release
  • 200k context and a large jump on practical coding tasks.
What it means

Open models kept closing the gap on the closed frontier, quarter after quarter.

Open weightsReasoning357B (32B active)Apache 2.0
·OpenAI

Sora 2

Synchronised audio and a social app attached. Made AI video a consumer feed.

What changed in this release
  • Added synchronised audio and much better physical realism.
  • Shipped with a social app for sharing generated clips.
What it means

AI video went from a professional tool to a consumer feed, with all the identity problems that implies.

ProprietaryVideo
·Anthropic

Claude Sonnet 4.5

30+ hours of autonomous coding without losing the thread.

What changed in this release
  • Held focus on a single coding task for more than 30 hours without losing the thread.
What it means

Agents that work overnight on a real task became a practical proposition.

ProprietaryReasoning200k tokens
·DeepSeek

DeepSeek-V3.1

Merged chat and reasoning into one hybrid-thinking checkpoint.

What changed in this release
  • Merged the separate chat and reasoning models into a single hybrid.
What it means

One model to host instead of two, which halves the cost of running your own.

Open weightsReasoning671B (37B active)MIT
·OpenAILandmark

GPT-5

A router that picks between fast and thinking modes per request, ending the model-picker era.

What changed in this release
  • A router picks between a fast model and a thinking model automatically for each message.
  • Removed the model picker that had confused most users.
  • Substantially lower rates of confidently stating false things.
What it means

You stopped needing to know which model to use. The reception was mixed — many people missed being able to choose.

ProprietaryReasoning400k tokens
·Anthropic

Claude Opus 4.1

Incremental agentic-coding upgrade over Opus 4.

What changed in this release
  • Incremental gains on agentic coding and multi-step tasks over Opus 4.
What it means

A steady refinement rather than a leap.

ProprietaryReasoning200k tokens
·OpenAILandmark

gpt-oss-120b / 20b

OpenAI's first open weights since GPT-2, six years later.

What changed in this release
  • OpenAI's first open weights since GPT-2 in 2019.
  • Two sizes, both with visible reasoning, under a permissive licence.
What it means

A notable reversal from the company that had led the industry into secrecy — largely a response to Chinese open releases.

Open weightsReasoning117B / 21BApache 2.0
·Z.ai

GLM-4.5

Agentic-first open MoE, MIT-licensed, priced far below the frontier.

What changed in this release
  • Designed around tool use and agents from the start, released under a fully permissive licence.
What it means

Strong agent capability at a fraction of frontier pricing.

Open weightsReasoning355B (32B active)MIT
·xAI

Grok 4

Heavy variant ran multiple agents in parallel and compared their answers.

What changed in this release
  • A 'heavy' mode runs several copies in parallel and compares their answers before replying.
What it means

Spending more compute per question, rather than per model, as a way to buy accuracy.

ProprietaryReasoning
·AnthropicLandmark

Claude Opus 4 / Sonnet 4

Sustained multi-hour autonomous coding. The generation that made long-running agents practical.

What changed in this release
  • Sustained focus on a single task for hours instead of minutes.
  • Substantially better at remembering what it had already tried.
What it means

Made autonomous coding agents genuinely practical — you can hand over a task and come back later.

ProprietaryReasoning200k tokens
·Google DeepMindLandmark

Veo 3

First major video model to generate synchronised dialogue and sound effects in one pass.

What changed in this release
  • Generates matching dialogue, sound effects and ambience along with the video, in one pass.
  • Previous video models produced silent footage.
What it means

Complete audiovisual scenes from a sentence. The point where AI video became hard to dismiss.

ProprietaryVideo
·AlibabaLandmark

Qwen3

Hybrid thinking across eight sizes, all Apache 2.0. Became the open default almost immediately.

What changed in this release
  • Eight sizes from 0.6B to 235B, all fully permissive.
  • Every one supports switching thinking on or off per request.
What it means

Frontier-adjacent reasoning became free and self-hostable across every hardware budget.

Open weightsReasoning0.6B–235B (22B active)Apache 2.0
·OpenAI

OpenAI o3 / o4-mini

First reasoning models that could use tools inside the chain of thought, including image manipulation.

What changed in this release
  • First reasoning models that use tools mid-thought — searching, running code, cropping images.
  • Can zoom into part of a photo while reasoning about it.
What it means

Thinking and acting merged. The model can go and check rather than guessing from memory.

ProprietaryReasoning
·OpenAI

GPT-4.1

API-only, tuned for instruction following and long-context coding.

What changed in this release
  • API-only, tuned for following instructions exactly and handling million-token codebases.
What it means

Built for developers wiring models into software, not for people chatting.

ProprietaryMultimodal1M tokens
·Meta AI

Llama 4

Meta's first MoE generation, with a 10M-token context claim on Scout.

What changed in this release
  • Meta's first mixture-of-experts generation.
  • Claimed a 10 million token context on the smaller variant.
What it means

Received a notably mixed reception, and marked the beginning of Meta's retreat from open releases.

Open weightsMultimodal109B–400B (17B active)Llama 4 Community
·Google DeepMindLandmark

Gemini 2.5 Pro

Thinking built into the base model. Took the top of the leaderboards and held it for months.

What changed in this release
  • Reasoning built into the base model rather than offered as a separate mode.
  • Took the top of the public leaderboards and stayed there for months.
What it means

Thinking before answering became a standard feature rather than a premium option.

ProprietaryReasoning1M tokens
·Google DeepMind

Gemma 3

Vision, 128k context and 140 languages, sized to run on a single GPU.

What changed in this release
  • Added vision, 128k context and 140 languages while still fitting on one GPU.
What it means

Serious multimodal capability for anyone with a single graphics card.

Open weightsMultimodal1B–27BGemma Terms
·OpenAI

GPT-4.5

The largest non-reasoning model OpenAI shipped, and the last of that line.

What changed in this release
  • The largest model OpenAI ever trained without a reasoning stage.
  • Warmer and more natural, but no better at hard problems.
What it means

Quietly marked the end of an era: making models bigger stopped being the way forward. Everything after this reasons instead.

ProprietaryText
·AnthropicLandmark

Claude 3.7 Sonnet

First hybrid reasoning model — instant or extended thinking from one set of weights. Shipped with Claude Code.

What changed in this release
  • One model that can answer instantly or think at length, controlled by the caller.
  • Previously these were separate models you had to choose between.
  • Launched alongside Claude Code, a terminal-based coding agent.
What it means

You stopped having to pick the right model for the job. The model decides how hard to think.

ProprietaryReasoning200k tokens
·xAI

Grok 3

Trained on the 100k-H100 Colossus cluster, built in 122 days.

What changed in this release
  • Trained on a 100,000-GPU cluster that xAI built in 122 days.
  • Added a reasoning mode competitive with the best available.
What it means

Demonstrated that the physical build-out, not the research, had become the bottleneck.

ProprietaryReasoning
·Google DeepMind

Gemini 2.0 (GA)

Flash, Flash-Lite and Pro to general availability.

What changed in this release
  • The 2.0 family moved from preview to full production availability.
What it means

Google's long-context, natively multimodal stack became something businesses could depend on.

ProprietaryMultimodal1M tokens
·DeepSeekLandmark

DeepSeek-R1

Open-weight o1-class reasoning with visible chain of thought. Wiped ~$600B off NVIDIA's market cap in a day.

What changed in this release
  • Matched OpenAI's o1 reasoning model, with the weights released free.
  • Showed its full thinking process, which o1 deliberately hid.
  • Trained mostly by rewarding correct answers rather than copying human-written reasoning.
What it means

Wiped roughly $600 billion off NVIDIA's value in a single day. It suggested frontier AI might not need the compute everyone had priced in.

Open weightsReasoning671B (37B active)MIT
·MiniMax

MiniMax-Text-01

Lightning attention pushed the usable context to 4M tokens.

What changed in this release
  • A modified attention mechanism pushed the usable context to about 4 million tokens.
What it means

Enough to hold an entire technical library in view at once.

Open weightsText456B (45.9B active)MiniMax Model License

2024

25 releases
·DeepSeekLandmark

DeepSeek-V3

GPT-4-class for a reported $5.6M in compute. The number that rewrote everyone's training budget.

What changed in this release
  • GPT-4-class performance for a reported $5.6 million of compute — roughly a tenth of the going rate.
  • Achieved on export-restricted, deliberately slower chips.
  • Weights released free under a permissive licence.
What it means

Broke the assumption that only companies spending billions could reach the frontier. The financial markets noticed a month later.

Open weightsText671B (37B active)MIT
·Microsoft

Phi-4

Out-reasoned much larger models on maths benchmarks, again on synthetic data.

What changed in this release
  • A 14B model out-reasoning much larger ones on maths, trained largely on synthetic data.
What it means

Reinforced that how you train matters more than how big you build.

Open weightsText14BMIT
·Google DeepMind

Gemini 2.0 Flash

Native audio and image output plus a live streaming API, at Flash pricing.

What changed in this release
  • Generates images and speech directly, rather than calling separate systems.
  • Added a live API you can stream audio and video into continuously.
What it means

Made continuous, real-time AI assistants cheap enough to deploy widely.

ProprietaryMultimodal1M tokens
·OpenAILandmark

Sora

Ten months after the teaser, text-to-video shipped as a product.

What changed in this release
  • Text-to-video producing up to a minute of footage with consistent characters and physics.
  • Ten months between the demo and anyone being able to use it.
What it means

Video generation became a real product, with immediate consequences for stock footage and advertising.

ProprietaryVideo
·Amazon

Amazon Nova

Micro / Lite / Pro / Premier, priced to undercut everyone else on Bedrock.

What changed in this release
  • A full model range priced aggressively below competitors on Amazon's cloud.
What it means

Cloud providers started treating models as a commodity to bundle rather than a premium product.

ProprietaryMultimodal
·Anthropic

Claude 3.5 Sonnet (new) + Haiku

Introduced computer use — the model driving a mouse and keyboard directly.

What changed in this release
  • Computer use: the model can move a cursor, click and type in a real desktop.
  • It looks at screenshots and decides what to do next.
What it means

The shift from AI that answers to AI that acts. Rough at first, but it is the direction everything has gone since.

ProprietaryMultimodal200k tokens
·Meta AI

Llama 3.2

Added vision at the top end and 1B/3B models small enough for on-device use.

What changed in this release
  • Added image understanding at the large sizes.
  • Added 1B and 3B models designed to run on phones.
What it means

Open models arrived on-device, where nothing you type leaves your handset.

Open weightsMultimodal1B–90BLlama 3.2 Community
·Alibaba

Qwen2.5

18T tokens across the full size ladder. The most-forked open family of its generation.

What changed in this release
  • 18 trillion training tokens across the full size range, with strong coding and maths.
What it means

Became the most-adapted open model family of its generation.

Open weightsText0.5B–72BApache 2.0 / Qwen
·OpenAILandmark

OpenAI o1

Spent inference-time compute on a hidden chain of thought. Opened a second scaling axis after pre-training.

What changed in this release
  • Thinks privately before answering, sometimes for minutes.
  • Trained to reason through problems rather than to produce a fluent answer immediately.
  • Jumped from roughly 13% to 83% on a qualifying exam for the International Mathematics Olympiad.
What it means

A second way to make AI smarter that has nothing to do with size: let it think longer at the moment you ask. This is the origin of every 'reasoning' model below.

ProprietaryReasoning
·xAI

Grok-2

xAI's first frontier-competitive model, with image generation attached.

What changed in this release
  • Reached the frontier tier and added image generation with few restrictions.
What it means

xAI went from newcomer to genuine competitor in under a year.

ProprietaryMultimodal
·Black Forest Labs

FLUX.1

Built by the original Stable Diffusion team. Took the open image-generation crown on release.

What changed in this release
  • Built by the people who created Stable Diffusion, at a new lab.
  • Markedly better at hands, faces and text within images.
What it means

Took the open image-generation crown and has held it since.

Open weightsImage12BApache 2.0 (schnell)
·Meta AILandmark

Llama 3.1 405B

31M H100-hours. The first open-weight model to credibly trade blows with the closed frontier.

What changed in this release
  • 405 billion parameters and 31 million hours of top-end GPU time, given away free.
  • Traded blows with GPT-4 and Claude 3.5 on standard tests.
What it means

The first time free, downloadable weights genuinely reached the commercial frontier. Estimated cost to train: hundreds of millions of dollars.

Open weightsText405BLlama 3.1 Community
·AnthropicLandmark

Claude 3.5 Sonnet

Beat Claude 3 Opus at a fifth of the price, and introduced Artifacts. Became the default coding model for a year.

What changed in this release
  • The mid-tier model beat the previous top-tier one at a fifth of the price.
  • Introduced Artifacts — generated code and documents appearing in a live side panel.
What it means

Became the default coding assistant for roughly a year, and showed the interface matters as much as the model.

ProprietaryMultimodal200k tokens
·NVIDIA

Nemotron-4 340B

Released explicitly as a synthetic-data generator for training other models.

What changed in this release
  • Released explicitly as a tool for generating synthetic training data for other models.
What it means

Models started being used to teach other models — an important step as high-quality human text runs short.

Open weightsText340BNVIDIA Open Model
·Alibaba

Qwen2

The point at which Chinese open-weight models became the default choice for many builders.

What changed in this release
  • A complete open family from 0.5B to 72B with strong multilingual coverage.
What it means

Chinese labs became a default choice for developers worldwide, not just domestically.

Open weightsText0.5B–72BApache 2.0 / Qwen
·OpenAILandmark

GPT-4o

One model for text, vision and audio end to end, with ~320ms voice latency. Halved the price of GPT-4.

What changed in this release
  • One model handling text, vision and audio directly, instead of three chained together.
  • Voice replies in about 320 milliseconds — roughly human conversational speed.
  • Half the price of GPT-4 and free for everyone.
What it means

Real-time spoken conversation with AI became normal, and frontier-grade capability became free.

ProprietaryMultimodal128k tokens
·Microsoft

Phi-3

Small enough to run on a phone while matching models 10× larger.

What changed in this release
  • A 3.8B model matching far larger ones, small enough to run offline on a phone.
What it means

AI that works with no internet connection and no data leaving your device.

Open weightsText3.8B–14BMIT
·Meta AI

Llama 3

15T training tokens — far past Chinchilla-optimal, because inference cost matters more than training cost.

What changed in this release
  • 15 trillion training tokens — far more than the 'optimal' amount.
  • Deliberately overtrained so the finished model would be cheap to run.
What it means

Spend more once on training to save on every query forever. Standard practice now.

Open weightsText8B / 70BLlama 3 Community
·Mistral AI

Mixtral 8x22B

Scaled the open MoE recipe up, still Apache-licensed.

What changed in this release
  • Scaled the open mixture-of-experts recipe up while staying fully permissive.
What it means

Kept efficient open models within reach of the frontier.

Open weightsText141B (39B active)Apache 2.0
·Cohere

Command R+

Built specifically for RAG and tool use rather than chat benchmarks.

What changed in this release
  • Built for retrieval — answering from documents you supply — with inline citations.
  • Tuned for calling external tools reliably.
What it means

Aimed at businesses who need answers grounded in their own documents with a source attached, not clever prose.

Open weightsText104BCC-BY-NC
·Databricks

DBRX

Fine-grained MoE that briefly led every open leaderboard.

What changed in this release
  • A finer-grained mixture-of-experts with 16 smaller experts rather than 8 large ones.
What it means

More specialisation per expert, and briefly the strongest open model available.

Open weightsText132B (36B active)Databricks Open
·AnthropicLandmark

Claude 3 (Opus, Sonnet, Haiku)

The first time a non-OpenAI model took the top spot outright. Established the three-tier naming.

What changed in this release
  • Three tiers — Opus, Sonnet, Haiku — trading cost against capability.
  • Opus beat GPT-4 across most standard tests.
What it means

The first time a company other than OpenAI clearly held the top spot, which turned a one-horse race into a real market.

ProprietaryMultimodal200k tokens
·Google DeepMind

Gemma

Google's first open-weight release, built from the Gemini research stack.

What changed in this release
  • Google's first open-weight release, built from the Gemini research.
What it means

Even the most closed labs concluded they needed an open offering to stay relevant with developers.

Open weightsText2B / 7BGemma Terms
·Google DeepMindLandmark

Gemini 1.5 Pro

First production model past a million tokens. Near-perfect needle-in-a-haystack recall at that length.

What changed in this release
  • Context jumped to one million tokens — hours of video, or an entire large codebase.
  • Retrieved specific details buried anywhere in that window almost perfectly.
What it means

Changed what people build. Rather than indexing documents into a search system, you can now just paste everything in.

ProprietaryMultimodal1M tokens
·Allen Institute for AI

OLMo 7B

Weights, data, training code and checkpoints all released — the fully reproducible option.

What changed in this release
  • Released the weights, the training data, the training code and every intermediate checkpoint.
What it means

'Open' usually means only the weights. This is the version researchers can actually audit and reproduce.

Open weightsText7BApache 2.0

2023

16 releases
·Microsoft

Phi-2

"Textbook-quality" synthetic data beat models 25× its size. Made data curation a first-class lever.

What changed in this release
  • Trained on textbook-quality synthetic material instead of scraped web text.
  • Matched models 25 times its size on reasoning.
What it means

Data quality turned out to substitute for raw scale — which is why capable models now fit on a phone.

Open weightsText2.7BMIT
·Mistral AILandmark

Mixtral 8x7B

Sparse mixture-of-experts with open weights — GPT-3.5 quality at a fraction of the inference cost.

What changed in this release
  • Mixture-of-experts: eight specialist sub-networks, only two of which run per word.
  • Has 47B parameters of knowledge but the running cost of a 13B model.
What it means

The trick that lets today's models be enormous and still affordable to run. Nearly every large model since uses it.

Open weightsText46.7B (12.9B active)Apache 2.0
·Google DeepMindLandmark

Gemini 1.0

Natively multimodal from pre-training rather than bolted on, in Ultra / Pro / Nano sizes.

What changed in this release
  • Trained on text, images, audio and video together from the start rather than adding vision later.
  • Shipped in three sizes, including one small enough for a phone.
What it means

Google's structural answer to GPT-4 — and the reason your phone can now describe what the camera sees.

ProprietaryMultimodal
·Carnegie Mellon / Princeton

Mamba

The most serious challenger to attention: a state-space model that scales linearly with length instead of quadratically. Influential, but the Transformer held the frontier.

What changed in this release
  • Attention compares every word with every other word, so doubling the input roughly quadruples the work.
  • Mamba keeps a running summary instead, so cost grows in a straight line with length — around 5× faster on long inputs.
What it means

The most credible attempt to replace the Transformer. It did not win — attention still holds the frontier — but the ideas turn up inside hybrid models, and it is worth knowing the field did try alternatives.

Open weightsText
·DeepSeek

DeepSeek LLM 67B

The first release from the lab that would go on to reset the field's cost assumptions.

What changed in this release
  • A strong bilingual English/Chinese model with open weights from a new lab.
What it means

Easy to overlook at the time. This lab would go on to halve the industry's assumed cost of frontier training.

Open weightsText67BDeepSeek License
·xAI

Grok-1

xAI's debut, with live access to X. Weights were open-sourced the following March.

What changed in this release
  • Live access to posts on X, so it could answer about events from minutes ago.
  • Deliberately looser content restrictions than competitors.
What it means

The first mainstream model positioned on real-time information and fewer refusals rather than raw capability.

Open weightsText314BApache 2.0
·OpenAI

DALL·E 3

Prompt adherence good enough to render legible text in images — a long-standing failure mode, solved.

What changed in this release
  • Built directly into ChatGPT, so the chatbot rewrote your rough prompt into a detailed one.
  • Could finally render readable text inside images.
What it means

Removed the need to learn 'prompt engineering' for images — you just describe it badly and the model fixes it.

ProprietaryImage
·Mistral AILandmark

Mistral 7B

Beat every 13B model on release. Announced via a bare magnet link.

What changed in this release
  • Beat every 13B model available, at roughly half the size.
  • Used attention tricks that cut memory use on long inputs.
  • Announced with nothing but a download link on social media.
What it means

Made small models respectable. A 7B model you can run on a laptop stopped being a toy.

Open weightsText7.3BApache 2.0
·Stability AI

SDXL 1.0

The Stable Diffusion generation that closed most of the gap to Midjourney.

What changed in this release
  • Roughly three times the image-understanding capacity, with a second model that refines the output.
  • Native 1024×1024 output instead of 512×512.
What it means

Open image generation caught up with the paid tools on quality.

Open weightsImageCreativeML OpenRAIL++
·Meta AILandmark

Llama 2

The first genuinely commercially usable open-weight frontier-adjacent model. 3.3M GPU-hours.

What changed in this release
  • Licensed for commercial use, unlike the original Llama.
  • Added a properly tuned chat version and published the safety testing.
What it means

The moment companies could build products on free model weights without legal risk. Thousands did.

Open weightsText7B–70BLlama 2 Community
·Anthropic

Claude 2

A 100k-token window when rivals offered 8k. Long context became a competitive axis.

What changed in this release
  • Context window jumped to 100,000 tokens — roughly a 300-page book in one go.
  • Rivals were offering 8,000 tokens at the time.
What it means

You could finally hand a model an entire contract or codebase instead of feeding it in pieces and hoping it remembered.

ProprietaryText100k tokens
·Technology Innovation Institute

Falcon-40B

Apache-licensed and trained on the RefinedWeb corpus — briefly the top open model on every leaderboard.

What changed in this release
  • Trained on a carefully filtered web corpus rather than curated books and papers.
  • Released under a fully permissive licence allowing commercial use.
What it means

Proved that well-cleaned web text beats smaller hand-picked datasets, and that serious models could come from outside the US and China.

Open weightsText40BApache 2.0
·Google

PaLM 2

Powered Bard's relaunch and Google's first serious answer to GPT-4.

What changed in this release
  • Better multilingual and reasoning performance in a smaller, cheaper package than PaLM.
What it means

Google's first genuinely competitive answer to GPT-4, and the relaunch of Bard.

ProprietaryText340B
·Anthropic

Claude 1

Anthropic's first public model, trained with Constitutional AI rather than pure RLHF.

What changed in this release
  • Trained against a written set of principles rather than case-by-case human ratings ('Constitutional AI').
  • The model critiques and revises its own answers against those rules.
What it means

A different bet on how to keep AI well-behaved: write down the values explicitly instead of hoping they emerge from crowd-worker preferences.

ProprietaryText9k tokens
·OpenAILandmark

GPT-4

Bar-exam-passing multimodal reasoning, with architecture details withheld — the moment frontier labs stopped publishing.

What changed in this release
  • Could accept images as well as text.
  • Passed the US bar exam around the 90th percentile, where GPT-3.5 was near the bottom.
  • OpenAI published no size, architecture or training data details.
What it means

The capability jump that triggered the corporate AI rush — and the moment frontier labs stopped telling anyone how their models work.

ProprietaryMultimodal8k–32k tokens
·Meta AILandmark

LLaMA

Leaked within a week of its research-only release, and the local-LLM movement started the same month.

What changed in this release
  • Applied the Chinchilla lesson: modest size, enormous amounts of training data.
  • The 13B version matched GPT-3 at a fraction of the running cost.
  • Released to researchers only — and leaked publicly within a week.
What it means

The leak accidentally created the entire local-AI movement. Within months people were running capable models on laptops.

Open weightsText7B–65BNon-commercial

2022

13 releases
·OpenAILandmark

ChatGPT (GPT-3.5)

The chat wrapper that took the field from research to 100M users in two months.

What changed in this release
  • Wrapped an instruction-tuned GPT-3.5 in a simple chat box, free to use.
  • No new capability — the change was that anyone could try it without reading documentation.
What it means

100 million users in two months, the fastest adoption of any consumer product to that point. The interface, not the model, is what changed the world.

ProprietaryText
·OpenAI

Whisper

Near-human multilingual speech recognition, released free. Effectively ended paid ASR as a default.

What changed in this release
  • Trained on 680,000 hours of varied audio rather than clean studio recordings.
  • Handled 99 languages, accents and background noise, and was released free.
What it means

Accurate transcription went from a paid service to a free download, more or less overnight.

Open weightsAudioMIT
·Stability AILandmark

Stable Diffusion

Open weights that ran on a consumer GPU. Spawned an entire ecosystem of forks, LoRAs and UIs.

What changed in this release
  • Ran the diffusion process in a compressed space, cutting the compute needed by roughly 50×.
  • Released the full weights publicly, so it ran on an ordinary gaming graphics card.
What it means

Image generation stopped being something you rented from a company and became something you owned. Both the creative explosion and the flood of misuse start here.

Open weightsImageCreativeML OpenRAIL-M
·BigScience

BLOOM

A 1,000-researcher collaboration covering 46 languages, trained on a public supercomputer.

What changed in this release
  • A thousand researchers across 70 countries trained one model on a public supercomputer.
  • Covered 46 human languages and 13 programming languages by design.
What it means

A deliberate push back against AI that only works well in English and only exists inside private companies.

Open weightsText176BRAIL
·Google

Imagen

Showed that a big frozen text encoder mattered more than a big image model for prompt fidelity.

What changed in this release
  • Found that a large frozen text model mattered more for image quality than a larger image model.
What it means

Explained why some image generators follow complicated instructions and others ignore half of them.

ProprietaryImage
·Meta AI

OPT-175B

Meta's GPT-3-scale replication, released with the training logbook — unusual transparency for the time.

What changed in this release
  • Meta matched GPT-3's scale and released the weights to researchers.
  • Shipped with the raw training logbook, including the failures.
What it means

Unusual honesty about how messy training a huge model actually is — most labs publish only the wins.

Open weightsText175BNon-commercial
·Google DeepMind

Flamingo

Bolted a vision encoder onto a language model, so images could be handled with a few examples rather than retraining.

What changed in this release
  • Connected an existing vision model to an existing language model rather than training one giant system.
  • Could learn a new visual task from a handful of image–text examples in the prompt.
What it means

The practical recipe for giving a chatbot eyes, and why you can paste a screenshot into a chat today.

ProprietaryMultimodal80B
·OpenAILandmark

DALL·E 2

Diffusion-based image generation good enough to make text-to-image a mainstream product category.

What changed in this release
  • Switched to diffusion — start from noise and repeatedly refine it into an image.
  • Four times the resolution of the original DALL·E with far better prompt fidelity.
What it means

The point where AI images went from a curiosity to something people actually used, and where the argument with working artists began in earnest.

ProprietaryImage
·Google

PaLM

~6,000 TPU v4 chips for ~60 days. The high-water mark for dense scaling.

What changed in this release
  • 540 billion parameters trained across two data centres at once.
  • Started showing step-by-step reasoning when asked to 'think it through'.
What it means

First strong hint that prompting a model to show its working makes it genuinely more accurate — the seed of today's reasoning models.

ProprietaryText540B
·Google DeepMindLandmark

Chinchilla

Showed the field had been training models far too large on far too few tokens. Redirected everyone's compute budgets.

What changed in this release
  • Found that everyone had been building models too big and feeding them too little text.
  • A 70B model trained on far more data beat a 280B model trained the old way.
What it means

Redirected the industry's spending overnight. Bigger stopped being automatically better, and smaller-but-better-fed models made AI dramatically cheaper to run.

ProprietaryText70B
·EleutherAI

GPT-NeoX-20B

Open weights at a scale that had previously been API-only.

What changed in this release
  • Open weights at 20 billion parameters, a scale that had been API-only.
What it means

Kept the open ecosystem within reach of the commercial frontier.

Open weightsText20BApache 2.0
·OpenAILandmark

InstructGPT

Trained on human preference rankings (RLHF) so the model followed instructions instead of just continuing text. The technique behind ChatGPT.

What changed in this release
  • Humans ranked competing model answers, and the model was trained to produce what they preferred.
  • This is 'RLHF' — reinforcement learning from human feedback.
  • A 1.3B model trained this way was preferred over the 175B GPT-3.
What it means

The step that turned a text-continuation engine into an assistant that does what you ask. Being helpful turned out to be a training choice, not a size problem.

ProprietaryText
·Google

LaMDA

Dialogue-tuned model that later became the basis for Bard.

What changed in this release
  • Tuned specifically for open-ended conversation rather than one-off answers.
What it means

The model at the centre of the 2022 story about a Google engineer who believed it was sentient. It was not.

ProprietaryText137B

2021

6 releases
·Google DeepMind

Gopher

DeepMind's first large-scale LM study, published alongside a detailed harms analysis.

What changed in this release
  • DeepMind's systematic study of what does and does not improve as models get bigger.
  • Published alongside an unusually frank analysis of the harms.
What it means

Showed scale helps enormously with knowledge and comprehension, but barely at all with logical reasoning — a gap that took another three years to close.

ProprietaryText280B
·NVIDIA

Megatron-Turing NLG

Three months on 2,000+ A100s. A demonstration that scale was becoming an infrastructure problem.

What changed in this release
  • 530 billion parameters, trained across more than 2,000 top-end chips for three months.
What it means

Made it obvious that frontier AI was becoming an infrastructure business only a handful of organisations could afford.

ProprietaryText530B
·OpenAI

Codex

GPT-3 fine-tuned on public code — the model behind the first GitHub Copilot.

What changed in this release
  • Took GPT-3 and continued training it on public source code.
  • Could turn a plain-English comment into a working function.
What it means

Became GitHub Copilot — the first AI product millions of professionals used in their daily work.

ProprietaryCode12B
·EleutherAI

GPT-J

For a year, the strongest freely downloadable model anyone could run.

What changed in this release
  • Scaled the open replication to 6 billion parameters under a fully permissive licence.
What it means

For about a year this was the best model an individual could download and run themselves.

Open weightsText6BApache 2.0
·EleutherAI

GPT-Neo

First credible open replication of the GPT-3 recipe, built by a volunteer collective.

What changed in this release
  • A volunteer collective rebuilt the GPT-3 approach and gave the weights away free.
What it means

The start of the open-source counterweight to closed labs — a theme running the length of this timeline.

Open weightsText2.7BMIT
·OpenAILandmark

CLIP

Trained on 400M image–caption pairs to place pictures and the words describing them in the same space. This is what lets you type a prompt and get a matching image.

What changed in this release
  • Trained on 400 million image–caption pairs scraped from the web.
  • Learned to put a picture and the words describing it in the same mathematical space, so 'a dog on a skateboard' sits near an actual photo of one.
  • Could classify images it was never explicitly trained on, just by comparing against text descriptions.
What it means

The missing link between language and pictures. Typing a prompt and getting a matching image only works because something taught the machine which words go with which visuals — this is that something.

Open weightsMultimodal

2020

3 releases
·GoogleLandmark

Vision Transformer (ViT)

Chopped an image into patches and fed them to a Transformer like words. One architecture now covered both text and vision.

What changed in this release
  • Cut an image into a grid of small patches and treat each patch like a word in a sentence.
  • The same Transformer built for text then works on pictures, with no vision-specific machinery.
What it means

One architecture for everything. This is why a single model can now handle text, images, audio and video together instead of needing a separate specialist per sense.

Open weightsImage
·UC BerkeleyLandmark

Diffusion Models (DDPM)

Generate an image by starting from pure noise and repeatedly removing a little of it. The method behind DALL·E 2, Stable Diffusion, Midjourney, Sora and Veo.

What changed in this release
  • Take a real image, add noise step by step until it is static, then train a network to undo one step at a time.
  • To generate something new, start from pure static and run the undoing process.
  • More stable to train and higher quality than GANs, which often collapsed.
What it means

This is the engine inside essentially every AI image and video tool you have used — DALL·E 2, Stable Diffusion, Midjourney, Sora, Veo. The pictures come from noise.

Open weightsImage
·OpenAILandmark

GPT-3

175B parameters made few-shot prompting work without fine-tuning — the result that started the scaling race.

What changed in this release
  • More than 100× the size of GPT-2 at 175 billion parameters.
  • Could do new tasks from a couple of examples typed into the prompt, with no retraining at all.
What it means

The birth of prompting. Instead of hiring engineers to retrain a model, you just describe what you want — which is why anyone can use AI today.

ProprietaryText175B2k tokens

2019

3 releases
·Google

T5

Recast every NLP task as text-to-text. Still the base for later Google work including Imagen.

What changed in this release
  • Reframed every task — translation, summarising, answering — as text in, text out.
  • One model and one training method covering jobs that previously needed separate systems.
What it means

The idea that one general model can replace dozens of specialised ones, made explicit.

Open weightsText11BApache 2.0
·Google

XLNet

Permutation language modelling, pitched as the successor to BERT's masked objective.

What changed in this release
  • Predicted words in a random order rather than a fixed left-to-right or blanked-out pattern.
  • Aimed to get bidirectional understanding without BERT's artificial blank tokens.
What it means

A serious technical rival to BERT that showed the field was still exploring how models should read.

Open weightsText340MApache 2.0
·OpenAI

GPT-2

Held back at launch as "too dangerous to release" — the first time staged release became an industry talking point.

What changed in this release
  • Roughly 10× larger than GPT-1 and trained on web pages rather than books.
  • Produced paragraphs coherent enough that OpenAI withheld the full model at first.
What it means

The first serious public argument about whether an AI model was too dangerous to release — a debate that has never stopped.

Open weightsText1.5BMIT

2018

4 releases
·GoogleLandmark

BERT

Bidirectional encoder that dominated NLP benchmarks for years. Encoder-only, so it was never built to be prompted.

What changed in this release
  • Read text in both directions at once by hiding random words and learning to fill the blanks.
  • Set new records across essentially every language-understanding benchmark on release.
What it means

Quietly went into Google Search in 2019 and improved results for millions of queries. Great at understanding text, but it cannot write — it was never built to be prompted.

Open weightsText340MApache 2.0
·OpenAILandmark

GPT-1

The first generative pre-trained transformer — decoder-only, and the template every GPT since has followed.

What changed in this release
  • Applied the Transformer to plain next-word prediction over a large book corpus.
  • Used only the decoder half of the Transformer — predict forward, never peek ahead.
What it means

The blueprint. Everything OpenAI shipped afterwards is this same idea with more data and more compute.

Open weightsText117MMIT
·Allen Institute for AI

ELMo

First widely used contextual embeddings — the same word finally got different vectors in different sentences.

What changed in this release
  • Gave each word a representation that depends on the sentence around it.
  • 'Bank' in 'river bank' and 'bank account' finally got different representations.
What it means

Context-awareness arrived, which is the difference between keyword matching and actual reading comprehension.

Open weightsText94M
·fast.ai

ULMFiT

Showed that pre-training then fine-tuning beats training from scratch — the transfer-learning recipe for text.

What changed in this release
  • Pre-train one model on generic text, then fine-tune it cheaply for each specific task.
  • Beat purpose-built models using a fraction of the labelled examples.
What it means

Established the train-once, adapt-many pattern that makes today's models economically viable.

Open weightsText

2017

1 release
·GoogleLandmark

Transformer ("Attention Is All You Need")

Threw out recurrence and kept only attention. Every model below this line — GPT, BERT, Claude, Gemini, Llama — is a Transformer.

What changed in this release
  • Removed recurrence — the old approach of reading text strictly word by word, in order.
  • Used only 'attention': every word looks at every other word simultaneously and decides what is relevant.
  • Because it no longer had to wait for the previous word, training could run massively in parallel on many chips at once.
What it means

The single most important paper on this page. Every major model after 2018 — GPT, BERT, Claude, Gemini, Llama, all of them — is a Transformer. The parallelism it unlocked is what made training on the entire internet financially possible.

Open weightsText

2016

1 release
·Google DeepMind

WaveNet

Generated raw audio sample by sample. Synthetic speech stopped sounding robotic.

What changed in this release
  • Generated audio one sample at a time — tens of thousands per second — instead of stitching together recorded speech fragments.
What it means

Why voice assistants stopped sounding like a chopped-up recording and started sounding like a person.

ProprietaryAudio

2015

1 release
·Microsoft

ResNet

Skip connections let networks go hundreds of layers deep without collapsing. Depth stopped being the ceiling.

What changed in this release
  • Added shortcut connections that let information skip layers.
  • Networks could go past 100 layers deep without training collapsing, which had been a hard wall.
What it means

Depth stopped being the limit on how capable a network could get — a prerequisite for everything larger that followed.

Open weightsImageup to 152 layers

2014

2 releases
·Google

Seq2Seq

Encoder–decoder networks that map one sequence to another. Made neural machine translation work.

What changed in this release
  • One network (a stack of LSTMs) read a whole input sentence into a compressed summary; a second wrote the output from it.
  • Replaced hand-built translation rules with a single system trained end to end.
What it means

Machine translation stopped being obviously robotic. This is roughly when Google Translate became usable.

Open weightsText
·Université de MontréalLandmark

GANs

Two networks competing — one generating, one detecting fakes. The first credible route to synthetic images.

What changed in this release
  • Set two networks against each other: one inventing fake images, one trying to spot them.
  • Both improved by competing, with no need for a human to score the output.
What it means

The original engine behind convincing fake faces and, eventually, the deepfake problem.

Open weightsImage

2013

1 release
·GoogleLandmark

word2vec

Turned words into vectors where meaning became arithmetic — king − man + woman ≈ queen.

What changed in this release
  • Represented each word as a list of numbers learned from the company it keeps in real text.
  • Words with similar meanings ended up close together, and relationships became arithmetic you could actually do.
What it means

The first time machines had a usable notion of meaning rather than just matching characters. It is why search engines started understanding synonyms.

Open weightsText

2012

1 release
·University of TorontoLandmark

AlexNet

Won ImageNet by such a margin that the entire field switched to deep neural networks within a year.

What changed in this release
  • Trained a deep neural network on two consumer graphics cards instead of ordinary processors.
  • Cut the ImageNet photo-labelling error rate from about 26% to 15% — a bigger jump than the previous several years combined.
What it means

This is the moment computers got genuinely good at seeing. Almost every AI advance since traces back to the realisation that gaming graphics cards could train neural networks.

Open weightsImage60M

1997

1 release
·IDSIALandmark

LSTM — the RNN that worked

Recurrent neural networks (RNNs) read text one step at a time, but forgot the start of a paragraph by the end. LSTM gave them a memory that lasted, and ran speech recognition, translation and text generation for two decades — until the Transformer removed recurrence entirely.

What changed in this release
  • Plain recurrent networks (RNNs) read a sequence one step at a time, carrying a running memory — but that memory faded within a few dozen steps, so they forgot the start of a paragraph by the end of it.
  • LSTM added gates that explicitly decide what to keep, what to discard and what to output, so information could survive hundreds of steps.
What it means

The architecture behind roughly two decades of sequence AI — Google Translate, Siri, autocomplete. It is also what the Transformer replaced: reading strictly one step at a time cannot be parallelised, which capped how large these models could ever get.

Open weightsText

Dates are public announcement dates, compiled 28 Jul 2026 from lab documentation, release trackers and public reporting. Anthropic and the pre-2019 research papers are checked against primary sources; the rest are corroborated across independent trackers, and a handful of 2026 dates differ by days between them.
Covers the foundational architectures (2012–2017) and foundation-model releases from named labs since — not fine-tunes, distills, quantizations, or per-size variants of a family already listed.