TL;DR: The history of AI starts in 1943 with the first mathematical model of a neuron, gets its name at the 1956 Dartmouth workshop, and then survives two funding collapses known as AI winters before deep learning revives it in 2012. This timeline covers every major artificial intelligence milestone from 1943 to 2026, including the exact dates, the people behind them, and the pattern that keeps repeating.
Here is the thing almost every AI history article gets wrong: it reads like a victory lap.
The real story is that artificial intelligence has failed publicly, twice, hard enough that researchers stopped using the words “artificial intelligence” on grant applications because the term had become poison. Between 1974 and 1980, and again between 1987 and 1993, funding evaporated, labs closed, and the field’s biggest names were treated like people who had oversold a dream.
I have been tracking AI news daily on this site since July 2026, and in one 34-day window our pipeline logged 2,234 unique AI stories across 20 active sources. That is roughly 66 new AI stories a day. When you are drinking from that firehose, it is easy to believe AI just appeared in November 2022 and has been going straight up ever since.
It did not. And knowing the actual history of AI is the single best defence against getting fooled by the current hype, in either direction.
So in this guide I am going to walk you through the complete artificial intelligence timeline. Every era, every major milestone, the two winters nobody mentions in the LinkedIn posts, and the honest pattern that connects 1956 to 2026. I will also show you the live data we collect on how fast the field is moving right now, which is the part no static timeline can give you.
Let’s get into it.
Who Created AI and When Was AI Invented?
No single person created AI. The field was formally founded at the Dartmouth Summer Research Project in 1956, where John McCarthy coined the term “artificial intelligence.” But the technical foundations were laid in 1943 by Warren McCulloch and Walter Pitts, who built the first mathematical model of an artificial neuron, and in 1950 by Alan Turing, whose paper “Computing Machinery and Intelligence” asked whether machines could think.
If you want one date for when AI was invented, use 1956. That is when the field got its name, its founding document, and its first generation of researchers in one room.
The room mattered. In the summer of 1955, John McCarthy, Marvin Minsky, Nathaniel Rochester of IBM, and Claude Shannon proposed a summer workshop at Dartmouth College built on a claim that still sounds bold: that “every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it.”
They asked the Rockefeller Foundation for $14,000. The foundation was not impressed and awarded roughly half of it. So the entire academic field of artificial intelligence was launched on about seven thousand dollars and a two-page proposal.
The attendee list is the founding cast of the discipline: McCarthy, Minsky, Shannon, Rochester, Ray Solomonoff, Oliver Selfridge, Trenchard More, Arthur Samuel, Allen Newell, and Herbert Simon. Between them they would produce the first machine learning program, the first AI programming language, the first reasoning systems, and two Turing Awards.
Who invented AI before it was called AI?
Three names matter before 1956.
Warren McCulloch and Walter Pitts (1943) showed that networks of simplified artificial neurons could compute logical functions. Every neural network running today traces back to that paper.
Alan Turing (1950) published “Computing Machinery and Intelligence,” which proposed the Imitation Game, now called the Turing Test. Turing did not ask “can machines think?” He replaced it with a question you can actually test: can a machine convince a human it is human?
Frank Rosenblatt (1958) built the Perceptron, the first trainable neural network, and demonstrated it on IBM hardware before building the physical Mark I Perceptron machine. He is the reason “training a model” is a phrase that exists.
If you are new to the vocabulary here, our AI glossary has plain-English definitions for 264 terms, from perceptron to transformer, each one cited.
The Complete AI Timeline at a Glance
Here is the full artificial intelligence timeline in one table. Every entry below is expanded in the era sections that follow.
| Year | Milestone | Why it mattered |
|---|---|---|
| 1943 | McCulloch and Pitts model the artificial neuron | First mathematical basis for neural networks |
| 1950 | Turing publishes “Computing Machinery and Intelligence” | Introduces the Turing Test |
| 1952 | Arthur Samuel’s checkers program | First program that improved through self-play |
| 1956 | Dartmouth Summer Research Project | The term “artificial intelligence” is coined |
| 1956 | Logic Theorist by Newell, Simon and Shaw | Proved 38 of the first 52 theorems in Principia Mathematica |
| 1958 | Rosenblatt’s Perceptron | First trainable neural network |
| 1958 | McCarthy creates Lisp | The dominant AI language for 30 years |
| 1966 | ELIZA at MIT | First widely known chatbot |
| 1969 | Minsky and Papert publish Perceptrons | Exposed single-layer limits, chilled neural network research |
| 1972 | MYCIN begins at Stanford | Landmark medical expert system |
| 1973 | The Lighthill Report | Triggers collapse of UK AI funding |
| 1974 to 1980 | The first AI winter | Funding and credibility collapse |
| 1980 | XCON deployed at Digital Equipment Corporation | Expert systems prove commercial value |
| 1982 | Japan launches the Fifth Generation project | Sparks a global AI funding race |
| 1986 | Backpropagation popularised by Rumelhart, Hinton and Williams | Made multi-layer networks trainable |
| 1987 to 1993 | The second AI winter | Lisp machine market collapses, expert systems disappoint |
| 1997 | Deep Blue beats Garry Kasparov | First computer to beat a reigning world chess champion |
| 1997 | LSTM introduced by Hochreiter and Schmidhuber | Solved long-range memory in sequence models |
| 2009 | ImageNet released by Fei-Fei Li’s team | 14 million labelled images, the fuel for deep learning |
| 2011 | IBM Watson wins Jeopardy! | Natural language question answering goes mainstream |
| 2012 | AlexNet wins ImageNet with 15.3% top-5 error | Starts the deep learning era |
| 2014 | Generative Adversarial Networks introduced | Breakthrough in generative modelling |
| 2016 | AlphaGo defeats Lee Sedol 4-1 | Landmark for reinforcement learning |
| 2017 | “Attention Is All You Need” introduces the Transformer | The architecture behind every modern LLM |
| 2018 | GPT-1 and BERT released | Pretraining becomes the default method |
| 2020 | GPT-3 ships with 175 billion parameters | Few-shot learning at scale |
| 2020 | AlphaFold 2 solves protein structure prediction | AI delivers a genuine scientific result |
| 2021 | DALL-E and GitHub Copilot | Generative images and AI pair programming |
| 2022 | Stable Diffusion goes open source | Open weights arrive for image generation |
| 2022 | ChatGPT launches on 30 November | 100 million users in two months |
| 2023 | GPT-4, Claude, Bard, Llama 2 | The frontier model race begins |
| 2024 | o1 reasoning models, EU AI Act, Nobel Prizes for AI | Reasoning, regulation, recognition |
| 2025 | DeepSeek R1, GPT-5, Gemini 3 | Open reasoning models close the gap |
| 2026 | GPT-5.5, Claude Opus 4.8, Claude Fable 5 | Release cycles compress to weeks |
That table is the short version. The interesting part is what happened between the rows.
The People Behind the AI Timeline
The history of AI is the work of a few dozen people, most of whom worked on it for decades before it paid off. This table maps the major figures to the specific contribution they are remembered for.
| Person | Contribution | Year | Recognition |
|---|---|---|---|
| Alan Turing | “Computing Machinery and Intelligence,” the Turing Test | 1950 | Turing Award named after him |
| Warren McCulloch & Walter Pitts | First mathematical model of an artificial neuron | 1943 | Foundational to all neural networks |
| John McCarthy | Coined “artificial intelligence,” created Lisp | 1955, 1958 | Turing Award 1971 |
| Marvin Minsky | Co-founded MIT AI Lab, co-authored Perceptrons | 1959, 1969 | Turing Award 1969 |
| Allen Newell & Herbert Simon | Logic Theorist, General Problem Solver | 1956, 1957 | Turing Award 1975; Simon also won a Nobel in Economics |
| Arthur Samuel | Self-learning checkers program, coined “machine learning” | 1952 | First working ML system |
| Frank Rosenblatt | The Perceptron, first trainable neural network | 1958 | Died 1971, before vindication |
| Joseph Weizenbaum | ELIZA, then became a leading AI critic | 1966 | Wrote Computer Power and Human Reason |
| Geoffrey Hinton | Backpropagation, AlexNet, deep learning | 1986, 2012 | Turing Award 2018; Nobel Prize in Physics 2024 |
| Yann LeCun | Convolutional neural networks, LeNet | 1989 | Turing Award 2018 |
| Yoshua Bengio | Deep learning theory, attention mechanisms | 1990s-2010s | Turing Award 2018 |
| Jürgen Schmidhuber & Sepp Hochreiter | LSTM networks | 1997 | Powered sequence models for two decades |
| Fei-Fei Li | Created the ImageNet dataset | 2009 | Enabled the deep learning breakthrough |
| Ian Goodfellow | Generative Adversarial Networks | 2014 | Started modern generative modelling |
| Ashish Vaswani and seven co-authors | “Attention Is All You Need,” the Transformer | 2017 | The architecture behind every modern LLM |
| Demis Hassabis & John Jumper | AlphaGo, AlphaFold | 2016, 2020 | Nobel Prize in Chemistry 2024 |
Two things stand out when you lay it out like this.
Rosenblatt died in 1971, two years after Perceptrons was published and 41 years before AlexNet proved the approach worked. He never saw it.
And Hinton, LeCun and Bengio shared the 2018 Turing Award for work most of the field had written off in the 1990s. All three kept going through a period when neural networks were an unfashionable research choice. Hinton then won a Nobel in 2024 for the same body of work.
That is what conviction looks like when you compress it into a table.
1943 to 1955: The Foundations Before AI Had a Name
The foundations of AI were built by people who had no word for what they were doing. Between 1943 and 1955, researchers produced the artificial neuron, the stored-program computer, information theory, the Turing Test, and the first self-improving program, all before the field was formally named.
Warren McCulloch, a neurophysiologist, and Walter Pitts, a logician who had run away from home as a teenager and never earned a degree, published “A Logical Calculus of the Ideas Immanent in Nervous Activity” in 1943. Their claim was that a network of simple threshold units could compute any logical function. That is the seed of every neural network since.
Claude Shannon’s 1948 information theory gave the field a way to measure information itself. In 1950 he published a paper on programming a computer to play chess, which set the template for AI research for the next fifty years: pick a game, define the search space, and see how far you can push a machine.
Then Turing, in October 1950, published the paper that framed the entire debate. He predicted that by the year 2000 a machine would fool an average interrogator 30% of the time in a five-minute conversation. He was roughly right on timing and completely wrong about what would make it possible, since nothing in his paper anticipates a 175-billion-parameter statistical model trained on the open web.
Arthur Samuel started work on his checkers program at IBM in 1952. By 1955 it could improve its own play by learning from games it had played. Samuel later coined the term “machine learning,” and his program is the first real example of it.
The honest read on this era: the ideas were mostly right and the hardware was hopeless. The IBM 701 Samuel worked on had a few thousand words of memory. The theory was decades ahead of the silicon, and that gap is the root cause of everything that goes wrong in the next section.
1956 to 1973: The Golden Age and the First Big Promises
The period from 1956 to 1973 produced genuine breakthroughs alongside predictions that were spectacularly wrong. Symbolic AI, which represents knowledge as explicit rules and symbols, dominated the era, and it worked brilliantly on toy problems and collapsed on real ones.
The wins were real. Logic Theorist, built by Allen Newell, Herbert Simon and J. C. Shaw, proved 38 of the first 52 theorems in Russell and Whitehead’s Principia Mathematica, and found a shorter proof for one of them. John McCarthy created Lisp in 1958, which stayed the default AI language into the 1990s. In 1966, Joseph Weizenbaum built ELIZA at MIT, a pattern-matching script that imitated a psychotherapist.
ELIZA is worth pausing on. Weizenbaum built it to demonstrate how shallow machine conversation was. Instead, his own secretary asked him to leave the room so she could talk to it privately. That reaction has a name now, the ELIZA effect, and if you have ever watched someone thank a chatbot, you have seen it. We are still building products on top of that same human tendency, which is worth remembering when you read a breathless review of a new AI assistant.
Stanford’s Shakey the Robot, developed between 1966 and 1972, was the first mobile robot that could reason about its own actions rather than just execute commands.
Where it went wrong: the predictions
The technical progress was matched by predictions that were, to be blunt, indefensible.
In 1965 Herbert Simon stated that “machines will be capable, within twenty years, of doing any work a man can do.” In 1970 Marvin Minsky told Life magazine that “in from three to eight years we will have a machine with the general intelligence of an average human being.”
Neither happened. And when they did not happen, the funders noticed.
The technical wall was the combinatorial explosion. Symbolic AI works by searching through possible states. Add a few variables and the number of states explodes beyond any computer’s ability to search. A program that could solve a puzzle with five blocks could not solve one with fifty. Researchers called their test environments “microworlds,” which was honest, and the problem was that nothing scaled out of them.
1969: the Perceptrons book
In 1969 Marvin Minsky and Seymour Papert published Perceptrons, which proved that a single-layer perceptron cannot compute the XOR function. The mathematics was correct. The effect was that neural network research lost most of its funding for over a decade, because the book was widely read as proof that the whole approach was a dead end.
It was not. Multi-layer networks solve XOR fine. The field just did not have a good way to train them yet, and would not until backpropagation was popularised in 1986. Seventeen years of neural network research were lost to a correct result about the wrong thing.
That is a useful lesson for reading AI criticism today. A precise technical limitation is not the same as a fundamental ceiling, and the difference is usually the thing nobody has invented yet.
1974 to 1980: The First AI Winter
The first AI winter ran from roughly 1974 to 1980. It was triggered by the 1973 Lighthill Report in the UK and by DARPA funding cuts in the US, and it happened because AI researchers had promised general intelligence and delivered systems that only worked in laboratory conditions.
Sir James Lighthill, an applied mathematician, was commissioned by the British Science Research Council to assess the state of AI. His 1973 report was brutal. His central finding: “In no part of the field have the discoveries made so far produced the major impact that was then promised.”
Lighthill’s core argument was the combinatorial explosion. AI results that looked impressive in microworlds did not survive contact with real-world scale, and he saw no path from one to the other.
The consequences were immediate and severe:
- UK funding was gutted. AI research in Britain was cut back to a handful of universities. Edinburgh’s AI department, one of the largest in the world at the time, was hit hard.
- DARPA pulled back in 1973 and 1974. The 1973 Mansfield Amendment restricted DARPA appropriations to projects with direct military application, which excluded most open-ended AI research. DARPA’s frustration with Carnegie Mellon’s Speech Understanding Research programme had already soured the relationship.
- The words became toxic. Researchers rebranded their work as “informatics,” “knowledge engineering,” or “machine learning” to avoid the phrase artificial intelligence on funding applications.
Read that last point again, because it happened. The term “artificial intelligence” was so damaged by unmet promises that serious scientists avoided writing it down.
If you want a sense of scale for how quickly sentiment can flip: in 1970 Minsky was telling a mass-market magazine that human-level machine intelligence was three to eight years away. By 1974 you could not reliably get funded to work on it in Britain.
1980 to 1987: Expert Systems and the Second Boom
Expert systems ended the first AI winter by doing something the field had never managed before: making money. An expert system encodes a specialist’s knowledge as a large set of if-then rules, and in narrow domains like configuring computer hardware or diagnosing infections, that approach worked well enough to deploy commercially.
The flagship case was XCON, originally called R1, written in OPS5 by John P. McDermott at Carnegie Mellon in 1978 and put into production at Digital Equipment Corporation in 1980. XCON configured VAX computer orders, checking that every component a customer ordered would actually work together.
The numbers were real. By 1986 XCON had processed around 80,000 orders at 95 to 98% accuracy, and DEC estimated it was saving roughly $25 million a year in avoided errors, faster assembly, and fewer free replacement components.
Stanford’s MYCIN, developed from 1972, diagnosed bacterial infections and recommended antibiotic dosages. In evaluations it performed at or above the level of junior physicians. It was never deployed clinically, partly because of liability questions nobody had answers for in 1976, which is a debate the medical AI field is still having in 2026.
The Fifth Generation and the funding race
In 1982 Japan’s Ministry of International Trade and Industry launched the Fifth Generation Computer Systems project, a ten-year national programme to build machines based on massively parallel computing and logic programming. MITI committed roughly ¥57 billion, about US$320 million at contemporary exchange rates.
The West panicked. The US responded with DARPA’s Strategic Computing Initiative. The UK launched the Alvey programme. Europe launched ESPRIT. A national AI project became a matter of industrial policy, and by the mid-1980s the AI industry was worth billions.
1986: backpropagation returns
While expert systems took the money and the headlines, the quiet fix for the 1969 Perceptrons problem arrived. In 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams published a paper that popularised backpropagation, an efficient way to train multi-layer neural networks by propagating errors backwards through the layers.
Almost nobody outside the connectionist research community cared at the time. Backpropagation is the training algorithm behind essentially every neural network built since, including the one behind ChatGPT. It sat mostly unexploited for 26 years because the hardware and the data to make it pay off did not exist yet.
Hinton eventually shared the 2024 Nobel Prize in Physics for this line of work. He waited nearly four decades for that.
1987 to 1993: The Second AI Winter
The second AI winter began in 1987 with the sudden collapse of the specialised AI hardware market, and it happened for a boring commercial reason: general-purpose workstations from Apple, IBM and Sun got cheap enough and fast enough to beat dedicated Lisp machines.
A Lisp machine could cost more than $100,000. When a Sun workstation running Lisp software matched it for a fraction of the price, an entire hardware category disappeared inside about a year. Lisp Machines Inc. shut down in 1987, and Symbolics filed for Chapter 11 bankruptcy in 1993.
The software side failed for a different reason. Expert systems worked in narrow domains and turned out to be brutally expensive to maintain. Every rule change risked breaking other rules. Nobody could keep a 10,000-rule knowledge base current as the underlying business changed. Even XCON, the poster child, became a maintenance burden.
Funding followed. In 1987 Jack Schwartz, then heading DARPA’s Information Science and Technology Office, cut the Strategic Computing Initiative’s AI budget hard, dismissing expert systems as clever programming rather than intelligence.
Japan’s Fifth Generation project ended in 1992 having produced real research but not the commercial systems that justified the spend.
Here is what I find most useful about this winter: nothing about the underlying science had been disproven. Backpropagation worked in 1986 exactly as well as it worked in 2012. What was missing was data and compute. The field spent six years in disgrace waiting for hardware that had not been built yet.
1993 to 2011: Quiet Progress and Public Wins
Between 1993 and 2011 AI stopped promising general intelligence and started shipping narrow systems that worked. This period produced Deep Blue, LSTM, the statistical machine learning revolution, ImageNet, and IBM Watson, and most of it happened without the word “AI” being used in marketing.
Deep Blue beat Garry Kasparov on 11 May 1997, the first time a computer defeated a reigning world chess champion in match play. It was mostly brute-force search running on custom hardware, not learning, and Kasparov’s reaction was to accuse IBM of cheating. But the symbolic weight was enormous.
The same year, Sepp Hochreiter and Jürgen Schmidhuber introduced the Long Short-Term Memory (LSTM) network, which solved the vanishing gradient problem that had made it impossible for recurrent networks to remember anything over long sequences. LSTM powered Google Translate and speech recognition for years and was the state of the art in sequence modelling until the Transformer arrived in 2017.
The bigger shift was philosophical. AI research moved from hand-coded rules to statistical methods that learn from data. Support vector machines, random forests, and probabilistic graphical models delivered results in spam filtering, credit scoring, recommendation engines, and search ranking. Google’s PageRank is an AI system that nobody called AI.
2009: ImageNet, the boring milestone that mattered most
In 2009 Fei-Fei Li’s team at Princeton and later Stanford released ImageNet, a labelled dataset of over 14 million images across thousands of categories, largely labelled through Amazon Mechanical Turk.
It is the least glamorous entry on this timeline and arguably the most consequential. Every algorithmic idea needed for the 2012 breakthrough already existed. What did not exist was a dataset large enough to prove it. Li’s insight was that the bottleneck was data, not algorithms, and she was right.
IBM Watson won Jeopardy! against champions Ken Jennings and Brad Rutter on 16 February 2011, demonstrating natural language question answering to a prime-time audience. IBM then spent a decade struggling to turn it into a business, which is its own lesson about the gap between a demo and a product.
2012 to 2017: The Deep Learning Revolution
The deep learning era began on 30 September 2012, when AlexNet won the ImageNet challenge with a top-5 error rate of 15.3%, roughly 10.8 percentage points better than the runner-up. It proved that deep convolutional neural networks trained on GPUs could beat every hand-engineered computer vision method at once.
AlexNet was built by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton. The three ingredients that made it work were the algorithm from 1986, the dataset from 2009, and consumer gaming GPUs. That is it. No new theory. The 26-year gap between backpropagation and AlexNet is the clearest evidence in this entire timeline that AI progress is usually gated by compute and data rather than ideas.
What followed came fast:
- 2014: Generative Adversarial Networks. Ian Goodfellow and colleagues introduced GANs, where two networks compete, one generating and one judging. This is the direct ancestor of AI image generation.
- 2014: Sequence-to-sequence learning made neural machine translation viable, and Google Translate’s quality jumped visibly.
- 2015: ResNet made networks hundreds of layers deep trainable and passed human-level accuracy on ImageNet classification.
- March 2016: AlphaGo defeats Lee Sedol 4-1. Go has more legal board positions than atoms in the observable universe, so brute-force search was never an option. DeepMind’s system combined deep neural networks with Monte Carlo tree search, and move 37 in game two was described by professional players as something no human would play. It won.
- June 2017: “Attention Is All You Need.” Eight researchers at Google introduced the Transformer, which dropped recurrence entirely in favour of a self-attention mechanism. Transformers parallelise across a sequence, which means they scale with hardware in a way LSTMs never could.
That last one is the single most important architectural milestone in the modern history of AI. Every model you use today, GPT, Claude, Gemini, Llama, is a Transformer or a descendant of one. If you want the mechanics of how these models actually process text, our guide on what tokens are in AI breaks it down without the maths.
2018 to 2022: The Language Model Era and the History of Generative AI
The history of generative AI as a mass phenomenon starts in 2018, when OpenAI’s GPT-1 and Google’s BERT proved that pretraining a Transformer on huge amounts of unlabelled text, then fine-tuning it, beat every task-specific model. Scaling that recipe up produced GPT-3 in 2020 and ChatGPT in 2022.
The scaling story in one column of numbers:
| Model | Year | Parameters | What changed |
|---|---|---|---|
| GPT-1 | June 2018 | 117 million | Pretrain then fine-tune works |
| BERT | October 2018 | 340 million | Bidirectional pretraining reshapes NLP benchmarks |
| GPT-2 | February 2019 | 1.5 billion | Fluent long-form text; staged release over misuse fears |
| GPT-3 | May 2020 | 175 billion | Few-shot learning without fine-tuning |
| GPT-4 | March 2023 | Undisclosed | Multimodal input, major reasoning jump |
GPT-2’s release is a moment worth remembering. OpenAI initially withheld the full model, arguing it was too dangerous to publish because of the risk of mass-produced disinformation. That decision was mocked at the time as a publicity stunt. In hindsight it was the first serious public argument about AI release policy, and versions of that argument are now written into the EU AI Act.
GPT-3 in May 2020 was the moment the scaling hypothesis stopped being a theory. At 175 billion parameters it could perform tasks it was never trained on, given a few examples in the prompt. That capability, few-shot learning, is why prompting became a skill.
Then the generative wave broke across every medium:
- November 2020: AlphaFold 2 achieved near-experimental accuracy on protein structure prediction at CASP14, solving a 50-year-old problem in biology. Its creators shared the 2024 Nobel Prize in Chemistry.
- January 2021: DALL-E generated images from text prompts.
- June 2021: GitHub Copilot brought AI pair programming to millions of developers.
- April 2022: DALL-E 2 made generated images good enough to be commercially useful.
- July 2022: Midjourney opened to the public through Discord and put AI art in front of a mass audience.
- August 2022: Stable Diffusion was released with open weights, which meant anyone could run image generation on their own hardware. The open model movement starts here.
If you want to understand how these image and video systems actually work under the hood, we covered the mechanics in how AI creates images and videos.
30 November 2022: ChatGPT
OpenAI released ChatGPT as a research preview on 30 November 2022. It was not a new model. It was GPT-3.5 with a chat interface and reinforcement learning from human feedback, wrapped in a free web page anyone could use.
It reached 100 million users in about two months, the fastest consumer adoption in internet history at that point. TikTok took nine months. Instagram took more than two years.
The lesson I keep coming back to is that the breakthrough was the interface, not the model. GPT-3 had been available through an API since 2020 and almost nobody outside developer circles noticed. Put the same capability behind a text box with no signup friction and it changes the world in eight weeks. I have seen this pattern in the SaaS tools I review constantly: the product that wins is rarely the most capable one, it is the one with the shortest path from curiosity to first result.
2023 to 2026: The Mainstream AI Era
From 2023 onward the history of AI stops being a research story and becomes an industry story. Frontier labs began shipping competing models on a cycle measured in weeks, regulation arrived, and open-weight models closed most of the capability gap with closed frontier models.
2023: the frontier race opens
- February 2023: Microsoft launched Bing Chat, putting an OpenAI model into a search engine and starting the AI search race.
- February 2023: Meta released LLaMA to researchers, which leaked almost immediately and seeded the entire open-weight ecosystem.
- 14 March 2023: OpenAI shipped GPT-4, and on the same day Anthropic launched Claude.
- 21 March 2023: Google opened Bard to the public.
- July 2023: Llama 2 shipped with open weights and a commercial licence, which is the moment open models became a serious business option rather than a research curiosity.
- December 2023: Google announced Gemini, a natively multimodal family.
2024: reasoning, regulation, recognition
- February 2024: OpenAI previewed Sora, generating minute-long coherent video from text.
- March 2024: Anthropic shipped the Claude 3 family, and NVIDIA announced the Blackwell GPU platform.
- May 2024: GPT-4o unified real-time voice, vision and text in one model.
- 1 August 2024: The EU AI Act entered into force, the first broad legal framework for artificial intelligence anywhere.
- September 2024: OpenAI released the o1 models, trained to reason step by step before answering. This shifted the scaling story from training compute to inference compute.
- October 2024: The Nobel Prizes recognised AI twice in one week. Physics went to John Hopfield and Geoffrey Hinton for neural network foundations. Chemistry went in part to Demis Hassabis and John Jumper for AlphaFold.
2025: the cost floor drops
January 2025 brought DeepSeek R1, an open reasoning model that matched frontier quality at a fraction of the training cost and briefly wiped hundreds of billions off US tech valuations. Whatever you think of the reported numbers, it ended the assumption that frontier capability required a US hyperscaler budget.
August 2025 brought GPT-5, unifying fast responses and deep reasoning. November 2025 brought Gemini 3, which retook the top of several benchmarks. December 2025 brought GPT-5.2 as a direct response, three weeks later.
2026: release cycles compress to weeks
The current year is the fastest in the history of AI by release cadence. GPT-5.5 landed on 23 April 2026. Claude Opus 4.8 shipped on 28 May 2026, replacing Opus 4.7 at the same price. Claude Fable 5, the first publicly available Mythos-class model, arrived on 9 June 2026 and topped the Artificial Analysis Intelligence Index.
Between February and April 2026 alone, Anthropic, OpenAI and Google collectively released seven frontier models. The gap between “best available model” and “second best” now closes in days.
The 2026 Stanford AI Index puts numbers on where that leaves us:
- Generative AI reached 53% global population adoption within three years, faster than the personal computer or the internet.
- Global corporate AI investment hit $581.7 billion, up 130% year over year, with generative AI investment at $170.9 billion, up 404%.
- Performance on SWE-bench Verified, a real-world coding benchmark, rose from 60% to near 100% in a single year.
- As of March 2026, Anthropic, xAI, Google, OpenAI, Alibaba and DeepSeek were all clustered within 25 Elo points on the Arena leaderboard.
- AI incidents rose 55%, from 233 in 2024 to 362.
That fifth bullet is the one most timelines skip. If you want the full adoption picture, we maintain a separate breakdown of AI adoption statistics with sourcing on each figure.
How Fast Is AI Moving Right Now? Live Data From Our Own Tracker
AI is currently producing about 66 significant news stories per day across major research, industry and community sources. That figure comes from our own monitoring pipeline, which logged 2,234 unique AI stories in 34 days between 4 July and 6 August 2026 across 20 active sources.
I built this because I got tired of arguing about whether AI news volume was actually increasing or whether it just felt that way. Here is what the data says.
Where AI news actually comes from:
| Source | Unique stories (34 days) | Share |
|---|---|---|
| arXiv cs.AI | 1,261 | 56.4% |
| TechCrunch AI | 199 | 8.9% |
| ZDNet AI | 103 | 4.6% |
| The Verge AI | 84 | 3.8% |
| arXiv cs.LG | 79 | 3.5% |
| Reddit r/artificial | 71 | 3.2% |
| MarkTechPost | 68 | 3.0% |
| arXiv cs.CL | 68 | 3.0% |
| OpenAI blog | 44 | 2.0% |
| Wired AI | 39 | 1.7% |
| MIT Technology Review | 36 | 1.6% |
What the stories are about:
| Category | Unique stories | Share |
|---|---|---|
| Research | 1,381 | 61.8% |
| General industry | 437 | 19.6% |
| Model releases | 156 | 7.0% |
| Tools | 83 | 3.7% |
| Hardware | 76 | 3.4% |
| Policy | 66 | 3.0% |
| Funding | 35 | 1.6% |
Two things jump out.
First, 62% of AI news is research papers, and arXiv alone accounts for 56% of everything our pipeline sees. What reaches you through tech media is the thin top layer. The 826 non-arXiv stories in that window are what most people would call “AI news,” which is about 24 a day. The rest is the field talking to itself.
Second, model releases are only 7% of the volume. The stories that dominate your feed and your group chats are a small slice of what is actually happening. If your mental model of AI progress is “a new model dropped,” you are tracking the least representative 7%.
For the record, the company mention leaderboard over the same window ran OpenAI 21, Google 13, Anthropic 8, Meta 7, NVIDIA 6, Microsoft 4. Six companies account for the overwhelming majority of named-entity coverage, which tells you something about concentration that no timeline entry can.
If you want to follow this daily rather than read about it, I put together a full breakdown of the best AI news sites, sources and subreddits covering exactly which of these feeds are worth your time and which are noise.
AI Milestone Timelines by Field
The general history of AI hides the fact that each field adopted it on a completely different schedule. Games got there first, medicine started early and moved slowest, and art went from research curiosity to mass consumer product in under four years.
History of AI in games and competition
Games were AI’s proving ground because they have clear rules and an unambiguous score.
| Year | Milestone |
|---|---|
| 1952 | Arthur Samuel’s checkers program learns from self-play |
| 1994 | Chinook wins the world checkers championship |
| 1997 | Deep Blue defeats Garry Kasparov at chess |
| 2011 | IBM Watson wins Jeopardy! |
| 2016 | AlphaGo beats Lee Sedol at Go, 4-1 |
| 2017 | AlphaZero learns chess, shogi and Go from self-play alone |
| 2019 | AlphaStar reaches Grandmaster level at StarCraft II |
AlphaZero in 2017 is the underrated one. Deep Blue needed a library of human grandmaster games. AlphaZero was given only the rules and taught itself to superhuman level in hours. That is the shift from encoding human knowledge to generating it.
History of AI in healthcare
Medicine has the longest gap between capability and deployment of any field on this list.
| Year | Milestone |
|---|---|
| 1972 | MYCIN begins at Stanford, diagnosing bacterial infections |
| 1976 | MYCIN outperforms junior physicians in evaluation, never deployed |
| 1980s | Expert systems for diagnosis proliferate, then fail on maintenance |
| 2018 | FDA clears IDx-DR, the first autonomous AI diagnostic device |
| 2020 | AlphaFold 2 solves protein structure prediction at CASP14 |
| 2024 | Nobel Prize in Chemistry awarded in part for AlphaFold |
MYCIN was good enough in 1976 and still never treated a patient. The blockers were liability and integration, not accuracy. Forty-two years passed between MYCIN’s evaluation results and the FDA clearing the first autonomous diagnostic AI. If you are wondering whether the medical AI you read about this month will be in your clinic next year, that gap is your baseline.
When was AI art invented?
AI art in its modern form dates to 2014, when Generative Adversarial Networks made machine-generated images convincing. Earlier work exists, including Harold Cohen’s AARON drawing program from the 1970s, but the current wave starts with GANs and became a consumer product between 2021 and 2022.
| Year | Milestone |
|---|---|
| 1973 | Harold Cohen begins AARON, an autonomous drawing program |
| 2014 | Generative Adversarial Networks introduced |
| 2015 | Google’s DeepDream makes generated imagery go viral |
| 2021 | DALL-E generates images from text prompts |
| 2022 | DALL-E 2, Midjourney public release, Stable Diffusion open weights |
| 2024 | Sora generates minute-long coherent video from text |
Eight years from research paper to Discord bot that anyone could use. That is the fastest research-to-mass-adoption cycle anywhere in this timeline.
History of AI in education and work
This is the most recent and the fastest-moving branch. ChatGPT launched on 30 November 2022 and by the 2026 Stanford AI Index, four in five university students were using generative AI. Meanwhile AI skills now appear in 2.5% of all US job postings, up 55% in a year and 297% over a decade, and mentions of agentic AI skills jumped from 0.06% of postings in 2024 to 0.23% in 2025.
No previous technology on this timeline moved from launch to majority student adoption in under four years. If you are trying to work out which roles that pressure lands on first, we covered the evidence in will AI replace software engineers and what jobs are safe from AI.
What 80 Years of AI History Actually Teaches You
The history of AI shows four patterns that repeat: capability is gated by compute and data rather than ideas, breakthroughs are usually old ideas meeting new hardware, hype cycles overshoot in both directions, and the interface matters more than the model.
Pattern 1: the ideas usually arrive decades early. Backpropagation was published in 1986 and did not pay off until 2012. Neural networks were modelled in 1943. Reinforcement learning ideas from the 1980s powered AlphaGo in 2016. When someone tells you a technique is a dead end, check whether they mean the theory is wrong or the hardware is not ready. Those are very different claims.
Pattern 2: winters follow overpromising, not underdelivering. Both AI winters came after specific, dated, public predictions failed. Simon’s twenty-year prediction in 1965. Minsky’s three-to-eight-year prediction in 1970. The Fifth Generation project’s ten-year plan. The technology in each case kept improving. What collapsed was credibility.
This one is personal for me. I got scammed repeatedly as a teenager chasing online money-making schemes, burning through about $300 of my savings on fake pay-per-click sites and paid-to-click programmes. What I learned the expensive way is that the pitch and the product are separate things, and the louder the pitch, the more carefully you check the product. Every AI winter in this timeline is that same lesson at national scale.
Pattern 3: the boring infrastructure milestone beats the flashy demo. ImageNet is a labelled dataset. It is the least exciting entry in the entire timeline and it enabled everything after 2012. Lisp, published in 1958, outlasted every AI system built with it. Transformers are an architecture paper, not a product.
Pattern 4: distribution beats capability. GPT-3 was available for two and a half years before ChatGPT. Same underlying capability, almost no cultural impact. Add a chat box and remove the signup friction and you get 100 million users in eight weeks.
Are we in another AI bubble?
Honestly, I do not know, and anyone who tells you they do with confidence is selling something.
What I can tell you is what is different this time. In both previous winters, AI had no significant paying commercial market outside a handful of expert system deployments. Today generative AI has reached 53% global adoption in three years and $581.7 billion in annual corporate investment. That is a real revenue base, not a research grant.
What is the same: predictions with dates attached. Every claim about AGI arriving by a specific year is structurally identical to Simon in 1965 and Minsky in 1970. Those men were not fools. They were among the smartest people in computing, and they were wrong by decades because they underestimated how hard the last 10% is.
My working assumption is that the technology keeps improving and the expectations correct. That is not a winter. It is a repricing, and it is survivable if you did not build your business on the most aggressive version of the forecast.
Frequently Asked Questions About the History of AI
When was AI invented?
AI was formally invented in 1956 at the Dartmouth Summer Research Project, where John McCarthy coined the term “artificial intelligence.” The technical groundwork came earlier: McCulloch and Pitts modelled the artificial neuron in 1943, and Alan Turing published his test for machine intelligence in 1950.
Who is the father of AI?
John McCarthy is most commonly called the father of AI. He coined the term “artificial intelligence” in 1955, organised the 1956 Dartmouth workshop, and created the Lisp programming language in 1958. Alan Turing is often called the father of computer science and theoretical AI, and Marvin Minsky is credited as a co-founder of the field.
What is an AI winter?
An AI winter is a period of reduced funding, research activity and public interest in artificial intelligence following a collapse in confidence. There have been two: 1974 to 1980, triggered by the Lighthill Report and DARPA cuts, and 1987 to 1993, triggered by the collapse of the Lisp machine market and disappointing expert systems.
When was generative AI invented?
Modern generative AI traces to 2014, when Ian Goodfellow introduced Generative Adversarial Networks. The Transformer architecture in 2017 and GPT-1 in 2018 made text generation practical, GPT-3 in 2020 made it useful without fine-tuning, and ChatGPT in November 2022 made it mainstream.
What was the first AI chatbot?
ELIZA, built by Joseph Weizenbaum at MIT in 1966, was the first widely known chatbot. It used simple pattern matching to imitate a psychotherapist, and users formed emotional attachments to it despite Weizenbaum designing it to demonstrate how shallow machine conversation was.
What is the most important milestone in AI history?
The 2017 Transformer paper, “Attention Is All You Need,” is arguably the most consequential technical milestone because every modern large language model is built on it. For cultural impact, ChatGPT’s launch on 30 November 2022 changed public understanding of AI more than any prior event.
The Honest Takeaway
Eighty-three years separate McCulloch and Pitts from Claude Fable 5. In that time the field has been declared dead twice, renamed to avoid stigma, and rescued each time by hardware catching up to ideas that were already sitting in published papers.
Three things are worth carrying out of this timeline.
The gap between an idea and its payoff is usually measured in decades, not quarters. Backpropagation waited 26 years. If you are evaluating an AI tool today and the vendor claims a research result from six months ago is already production-ready, that is the claim to test hardest.
Both winters were credibility failures, not technology failures. The science never stopped working. What broke was the distance between what was promised and what shipped. That is exactly the risk profile of the current market, and it is why I test tools with my own money before recommending them.
Nothing here happened in a straight line. The most important entry on this timeline is a labelled image dataset released in 2009 by a researcher who thought everyone else was solving the wrong problem. The next one probably looks equally unglamorous right now.
If you want the practical version of all this, the part where history turns into what you should actually buy, start with our AI tool reviews, where every verdict comes from a tool I paid for and used. And if you would rather keep an eye on where the field goes from here, the AI guides hub covers how the current generation of systems actually works.
I hope this timeline was useful. If I have missed a milestone you think belongs here, or got a date wrong, tell me and I will fix it. Cheers.