A Short History of Artificial Intelligence
From Turing's 1950 question to today's language models: the ideas, winters and breakthroughs.
Artificial intelligence is older than most people think. The field has a name, a founding meeting and seventy years of bold promises, long disappointments and sudden breakthroughs. This essay walks through that history in plain terms, from the first question about thinking machines to the language models of today.
At a glance
- 1950: Alan Turing asks whether machines can think and proposes the "imitation game."
- 1956: the Dartmouth workshop gives the field its name.
- 1970s and late 1980s: two "AI winters," when funding and interest collapsed.
- 2012: a deep neural network wins the ImageNet competition by a wide margin.
- 2017: the transformer architecture is published. It now underlies most large language models.
- 2024: Nobel Prizes in both physics and chemistry go to work built on AI.
The question: can machines think?
In 1950 the British mathematician Alan Turing published "Computing Machinery and Intelligence" in the journal Mind. Instead of arguing about what "thinking" means, he proposed a test. A judge holds typed conversations with a person and a machine. If the judge cannot reliably tell which is which, Turing suggested, we have little reason to deny that the machine thinks. He called it the imitation game. Most people now call it the Turing test.
Turing also predicted that by the end of the century, machines would play the game well enough to fool an average judge a good part of the time. He was early, but the question he asked still frames the field.
A name and a plan: Dartmouth, 1956
In 1955, John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon proposed a summer research project at Dartmouth College. Their proposal used the term "artificial intelligence" and set out a confident goal: every aspect of learning or intelligence could, in principle, be described precisely enough for a machine to simulate it. The workshop took place in the summer of 1956 and is usually counted as the birth of the field.
The early years produced real results. Programs proved theorems in logic, played checkers well and solved algebra problems. In 1958 Frank Rosenblatt at the Cornell Aeronautical Laboratory built the perceptron, a simple learning machine loosely inspired by neurons, the distant ancestor of today's neural networks.
The winters
The early optimism ran ahead of the hardware and the methods. In 1969 Minsky and Seymour Papert's book Perceptrons showed hard limits of single-layer networks, and interest in neural networks faded. In 1973 the Lighthill Report, written for the British Science Research Council, concluded that AI had failed to meet its promises, and British funding was cut sharply. American funding also tightened. This period is known as the first AI winter.
In the 1980s, "expert systems" brought a second boom. These programs stored the rules of human specialists, such as how to configure a computer order or diagnose an infection. Companies spent heavily on them. But expert systems were expensive to build, hard to update and brittle outside their narrow area. By the late 1980s the market for specialized AI hardware had collapsed, and a second winter followed.
Learning from data
Through the 1990s and 2000s, the field turned steadily toward machine learning: instead of writing rules by hand, let the program find patterns in data. Statistical methods improved speech recognition, spam filters and search engines. In 1997 IBM's Deep Blue defeated the world chess champion Garry Kasparov in a six-game match, mostly through specialized hardware and search rather than learning.
A small group of researchers kept working on neural networks with many layers, an approach that came to be called deep learning. For years it was out of fashion. Three things changed that: far more digital data, much faster graphics processors and better training methods.
The breakthrough: 2012
ImageNet was a large collection of labeled photographs used for an annual recognition contest. In 2012 a deep neural network called AlexNet, built by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton at the University of Toronto, won with a top-5 error rate of about 15 percent. The next-best entry scored about 26 percent. The gap was so large that most of the field switched to deep learning within a few years.
In 2016 DeepMind's AlphaGo beat Lee Sedol, one of the world's strongest Go players, four games to one. Go had long been considered too complex for computers to master. In 2019 Yoshua Bengio, Geoffrey Hinton and Yann LeCun received the 2018 Turing Award, computing's highest honor, for their work on deep learning.
Transformers and language models
In 2017 researchers at Google published "Attention Is All You Need," which introduced the transformer. It processes a whole sequence at once and learns which parts of the input matter most to each other. Transformers trained well on very large amounts of text and scaled well on modern chips.
Large language models are transformers trained to predict the next piece of text across a huge share of the written web, books and code. As they grew, they began to answer questions, write, translate, summarize and write software. OpenAI released ChatGPT to the public on November 30, 2022. It became one of the fastest-adopted consumer products in history and put AI into everyday conversation.
AI in science
In 2024 the Nobel Prize in Physics went to John Hopfield and Geoffrey Hinton for foundational work on artificial neural networks. Half of the Nobel Prize in Chemistry went to Demis Hassabis and John Jumper of Google DeepMind for AlphaFold, which predicts the shape of proteins from their sequence. The other half went to David Baker for computational protein design.
What is still unsolved
- Reliability. Language models can state false things with confidence.
- Reasoning. They are strong on many tests and still fail on some simple problems that people find easy.
- Cost and energy. Training and running the largest models takes large data centers and a lot of electricity.
- Alignment and safety. Making sure powerful systems do what people intend, and nothing harmful, is an open research field.
Each of these is the subject of active research and real investment. How fast they are solved shapes the most ambitious question in the field: whether machines will one day exceed human intelligence across the board. That idea, and the debate around it, is the subject of the next essay.
Sources
- A. M. Turing, "Computing Machinery and Intelligence," Mind 59, no. 236 (1950).
- J. McCarthy, M. L. Minsky, N. Rochester, C. E. Shannon, "A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence," August 31, 1955.
- F. Rosenblatt, "The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain," Psychological Review, 1958.
- M. Minsky and S. Papert, Perceptrons, MIT Press, 1969.
- J. Lighthill, "Artificial Intelligence: A General Survey," Science Research Council, 1973.
- A. Krizhevsky, I. Sutskever, G. E. Hinton, "ImageNet Classification with Deep Convolutional Neural Networks," NeurIPS 2012.
- D. Silver et al., "Mastering the game of Go with deep neural networks and tree search," Nature, 2016.
- Association for Computing Machinery, 2018 A.M. Turing Award announcement, March 2019.
- A. Vaswani et al., "Attention Is All You Need," NeurIPS 2017.
- OpenAI, "Introducing ChatGPT," November 30, 2022.
- The Nobel Prize in Physics 2024 and the Nobel Prize in Chemistry 2024, nobelprize.org.