Arthur Samuel’s checkers program at IBM learned from experience, including games against itself. Samuel popularised the term “machine learning” in 1959.
Archives
An AI system predicted the 3D shapes of nearly all known proteins.
DeepMind’s AlphaFold 2 made a leap in predicting how proteins fold, and its database grew to over 200 million predicted structures. Its creators shared the 2024 Nobel Prize in Chemistry.
The backpropagation paper that powers neural networks was barely four pages long.
The 1986 Nature paper by Rumelhart, Hinton and Williams, “Learning representations by back-propagating errors”, popularised the method used to train almost every neural network today. The core idea had appeared in earlier work too.
GPT-3 had 175 billion parameters.
Released in 2020, GPT-3’s 175 billion adjustable numbers made it more than 100 times larger than GPT-2. Each parameter is just a number nudged during training.
The 2012 neural network that changed computer vision was trained on two gaming graphics cards.
AlexNet won the 2012 ImageNet challenge by a large margin. It was trained on two NVIDIA GTX 580 GPUs — consumer gaming cards — helping start the move of AI onto graphics chips.
The image dataset that kick-started the deep-learning boom was labelled by crowd workers.
ImageNet contains over 14 million images hand-annotated with what they show, labelled largely through Amazon Mechanical Turk. The yearly ImageNet challenge became the benchmark that showed deep learning’s power.
Things that are easy for toddlers are hard for AI, and vice versa.
Moravec’s paradox: high-level reasoning like chess needs relatively little computation, while walking, grasping and recognising faces — which children do effortlessly — are extremely hard for machines.
Changing a single pixel can fool some image-recognition models.
Researchers showed that carefully chosen tiny changes — in some experiments just one pixel — can make an image classifier confidently mislabel a picture, for example calling a ship a car. These are called adversarial examples.
A horse called Clever Hans gave his name to a machine-learning problem.
In the early 1900s Hans appeared to do arithmetic by tapping his hoof. He was actually reading tiny, unconscious cues from people watching. When an AI model gets the right answers for the wrong reasons, researchers call it a “Clever Hans” effect.
Chatbots once struggled to count the r’s in “strawberry”.
Language models read tokens — chunks of text — rather than individual letters, so questions about spelling and letter counts are surprisingly hard for them. The strawberry question became a famous meme in 2024.
ChatGPT is estimated to have reached 100 million users in about two months.
Launched on 30 November 2022, ChatGPT was estimated by analysts to have hit 100 million monthly users by January 2023 — at the time described as the fastest-growing consumer app ever.
The “T” in ChatGPT stands for Transformer — from a 2017 paper about attention.
GPT means Generative Pre-trained Transformer. The transformer architecture was introduced by eight Google researchers in the 2017 paper “Attention Is All You Need”, now one of the most cited papers in computing.
AI research has gone through “winters”.
After early hype, funding and interest collapsed in the mid-1970s and again in the late 1980s, periods now called AI winters. Promises had run far ahead of what computers could actually do.
The A* pathfinding algorithm was invented for a wobbly robot named Shakey.
Shakey (SRI, 1966–1972) was the first mobile robot that could reason about its own actions. Its researchers developed the A* search algorithm, still used today in games, maps and robots.
AlphaGo played a move experts first thought was a mistake — and it won the game.
In game 2 against Lee Sedol in 2016, AlphaGo’s “move 37” was so unusual that commentators were puzzled; AlphaGo itself estimated a human would play it about 1 time in 10,000. It proved decisive.
A computer beat the world chess champion in a match in 1997.
IBM’s Deep Blue beat Garry Kasparov 3½–2½ in their 1997 rematch. It evaluated around 200 million positions per second using special chess chips — raw search, not learning.
Alan Turing’s famous test was originally called “the imitation game”.
In his 1950 paper “Computing Machinery and Intelligence”, Turing replaced the vague question “Can machines think?” with a game: can a machine’s typed answers be told apart from a person’s?
In 1958 a newspaper reported that a machine would one day walk, talk and be conscious.
After a US Navy press event about Frank Rosenblatt’s perceptron — a simple early neural network — The New York Times reported the Navy expected it to walk, talk, see, write, reproduce itself and be conscious of its existence. It could learn to tell simple shapes apart.
The phrase “artificial intelligence” was coined for a summer workshop.
John McCarthy and colleagues used the term in their 1955 proposal for a 1956 summer research project at Dartmouth College. The organisers hoped significant progress could be made in one summer. It took rather longer.
People confided in a 1960s chatbot that was only matching patterns.
Joseph Weizenbaum’s ELIZA (1966) mostly turned users’ sentences back into questions. Even so, some users became emotionally attached — Weizenbaum wrote that his secretary asked him to leave the room so she could talk to it privately. Treating a program as more understanding than it is became known as the “ELIZA effect”.
An 18th-century “chess-playing robot” was secretly a person hiding in a box.
The Mechanical Turk toured Europe from 1770 and beat players including, reportedly, Napoleon. A skilled human chess player was hidden inside the cabinet. Amazon’s crowd-work service is named after it.
The word “robot” comes from a Czech word for forced labour.
Karel Čapek’s 1920 play R.U.R. (Rossum’s Universal Robots) introduced “robot”, from the Czech “robota”, meaning drudgery or forced labour. His brother Josef suggested the word.
A ‘jiffy’ is a real unit of time.
In computing it is often the time between system timer ticks; in physics it has been defined as the time light takes to travel one centimetre — about 33.4 picoseconds.
The Indian rupee symbol ₹ was adopted in 2010.
Designed by D. Udaya Kumar, the symbol blends the Devanagari “र” and the Latin “R”. It was chosen through a public competition.
Pivot tables were popularised by Lotus Improv.
Lotus Improv (1991) introduced the idea of pivoting data. Excel added PivotTables in version 5.0 in 1993, the same release that brought VBA.
SAP’s S/4HANA “S” stands for “simple”.
S/4HANA, launched in 2015, is “Business Suite 4 SAP HANA”, with the S standing for Simple — a reference to its simplified data model on the HANA database.
PowerPoint was not invented by Microsoft.
It was created by Forethought Inc. and released for the Mac in 1987 as “PowerPoint”. Microsoft bought the company the same year.
Microsoft Word’s first version ran on DOS in 1983.
Word started as “Multi-Tool Word” for Xenix and MS-DOS. Early versions even supported a mouse, which was unusual for DOS programs.
Excel can nest up to 64 levels of functions.
Excel 2007 raised the limit from 7 to 64 nested levels — though a formula that deep usually belongs in LET or a helper column.
Excel’s maximum column is XFD.
Column 16,384 is labelled XFD. Press Ctrl+Right Arrow on an empty row and you will land there.