History of Generative AI: From 1943 to ChatGPT and Beyond

10/1/2026

By: Devessence Inc

From_1943_to_ChatGPT_cover.webp

Generative AI's real starting point is 1943, not 2022. That year, Warren McCulloch and Walter Pitts published the first mathematical model of an artificial neuron, decades before a computer existed that could run it. Getting from that paper to ChatGPT took 79 years, two AI winters, and a GPU company that nearly got talked out of the bet that made it the most valuable hardware maker on earth.

This timeline draws on a recent talk by Shaun Walker, founder of Devessence and a 17-year Microsoft MVP who also created the Oqtane and DotNetNuke frameworks. In our new article, we trace how generative AI actually developed, decade by decade, and close with where the token economy, infrastructure spending, and training-data supply are headed next.

Key Takeaways

  • Generative AI runs on machine learning, the bottom-up, non-deterministic branch of AI, not symbolic AI, the older rule-based branch still used in high-compliance systems.
  • The field's first mathematical model of a neuron was published in 1943, but it took until 2017's transformer architecture to produce the mechanism that actually powers today's large language models.
  • ChatGPT reached 100 million active users within two months of its November 2022 launch, the fastest-growing consumer application in history at that point.
  • Nvidia's parallel-computing bet, made years before deep learning existed as a business category, is why GPUs became the default hardware for training modern AI models.
  • Two AI winters, roughly 1970 to 1980 and a slower stretch through the late 1980s, show that generative AI's current momentum isn't guaranteed to be permanent or linear.
  • Goldman Sachs projects that AI-driven token usage will overtake human-driven usage by 2028, reshaping how much computing infrastructure the industry is racing to build.

From_1943_to_ChatGP_79_Years_to_ChatGPT.webp

Two Branches of AI, and Why Only One Powers ChatGPT

AI splits into two distinct traditions, and only one of them is what generative AI actually runs on.

Symbolic AI is rule-based and explainable

Symbolic AI is a top-down approach built on transparent, human-authored rules. It's deterministic, meaning the same input always produces the same output, which makes it well-suited to high-compliance environments that need to explain exactly why a system produced a given answer. It's a mature branch of AI, but it isn't the branch this article follows.

Machine learning is bottom-up and probabilistic

Machine learning works in the opposite direction: it's non-deterministic, so the same query can produce different results on different runs, and it operates closer to a black box. It requires enormous amounts of training data to build a model. In exchange, it generalizes well and handles perception tasks like vision, speech, and translation, exactly the territory generative AI lives in.

1943 to 1969: The Neural Network's Improbable Origins

Long before anyone could run one on a computer, the neural network started as pure mathematics.

The first artificial neuron existed only on paper

In 1943, McCulloch and Pitts at the University of Illinois at Chicago published "A Logical Calculus of the Ideas Immanent in Nervous Activity," the first mathematical model of an artificial neural network. McCulloch, a psychologist studying how the brain worked, described a function that could take binary inputs and perform logic on them, with no limit on how many inputs it handled. It wasn't implemented on any machine, since none capable of running it existed yet, and critically, it couldn't learn.

Alan Turing entered the story in 1950 with the Turing Test, still used today as a benchmark for machine intelligence. He'd already earned his reputation cracking the German Enigma machine during World War II.

The perceptron turned theory into a working machine

Frank Rosenblatt at Cornell changed that in 1957. His perceptron, implemented on an IBM 704, introduced weights and bias and could handle any type of input rather than just binary values. It was the first single-layer neural network capable of teaching itself to distinguish between categories, demonstrated live using punch cards.

The rest of the 1960s filled in the algorithmic toolkit that still underpins machine learning. John McCarthy at MIT coined the term "artificial intelligence" and designed Lisp in 1958, and Bernard Widrow and Ted Hoff at Stanford developed ADALINE and the Least Mean Squares algorithm in 1959, still used in machine learning today. Henry Kelley introduced backpropagation in 1960, and Joseph Weizenbaum built ELIZA at MIT in 1961, an early chatbot that gave rise to the "ELIZA effect": people attributing genuine feelings to a machine that has none, still visible in how people talk about ChatGPT today.

A critical book triggered the first AI winter

By 1969, momentum stalled. Marvin Minsky and Seymour Papert at MIT, both committed to symbolic AI, published Perceptrons, a book highlighting the XOR problem: single-layer networks using linear classifiers would need an effectively infinite number of neurons to solve certain problems. Many researchers read this as proof that neural networks were a dead end, and funding shifted almost entirely to symbolic AI. Minsky and Papert are widely credited with triggering what became known as the first AI winter.

1970 to 1989: The AI Winter and the Researchers Who Kept Going

Roughly 1970 to 1980 saw little research funding or activity in neural networks, deepened further in 1973 by the UK's Lighthill Report, a pessimistic government assessment of AI's progress. A small number of researchers kept the field alive anyway.

Two 1980 breakthroughs revived interest mid-winter

In 1980, Kunihiko Fukushima published the Cognitron, the first convolutional neural network architecture, while John Hopfield at Princeton created the Hopfield network, a recurrent design with content-addressable memory. Between the two, funding and research attention began returning to neural networks through the early 1980s. Rina Dechter coined the term "deep learning" itself in a 1986 paper.

Hinton, Bengio, and LeCun kept building through the downturn

Geoffrey Hinton and Yoshua Bengio at the University of Toronto and the Canadian Institute for Advanced Research kept producing foundational work on neural networks through the late 1980s, even while mainstream research favored symbolic AI.

Yann LeCun, at Bell Labs, developed LeNet during this period, later used by the U.S. Postal Service to read handwritten zip codes and by NCR to read digits on checks. These were genuinely practical deployments, though the companies involved were reluctant to share the underlying techniques, which slowed broader adoption.

George Cybenko's 1989 Universal Approximation Theorem showed a single hidden layer could theoretically approximate any function. The research community widely misread it as discouraging deeper, multi-layer architectures, a misunderstanding that arguably delayed real progress on deep networks by roughly two decades.

The 1990s: How Nvidia's Gaming Chips Became AI's Engine

From_1943_to_ChatGPT_From_Gamin_Cards_to_Ais_Engine.webp

Nvidia was founded in 1993 by Jensen Huang, Curtis Priem, and Chris Malachowski, all previously at Sun Microsystems, with a mission to build more powerful video cards. Nothing about that mission mentioned artificial intelligence.

Deep Blue and LSTM marked a symbolic milestone in a machine learning decade

In 1997, IBM's Deep Blue defeated reigning chess champion Garry Kasparov, running on specialized chess chips that could evaluate 200 million positions per second. It wasn't machine learning in the modern sense, but it publicly demonstrated that a machine could out-calculate a human expert. That same year, Jürgen Schmidhuber and Sepp Hochreiter invented Long Short-Term Memory (LSTM), a recurrent neural network design able to retain information across long sequences.

Nvidia coined "GPU" without knowing what it would become

In 1998, Nvidia released the GeForce 256, a card popular with gamers that introduced early parallel-computing capability, and coined the marketing term "GPU," graphics processing unit, along the way. Nobody at the company was thinking about neural networks yet. That would change within the next decade.

The 2000s: GPUs Enter Machine Learning

In 2005, Nvidia's Ian Buck, John Nickolls, and Bill Dally invented CUDA, a parallel computing platform letting developers program directly against GPUs. The same team accurately predicted that Dennard scaling, the trend that let chipmakers keep shrinking transistors, would hit a physical wall by the late 2000s, forcing workloads toward parallel architectures like GPUs. It happened roughly on schedule, and it caught CPU-focused companies like Intel and AMD somewhat off guard.

In 2009, Fei-Fei Li at Stanford released ImageNet, a freely available dataset of 14 million labeled images, at a time when nothing comparable existed for computer vision research. She also launched the ImageNet Large-Scale Visual Recognition Challenge, an annual public contest that would go on to define the next few years of the field.

This is a natural point to define training versus inference, two terms that come up constantly in this history. Training takes a large volume of data and categorizes it by attributes, work historically done by humans and increasingly automated over time.

Inference happens afterward, when new data is fed through an already trained model to classify or predict something about it; GPUs historically handle training, while inference often runs on high-powered CPUs.

The 2010s: The Deep Learning Boom That Built Modern AI

DeepMind was founded in London in 2010 by Demis Hassabis, Mustafa Suleyman, and Shane Legg, initially research-focused rather than aimed at near-term products.

AlexNet proved neural networks could win

2012 was the pivotal year. Alex Krizhevsky and Ilya Sutskever, working under Geoffrey Hinton at the University of Toronto, entered the ImageNet competition with what became known as AlexNet and achieved a 16.4% error rate, a dramatic improvement over prior approaches that weren't using neural networks. Krizhevsky drove the machine learning techniques while Sutskever optimized GPU processing through CUDA, proving neural networks could substantially outperform other image-recognition methods.

The following year, activist investor Starboard Value took a stake in Nvidia and pushed the board to scale back its investment in parallel computing and CUDA, which wasn't yet generating meaningful revenue. CEO Jensen Huang resisted, and Nvidia fended off the pressure, a decision that mattered enormously for the entire machine learning field in hindsight.

That same year, Nvidia engineers Bryan Catanzaro and Philippe Vandermersch invented cuDNN, a GPU-accelerated library that helped convince Huang that AI represented a once-in-a-lifetime opportunity.

OpenAI, AlphaGo, and the transformer arrived within three years

OpenAI was founded in 2015 by Elon Musk, Sam Altman, Ilya Sutskever, Greg Brockman, and Andrej Karpathy, incorporated as a nonprofit with a stated mission centered on researching artificial general intelligence. In 2016, DeepMind's AlphaGo defeated Go champion Lee Sedol using neural networks combined with Monte Carlo tree search, while Microsoft's Tay chatbot on Twitter had to be taken offline within 16 hours after being manipulated into posting offensive content.

Then, in 2017, Jakob Uszkoreit and Ashish Vaswani at Google Brain published "Attention Is All You Need," introducing the transformer, the architecture underlying essentially every large language model built since. A transformer works with tokens, embeddings, and probability: given a phrase like "a long time ago," it predicts the most likely next word based on patterns learned from massive training data, generating text one token at a time. Its output is entirely a function of learned probabilities, refined further each time a prediction is corrected.

Google let the paper be published openly, apparently without fully recognizing how significant it would become. OpenAI built on that transformer architecture in 2018 to create GPT-1. Microsoft made a billion-dollar investment in OpenAI in 2019, delivered largely as Azure compute capacity, the same year OpenAI released GPT-2.

2020 to 2022: GPT-3, Anthropic, and the ChatGPT Launch

GPT-3 arrived in 2020 with 175 billion training parameters, a dramatic jump from its predecessors, while Google DeepMind's AlphaFold solved the long-standing challenge of predicting protein 3D structures.

Anthropic split off from OpenAI in 2021

Dario Amodei and Daniela Amodei, siblings who had both worked at OpenAI, founded Anthropic in 2021 after disagreements over how OpenAI was approaching the risks of increasingly capable models. The same year, OpenAI announced the first version of DALL-E for text-to-image generation, and Microsoft launched GitHub Copilot as a technical preview.

ChatGPT became the fastest-growing consumer app in history

From_1943_to_ChatGPT_The_fastest_growing.webp

OpenAI released ChatGPT in November 2022, the first fully productized, publicly available version of the GPT platform it had spent years building. Adoption was immediate: a million users within five days and 100 million active users within two months, the fastest growth any consumer application had ever recorded. From that point forward, the only open questions were how fast generative AI would spread and in what direction.

2023 to 2026: The Generative AI Explosion

The years since ChatGPT's launch have moved fast enough that keeping track of them has become a job in itself.

2023 brought GPT-4, boardroom drama, and an open letter that didn't stop anything

OpenAI released GPT-4, and CEO Sam Altman was briefly fired by the board before being rehired days later. Meta released Llama as an open-source model, xAI released Grok, Google renamed Bard to Gemini, Mistral AI launched in France, and Anysphere released Cursor. The Future of Life Institute lobbied for a six-month pause on AI development over safety concerns; the pause never happened, and development accelerated instead.

2024 and 2025 brought multimodal models, regulation, and a wake-up call from China

OpenAI released GPT-4o, its first true multimodal model handling text, audio, and video, while Microsoft acquired Inflection AI and the European Union adopted the AI Act. In 2025, China's DeepSeek model, released as open source, proved far more capable than many expected, a clear signal that Chinese labs were keeping pace globally. OpenAI converted from a nonprofit into a public benefit corporation, Anthropic released Opus 4.5 late in the year, and Nvidia paid roughly $20 billion for AI chip startup Groq's inference technology.

2026 has been the year of agents, IPOs, and export controls

Early in the year, an autonomous agent that went through the names Claudebot, Moltbot, and Open Claw drew significant attention to AI systems that could run on a desktop and act on a user's behalf.

SpaceX absorbed xAI, then went public on the Nasdaq in June at a target valuation of $1.77 trillion, one of the largest public offerings in history, and acquired Cursor's parent company, Anysphere, days later. Microsoft and OpenAI's five-year exclusive partnership came to an end, and Nvidia introduced its Vera Rubin architecture as the successor to Blackwell.

Anthropic's Mythos model briefly had its distribution restricted by the U.S. Department of Commerce over export-control requirements before access was restored later in the year, a reminder that generative AI has become entangled with national policy, not just product roadmaps.

What Comes Next for Generative AI

A few trends are worth watching closely, based on where the infrastructure, capital, and data are actually pointing.

From_1943_to_ChatGPT_What_Comes_Next.webp

Token usage and infrastructure spending are both climbing

Goldman Sachs research projects that AI-driven token usage will overtake human-driven usage by 2028 and keep climbing well past that point, driven largely by agentic workloads rather than direct human prompting.

Hyperscaler infrastructure investment is expected to stay elevated for at least five more years, since demand for AI compute still outpaces supply. Energy providers and component makers, GPU, CPU, memory, storage, and networking vendors, stand out as the clearest beneficiaries of that spending.

Capital is consolidating around a small number of players

SpaceX went public earlier in 2026, Anthropic is expected to follow later this year, and OpenAI is reportedly targeting 2027, each a major capitalization event letting public markets invest directly in frontier AI labs.

That capital is likely to accelerate consolidation, as already seen in SpaceX's acquisition of Cursor, with larger players absorbing smaller ones. There's also growing scrutiny of how tightly interconnected the financing and investment relationships have become among OpenAI, Nvidia, Microsoft, Oracle, Amazon, and Meta.

Training data is running out faster than most people realize

Frontier models have now consumed most of the readily available digitized text and image data, a point some researchers call "peak data." As more content online is itself AI-generated and feeds back into future training runs, models risk stagnating on a diet of their own output.

There's a related risk too: as freely available data dries up, companies are increasingly looking to material users upload directly, and depending on a given service's terms, documents pasted into a chat interface for something as simple as a summary can end up contributing to that provider's training data.

Let's Talk
Teams weighing how much of their own proprietary data to expose to third-party AI tools face exactly this tradeoff. Let’s discuss where that line should sit for your platform.
Get in touch →

Final Thoughts

Every phase of this history looked, at the time, like either the beginning of something permanent or the end of a dead idea, and neither read turned out to be reliable. Two separate AI winters followed periods of real optimism. Two separate technical dead ends, the XOR problem and the misread Universal Approximation Theorem, each delayed progress by a decade or more before the field found its way back.

Expect the infrastructure story to matter as much as the model story over the next two years

Given how much capital is now tied up in data centers, GPU supply, and a small cluster of interdependent companies, the next 18 to 24 months are likely to be shaped as much by infrastructure economics as by any single model release.

Energy availability, chip supply, and data-center permitting are all live constraints now. The token-usage curve Goldman Sachs is tracking is worth watching as closely as benchmark scores, since it's a better early signal of where real deployment is heading.

Training-data scarcity will force a shift in what "state of the art" means

As peak data becomes a real constraint rather than a talking point, the next competitive edge in frontier models is likely to come less from finding more data and more from smarter training techniques, careful synthetic data generation, and proprietary data partnerships. The open internet's remaining supply of genuinely new training material is shrinking faster than most public conversation acknowledges.

Let's Talk
Ready to take the next step with AI?
Talk to our engineering team →

FAQs

  • The mathematical foundations go back to 1943, when Warren McCulloch and Walter Pitts published the first model of an artificial neuron, though the transformer architecture that powers today's large language models wasn't invented until 2017.

  • Symbolic AI is a top-down, rule-based approach that's deterministic and explainable, while machine learning is a bottom-up, non-deterministic approach that learns patterns from large amounts of training data rather than following explicit rules.

  • Two separate slowdowns, often called AI winters, followed periods when influential researchers or reports concluded neural networks had hit a fundamental limit, in 1969 and again through the 1970s, which pulled funding toward symbolic AI until later breakthroughs revived interest.

  • ChatGPT was the first widely accessible, free, productized version of the GPT platform, and it reached 100 million active users within two months of its November 2022 launch, a faster adoption curve than any consumer application before it.

  • Nvidia's early investment in parallel computing, originally built for graphics rendering, turned out to be exactly the kind of vector math that neural network training needed, and breakthroughs like CUDA and cuDNN made GPUs dramatically more efficient at training deep learning models than general-purpose CPUs.

Receive Notifications...



Free Whitepaper Yield prediction isn’t a modeling problem It’s a data architecture problem — and that’s where agtech ML initiatives stall. The reference architecture for .NET and Azure platforms. Get the whitepaper →