
One of the world’s most respected AI researchers just said something that might surprise you: ChatGPT and tools like it can’t actually create anything truly new. Richard Sutton, winner of the Turing Award (the Nobel Prize of computing), explained exactly why. And what real AI creativity actually looks like.
The Gist
- Turing Award winner Richard Sutton says generative AI lacks the ability to evaluate its own outputs
- Real discovery requires three steps: generating ideas, testing them, and keeping what works. AI only handles step one.
- Systems like AlphaGo, AlphaFold, and Claude Code are examples of AI that can actually discover
What Generative AI Can and Cannot Do
When you ask ChatGPT to write a poem, draft an email, or explain a concept, it generates text by predicting what words should come next based on patterns it learned from billions of pages of human writing. This type of AI is called “generative AI.” It generates outputs. It’s incredibly useful, and it can produce things that look creative.
But Richard Sutton, one of the founding figures of modern AI research, says there’s a critical limitation hiding behind all that impressive output. The problem isn’t what generative AI can produce. The problem is that it has no way of knowing whether what it produced is actually good, novel, or useful. It can generate a thousand ideas, but it can’t evaluate a single one.
Sutton illustrates this with a researcher’s joke that lands harder than it sounds: “This work is both novel and good. Unfortunately, the parts that are good are not novel, and the parts that are novel are not good.” That’s the trap generative AI is stuck in. It can combine and remix, but it can’t tell the difference between a genuinely new idea and a sophisticated-sounding mistake.
This matters beyond academia. If you’ve ever used ChatGPT for a complex task and felt like the answer was polished but somehow… off, you’ve experienced this limitation directly. The AI wasn’t lying. It just had no way to check whether its output was actually correct or valuable. It didn’t even know it didn’t know.

The Three Steps of Real Discovery
Sutton defines genuine discovery as a three-step process. First, variation: generating different options or ideas. Second, evaluation: actually testing those ideas to see which ones work. Third, selective retention: keeping the approaches that passed the test and building on them. Science (and any real creative process) requires all three.
Standard generative AI handles step one well. It produces endless variations. But it has no built-in mechanism for steps two and three. There’s no test it can run to verify its outputs. “The novelty flickers into existence,” as Sutton puts it, “but if its value is unrecognized, it flickers away and is lost.”
The AI systems that can do all three steps look very different. AlphaGo (the AI that beat world champions at the board game Go) has a built-in evaluation mechanism: winning or losing the game. AlphaFold (the AI that cracked protein structure prediction, one of biology’s hardest problems) could test its predictions against known molecular structures. Recursive self-improvement systems operate on a similar principle. The AI improves because it can measure whether it improved.
Claude Code, Anthropic’s programming AI, is also on Sutton’s list of systems that can truly discover, because code either works or it doesn’t. That binary test is what separates a creative system from a generative one. The ability to fail, and to know you failed, is what enables real learning.
What This Means for How You Use AI
This doesn’t mean ChatGPT or similar tools are useless. Far from it. They’re exceptional for tasks where you, the human, provide the evaluation: editing your writing, brainstorming ideas you then filter, drafting messages you review before sending. In those workflows, you become the “evaluation” step that generative AI is missing.
Where this matters most is in trusting AI outputs blindly. If you ask an AI to do your research, check a legal document, or give you medical advice, you’re asking it to do something it structurally cannot do well: evaluate whether its answer is actually correct. The more the task requires being right, the more human oversight you need.
In the short term, Sutton’s critique is likely to influence how the industry talks about AI capabilities. Calling a chatbot “creative” or “intelligent” without qualification is something AI companies will face more scrutiny over, especially from researchers and regulators.
Over the next three to six months, watch for a growing split in the AI market: tools that are honestly positioned as “generative” assistants versus systems designed around feedback loops and verifiable outcomes. Sutton’s vision, centered on AI that interacts with its environment and adapts based on results, points toward the next generation of tools. The AI that will actually change the world won’t just generate. It will evaluate, fail, learn, and try again.
Follow the story on AI Noobies.



