# Measuring Creativity in AI Agents: Why Novelty Alone Isn’t Enough for Scientific Discovery
The race to build AI systems capable of driving scientific breakthroughs has intensified dramatically in recent years. Reports of large language model agents proposing new algorithms, solving long-standing mathematical conjectures, and autonomously designing experiments have fueled enormous excitement and investment. Yet when we pit these agents against top-performing humans on real machine learning engineering challenges, a persistent gap remains. Understanding why requires looking beyond raw capability and examining how these systems think, explore, and ultimately deliver results.
## Redefining How We Think About AI Creativity
At the heart of the disconnect between agent promise and human-level performance lies a deceptively simple question: what does it mean for an AI to be creative? A widely accepted definition from creative psychology frames creativity as the simultaneous production of ideas that are both original and useful. An agent that generates wildly novel solutions that fail to perform well isn’t truly creative — and an agent that consistently produces solid but predictable solutions isn’t pushing boundaries either.
This definition becomes especially powerful when broken down further. Originality can be understood through two lenses. The first measures how novel an idea is relative to what the agent itself has already tried — this captures the internal exploration dynamics of the system. The second measures how novel an idea is compared to the entire body of human knowledge in a domain. On the usefulness side, researchers distinguish between the actual impact a solution delivers and whether that solution is practically feasible to implement.
Combining these dimensions gives us a four-part framework: originality relative to self-history, originality relative to human knowledge, impact on the task, and feasibility of execution. This framework turns creativity from an abstract concept into something that can be systematically measured and compared.
## Why Machine Learning Tasks Are the Ideal Testing Ground
Measuring creativity in AI systems is notoriously difficult. To do so meaningfully, three conditions must be met simultaneously. There needs to be a quantifiable way to assess how useful a solution is. There must be a rich set of human-produced solutions to compare against when measuring novelty. And the solution space must be large enough that genuinely new approaches are possible.
Machine learning engineering tasks — particularly those drawn from competitive data science platforms — satisfy all three criteria beautifully. These tasks have clearly defined performance metrics, thousands of publicly available human submissions that serve as reference points, and enormous combinatorial solution spaces where novel approaches regularly emerge.
The core research question becomes: given a fixed language model, how do different agent frameworks guide the generation of creative solutions over time, and can differences in creativity metrics explain why some frameworks or models outperform others?
## How Creativity Metrics Work in Practice
Each dimension of the creativity framework requires a specific measurement approach. When it comes to self-originality, automated scoring against all prior attempts in a given run provides a consistent scale. Cross-referencing solutions against large human corpora using retrieval and automated evaluation handles the human-knowledge novelty dimension. Impact is measured on a normalized scale showing how close a solution gets to the best human result. Feasibility is determined by whether the proposed code actually runs successfully.
A typical evaluation run unfolds across multiple episodes. In each episode, the agent proposes an approach, executes it, and receives feedback. A concrete example on an image classification task shows how this works in practice: an agent starts with a conventional method, then progressively experiments with different techniques. Early episodes often explore well-trodden paths, but occasionally a novel combination of methods emerges — one that looks unusual compared to the thousands of human approaches that came before it. The agent’s novelty scores spike during these moments, though translating that novelty into sustained performance improvement proves far more challenging.
## Exploration Gives Way to Exploitation — Quickly
Across all configurations tested, agents exhibit a remarkably consistent behavioral pattern. They begin runs with broad exploration, trying diverse approaches and scoring high on originality metrics. Within a handful of episodes, however, they converge on a particular strategy and begin refining it rather than seeking alternatives.
This shift from exploration to exploitation happens faster in AI agents than in human competitors. While human researchers tend to maintain some level of exploration throughout a competition, agents show a steep drop-off in novelty-seeking behavior early on. Strategic reasoning about exploration accounts for roughly three-quarters of the agent’s reasoning traces in the first few episodes, but drops to around one-quarter by the end of a run.
Perhaps the most striking finding is that novelty and performance impact are essentially uncorrelated. Optimizing for originality does not automatically produce better results, and high-impact solutions are not always the most novel ones. This decoupling explains a great deal about why agents struggle: they can reach territories that no human has visited, but they lack the mechanism to recognize which novel territories are worth investing in.
## Search Strategies Matter Less Than You’d Think
One might expect that different search strategies — whether greedy tree search, Monte Carlo tree search, or evolutionary approaches — would produce meaningfully different creativity profiles. Surprisingly, they do not. Different strategies converge to similar ranges of originality and impact scores over time, suggesting that the search algorithm itself is not the primary driver of creative behavior.
What does matter is everything surrounding the search: the prompts, the way context is passed between steps, the scaffolding that structures how the agent reasons, and the feedback mechanisms that guide future decisions. This finding shifts attention from designing clever search algorithms to designing the entire agent ecosystem as a cohesive creative system.
## High Novelty, Low Impact — The Core Paradox
One of the most illuminating findings comes from comparing agent solutions directly against human medalists. When measured against historical human solutions, certain agent configurations achieve novelty scores roughly twice as high as gold medal-winning human approaches. Yet the same configurations achieve medal-winning performance in only a small fraction of runs.
This paradox reveals a fundamental limitation: agents are excellent at finding unusual directions, but they struggle to identify which unusual directions are actually productive. The novelty they generate often consists of approaches that humans abandoned early on, methods that look interesting in isolation but fail to generalize, or combinations that are novel but infeasible at scale. The gap between “what’s new” and “what works” remains wide.
Temporal analysis of agent behavior shows that their novel ideas are spread throughout a competition’s timeline — they aren’t simply rediscovering early human approaches and getting stuck there. Stronger models do show a tendency to cluster around more mature human solutions, suggesting that better reasoning capabilities help agents converge toward proven strategies more effectively. But this convergence also means they explore less broadly overall.
## The Path Forward
The central insight from this line of research is that building AI systems capable of autonomous scientific discovery requires optimizing for both novelty and impact simultaneously. A dual objective — encouraging exploration of new ideas while rewarding solutions that actually move the needle on performance — may be the key to closing the gap between agent and human capabilities.
Several challenges remain. Current evaluations are constrained by episode limits and compute budgets, meaning longer, more sustained creative processes cannot yet be fully captured. Extending these frameworks to open-ended research settings, where success criteria are less clear and reference corpora are sparser, will require new approaches to measuring usefulness and relevance.
The most promising near-term direction appears to be human-AI collaboration, where the complementary strengths of human judgment and machine exploration combine to produce results neither could achieve alone. As model capabilities continue to advance rapidly, the prospect of AI agents conducting science with increasing autonomy is becoming more tangible — but getting there will require solving the creativity measurement problem first.
—
## Frequently Asked Questions
**What is the difference between P-Creativity and H-Creativity?**
P-Creativity measures how novel an idea is relative to the agent’s own prior attempts within a single run. H-Creativity measures how novel an idea is when compared against the full body of human-produced solutions in a given domain. An agent can score highly on self-originality while still producing ideas that closely mirror what humans have already discovered.
**Why do agents lose their novelty so quickly?**
Agents tend to shift from exploration to exploitation within a few episodes because their current framework incentives and reasoning patterns favor refining known strategies over venturing into unfamiliar territory. The scaffolding, prompts, and feedback loops in most agent systems do not explicitly reward continued exploration, leading to a rapid convergence on a narrow set of approaches.
**Is there a correlation between novelty and performance in AI agents?**
Research in this area has found that novelty and impact are essentially uncorrelated. High novelty does not predict high performance, and high performance does not require high novelty. This means that simply encouraging agents to be more creative will not automatically make them more effective.
**How was feasibility measured?**
Feasibility was treated as a binary gate: only episodes where the agent’s proposed code ran successfully were included in the creativity analysis. This means that a novel but non-functional solution would not contribute to impact scores, highlighting the practical barriers agents face when translating creative ideas into working results.
**Do stronger models always produce more creative or impactful results?**
Not necessarily. While stronger models like GPT-5 achieved higher novelty scores than weaker models in some configurations, they did not always translate that into better final performance. Model strength interacts with the broader agent framework, search strategy, and evaluation setup in complex ways.
**What does it mean that agent neighbors span the full timeline of human competition?**
This means that AI agents do not simply rediscover the earliest or most obvious approaches humans tried. Their novel ideas reference solutions from throughout a competition’s history, suggesting that stronger models can converge on both early and late-stage human strategies depending on the task and available context.
**Can creativity metrics be used to improve agent training?**
Yes, this is an active area of investigation. P-Creativity and impact scores offer promising signals for reinforcement learning and training setups, where encouraging both originality and usefulness could help agents learn to balance exploration with practical problem-solving.
—
## Conclusion
The journey toward AI systems that can genuinely contribute to scientific discovery requires us to move beyond simply measuring output quality and start understanding how those outputs are generated. Creativity — properly decomposed into originality, impact, and feasibility — provides a powerful framework for this understanding. The evidence clearly shows that current agents, even those backed by the most capable models available, struggle to convert their novelty into meaningful results. This gap is not a dead end but a roadmap: it tells us precisely what needs to be improved — the balance between exploration and exploitation, the design of the agent’s reasoning scaffolding, and the alignment of creative incentives with practical impact. As we refine our ability to measure and cultivate genuine creativity in AI systems, the prospect of autonomous scientific agents moves closer to reality.
Thank you for reading



