The date is the only number that tells you how much time the safety people actually have.
I defined the word here. Why a precise picture of the other side is already false is here.
The Uncertainty at the Heart of Every Prediction
Ask ten researchers and you get ten dates. 2027. 2050. Never. That spread is the story. Some people think what is left is an engineering bill. Some think you still need a conceptual jump that does not sit on a calendar.
The Metaculus forecasting platform, which aggregates thousands of individual predictions, has shifted its median AGI estimate from the mid-2040s to the early 2030s over the past three years. This shift reflects the impact of large language models (LLMs) demonstrating capabilities, from multi-step reasoning to code generation to scientific analysis, that many forecasters did not expect this soon.
But forecasting platforms capture sentiment, not physics. The actual timeline depends on several technical factors, each carrying its own uncertainty.
Prediction disagreement reflects genuine uncertainty about whether the remaining barriers to superintelligence are continuous (more of the same, approachable through scaling) or discontinuous (requiring breakthroughs in kind). Researchers who believe capacity scales smoothly with compute predict shorter timelines. Those who believe qualitative jumps in architecture or approach are required predict longer ones. Both positions are empirically defensible with current evidence.
The Scaling Hypothesis and Its Limits
The dominant theory in AI between 2020 and 2024 was the scaling hypothesis: intelligence emerges predictably from scale. Bigger models trained on more data with more compute become proportionally smarter. Chinchilla scaling laws (Hoffmann et al., 2022) formalized the relationship between model size, dataset size, and compute, showing that optimal performance requires scaling all three in proportion.
This hypothesis produced GPT-3, GPT-4, Claude, and Gemini. Each was substantially more capable than its predecessor, and each was substantially larger. The relationship between investment and output appeared predictable.
By 2025-2026, the picture has changed. The industry has begun encountering what can be described as a practical scaling wall: not a hard physical impossibility, but a steep increase in the cost required to achieve the next increment of capability.
The data wall. High-quality, human-generated text on the public internet has been largely exhausted as a training resource. The corpus of books, articles, forum posts, and code repositories that trained current frontier models is finite. Simply scraping more of the same yields diminishing returns. This has forced a shift toward curated datasets, expert-generated data, and synthetic data (using AI to generate training signal for other AI systems).
Training on AI-generated data introduces the risk of "model collapse": performance degradation over successive generations as errors and biases compound. Effective synthetic data requires external verification mechanisms (mathematical proofs, physics simulators, human expert review) to maintain quality. The synthetic data path is viable but not free. It requires infrastructure, verification, and significant engineering investment.
The cost curve. Training frontier models now requires multi-billion-dollar investments in compute infrastructure. The gap between successive model generations (GPT-4 to subsequent models) has often felt smaller to users than earlier leaps, even as costs have increased by orders of magnitude. This is the diminishing-returns pattern characteristic of approaching a performance ceiling: each increment of capability requires disproportionately more resources.
Architecture matters more than scale. The biggest jumps in AI capability have come from architectural innovations, not simply from more compute:
- Transformers (2017): The attention mechanism enabled language modeling at scale
- Scaling laws (2020): Chinchilla-optimal training formalized the relationship between parameters, data, and compute
- RLHF (2022): Reinforcement learning from human feedback made models useful and controllable
- Chain-of-thought / test-time compute (2024-25): Models that "think longer" at inference time outperform larger models on complex tasks
- Mixture of Experts (2024-25): Sparse architectures that activate only a fraction of parameters per query, reducing inference cost dramatically
Each of these was unpredicted. Each accelerated the timeline by years. The next architectural breakthrough, whatever it is, may compress or extend the timeline in ways current forecasts cannot capture.
The industry is transitioning from the "Age of Scaling" to the "Age of Research," where algorithmic breakthroughs in training recipes and reasoning architectures take precedence over stacking more GPUs.
Five Factors That Determine the Timeline
1. Can AI automate AI research? This is the most consequential variable. If AI systems can perform high-quality research (designing better architectures, optimizing training procedures, discovering new algorithms), the rate of progress decouples from the rate of human research output. Current AI systems already contribute meaningfully to coding, mathematical reasoning, and scientific literature review. Whether they can perform the creative, hypothesis-generating work that drives genuine breakthroughs remains unproven.
2. Do physical constraints bind? Chip fabrication (TSMC, Samsung, Intel advanced nodes), data center construction, and energy infrastructure impose real-world bottlenecks. These constraints operate on timescales of years, not weeks. Building a new semiconductor fab takes 3-5 years. Constructing the data center capacity for the next generation of frontier models requires power infrastructure that may not exist in the required locations. These physical bottlenecks may impose a de facto speed limit on any intelligence explosion scenario.
3. Is the scaling hypothesis locally or globally true? Scaling laws describe a predictable relationship between inputs and outputs within a given model. They do not guarantee that the model itself scales to superintelligence. Language modeling may approach human-level performance on many tasks through scale alone, but the gap between "human-level on benchmarks" and "genuinely superintelligent" may require qualitative shifts that scaling cannot produce.
4. Can alignment keep pace? Even if capability advances rapidly, deployment depends on alignment and safety. Regulatory frameworks (the EU AI Act, various national AI safety institutes) increasingly constrain how frontier models can be deployed. If alignment research lags behind capability, the most powerful systems may not be deployable, effectively extending the practical timeline regardless of what is technically achievable.
5. Does recursive self-improvement actually work? The theoretical argument for intelligence explosion (I.J. Good, 1965) assumes that a sufficiently intelligent system can improve itself, and that this improvement compounds. In practice, self-improvement may encounter diminishing returns, architectural constraints, or verification bottlenecks that prevent explosive takeoff. No evidence yet confirms or refutes this assumption at the relevant scale.
The Prediction Paradox
The transformer, scaling-on-text, RLHF, and chain-of-thought were not on the boards that now publish AGI years. If the developments that move the date are the ones nobody listed, a point forecast is a calibration tool, not a clock. The structural version of that claim, why a detailed picture of superintelligence is already wrong, is the paradox.
If the median is 2032, preparing as though it could be 2028 is cheap relative to preparing as though it is 2060. A median is a bound, not a schedule.
What the Transition May Look Like
The binary framing, "superintelligence exists or it doesn't," obscures the more likely reality: a gradual, uneven transition in which AI systems become superhuman in specific domains (mathematics, coding, materials science) while remaining sub-human in others (social reasoning, physical manipulation, common sense).
This "patchy superintelligence" may be the most likely near-term scenario. Not a single moment of transition, but a decade-long process in which the boundary between human and machine capability shifts domain by domain. Some professions are transformed early. Others are affected later. The economic and social disruption is real but distributed, not concentrated in a single shock.
It will not send a press release. Math goes first, then code, then materials. Each step looks manageable. Together they are the event. Prepare for a decade of that, not for a Friday.
The year is a spread: some CEOs say late 2020s, Metaculus says 2030-33, conservative surveys say 2050+. Five things move it: whether the models can do the research, whether the fabs and the power show up, whether scaling stalls, whether safety work delays the release, and whether self-improvement actually compounds. Budget for 2028 and stay willing to be wrong until 2050. The last four jumps were not on the board.