---
title: "The Boy Who Cried Wolf"
description: "AI is finally catching up with the promise. What if we’ve stopped listening?"
url: "https://reasonmachines.com/blog/the-boy-who-cried-wolf"
section: "Blog"
---

# The Boy Who Cried Wolf

_AI is finally catching up with the promise. What if we’ve stopped listening?_

September 21, 2026

The boy cried wolf. The village came running. Nothing. He cried again. Same story. When the wolf finally arrived, nobody moved.

For the past two years, AI has had a version of this problem. We were sold finished work and got impressive demos. We were promised a colleague and got another tool to supervise. It is understandable that some people stopped listening.

Our thesis: AI is finally approaching what people imagined it could do two years ago. Not everything, not reliably in every setting—but enough real work to change how we build. The danger is judging today’s systems by yesterday’s disappointment.

## Progress doesn’t automatically rebuild trust

There is evidence for both progress and skepticism. Stanford’s 2025 AI Index reports a jump in SWE-bench coding-problem performance from 4.4% in 2023 to 71.7% in 2024. Meanwhile, Pew found that half of U.S. adults were more concerned than excited about AI in 2025; only 10% were more excited. [1](#ref-1) [2](#ref-2)

These are different measures, not a shared scale. A benchmark is not an economy, concern is not disbelief, and American opinion is not global opinion. The charts below are historical snapshots—not September 2026 readings or proof that overpromising caused skepticism.

## Don’t believe the pitch. Try the work.

The evidence is still messy. METR’s early-2025 experiment found experienced open-source developers took 19% longer with AI. Its February 2026 update says that historical result likely no longer reflects current tools—but selection effects make the new speedup estimates unreliable. That is a reason to measure carefully, not declare victory. [3](#ref-3) [4](#ref-4)

Maybe the cost of the hype was not just disappointment. Maybe it trained people to look away at exactly the wrong moment. That is our interpretation, not a finding from these studies.

In the fable, the final warning was true. In AI, the answer should not be another promise. Give the system a real task. Check the result. Measure the time, including corrections. Update your opinion from the work—not the shouting.

## References

<a id="ref-1"></a>
[1] [Stanford HAI, 2025 AI Index Report](https://hai.stanford.edu/assets/files/hai_ai_index_report_2025.pdf). Reported SWE-bench results for 2023–2024; a historical frontier benchmark comparison, not a controlled productivity study.

<a id="ref-2"></a>
[2] [Pew Research Center, How Americans View AI and Its Impact on People and Society](https://www.pewresearch.org/science/2025/09/17/how-americans-view-ai-and-its-impact-on-people-and-society/). Published September 17, 2025; 5,023 U.S. adults surveyed June 9–15, 2025. Percentages shown do not sum to 100 because remaining responses are omitted.

<a id="ref-3"></a>
[3] [METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/). July 10, 2025; 16 developers, 246 tasks. Historical result, not a claim about today’s tools.

<a id="ref-4"></a>
[4] [METR, We Are Changing Our Developer Productivity Experiment Design](https://metr.org/blog/2026-02-24-uplift-update/). February 24, 2026. Explains why selection effects and concurrent-agent time measurement limit interpretation of the follow-up study.
