Articles › Essay 01

Cheap Illusion of Intelligence

Prompt and Circumstance: AI’s Houdini Heist

We are living through an era defined by a spectacular, trillion-dollar misunderstanding. Over the past few years, the tech industry has successfully conflated fluency with cognition, selling the world on the idea that because a machine can speak with absolute confidence, it must know what it is talking about. As a result, we are drowning in a sea of synthetic text that sounds remarkably smart but fundamentally means nothing. It is time to state a plain truth that the hype cycle desperately wants to ignore: the appearance of intelligence is incredibly cheap, but real intelligence remains profoundly rare. The ability to string together grammatically perfect, authoritative-sounding sentences is no longer a proxy for human thought. It is a commodity, generated at scale for fractions of a penny. But real intelligence—the ability to genuinely reason, to synthesize ground truth, to weigh consequences, and to know when you do not know—cannot be manufactured by simply predicting the next most statistically likely word.

The High-profile Era of AI Slop

If you want to see the gap between the appearance of intelligence and the reality of it, you only need to look at the mounting wreckage of high-profile AI failures. We have entered the era of "AI slop," where the consequences of delegating human judgment to statistical models are playing out in courts, boardrooms, and research institutions. Enterprise leaders are being sold the vision that models can digest thousands of pages of legal or financial documents to output reliable analysis. Instead, they are receiving statistically plausible hallucinations that collapse under scrutiny. The consulting industry is currently learning this the hard way. In late 2025, Deloitte was forced to issue a partial refund to the Australian government for an A$440,000 report analyzing the nation’s welfare compliance system. Academics quickly discovered that the report was riddled with AI-generated errors, including a completely fabricated quote attributed to a federal court judgment and citations to nonexistent academic studies. Months later, Ernst & Young (EY) Canada was forced to completely pull a 44-page cybersecurity report on loyalty rewards programs from its website. An investigation revealed that 16 of the 27 cited sources were entirely fabricated. The AI had invented references to nonexistent articles from Forbes and McKinsey, generating what researchers call "vibe citations"—text that flawlessly mimics the syntax and formatting of an industry reference without retrieving an actual document. This proves that models do not "research" or "read"; they merely probabilistically generate text to satisfy the formatting constraints of the user’s prompt.

Sycophancy and the Death of Objectivity

The failures extend well beyond corporate white papers into our most critical civic and legal systems. In a moment of profound irony, a federal judge in Minnesota recently threw out the expert testimony of Stanford University misinformation professor Jeff Hancock. Hancock admitted to using ChatGPT to help prepare a sworn declaration, instructing it to insert citations; instead, the bot fabricated academic authors (e.g., a fake study by "Huang, Zhang, Wang"), completely shattering his credibility in court. Because LLMs lack an internal mental model of truth, the model treated the command to "cite" simply as a prompt to generate characters that looked like a citation, proving it cannot differentiate between verifiable fact and creative fiction. Even purpose-built, highly expensive enterprise tools suffer from these architectural flaws. A landmark Stanford University study titled Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools benchmarked premium platforms and found staggering error rates: Lexis+ AI hallucinated 17% of the time, and Westlaw’s AI tool failed 33% of the time. Crucially, the Stanford researchers identified that these tools suffer from misgrounding and sycophancy. Misgrounding occurs when a model cites a real case but entirely fabricates what the judge actually ruled, providing a dangerous illusion of accurate research. Sycophancy occurs when the system prioritizes user alignment over objective truth; if a user asks a model to support an incorrect legal proposition, the AI will happily generate plausible-sounding arguments using mischaracterized authorities rather than correcting the user’s mistaken premise.

The "Reasoning" Mirage vs. Stochastic Generation

The tech industry insists that models are developing genuine reasoning capabilities to solve these issues, but rigorous testing reveals they are merely engaged in advanced pattern matching. Research demonstrates a "lost in the middle" phenomenon where models reliably extract information only from the very beginning or end of a prompt, completely ignoring critical data buried in the middle of long contexts. Feeding a model massive amounts of text does not equate to comprehensive reading or synthesis, but rather a superficial skimming operation that fails dangerously in complex retrieval tasks. Furthermore, Apple researchers recently introduced the GSM-Symbolic benchmark, demonstrating that when the superficial numerical values in a logic puzzle are slightly altered, model accuracy plummets. This proves that models rely on memorized pattern matching rather than true formal logic. There is a fundamental difference between the rigorous, multi-step measurement of quantities required for genuine reasoning and the stochastic generation of text; the models are simply predicting the shape of a correct answer rather than actually computing it.

Revaluing Real Intelligence

The core problem is architectural. Large language models are, at their foundation, sophisticated pattern matchers. They do not hold a consistent mental model of the physical world, they do not understand the stakes of a bad decision, and they do not actually compute math or logic. Believing that simply feeding these models more data and giving them more compute power will suddenly spark genuine, reliable reasoning is a leap of faith, not a scientific certainty. For low-stakes drafting or brainstorming, that cheap illusion of intelligence is useful. For critical, high-stakes decision-making, it is actively dangerous. Real intelligence requires the capacity for doubt, the ability to self-correct based on physical reality, and the moral weight of accountability. As the digital world floods with cheap, authoritative-sounding slop, the organizations and leaders who thrive will not be the ones who blindly offload their reasoning to predictive text generators. They will be the ones who recognize that while the appearance of intelligence can be bought for pennies, real intelligence is entirely irreplaceable.

Sources

  1. 01

    Alexis, A. (2025, October 21). Deloitte refunds over $60K for report with AI errors, Australian government says. CFO Dive. Documents Deloitte’s partial refund to the Australian Department of Employment and Workplace Relations due to hallucinated GPT-4o quotes, including fabricated federal court judgments.

  2. 02

    Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. Stanford University. Identifies baseline legal hallucination rates and the architectural failure of models generating fake case law.

  3. 03

    Ho, D. E., et al. (2025). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Stanford University. Details the 17% to 33% hallucination rates in premium enterprise legal tools, specifically identifying the dangers of "misgrounding" and AI "sycophancy."

  4. 04

    Kundaliya, D. (2026, May 22). EY cybersecurity report pulled after probe finds ‘AI hallucinations’. Computing UK. Outlines the retraction of the EY Canada "Points of Attack" report after GPTZero identified 16 fabricated "vibe citations" mimicking Forbes and McKinsey.

  5. 05

    Liu, N. F., et al. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. Proves the architectural limitation where models fail to retrieve relevant information buried in the middle of long-context prompts.

  6. 06

    Mirzadeh, I., et al. (2025). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. Apple Research / arXiv. Demonstrates that large language models experience cognitive collapse when logic variables are altered, proving they rely on pattern matching rather than true mathematical reasoning.

  7. 07

    Tribune News Service. (2025, January 16). Judge Blasts Stanford AI Expert’s Credibility Over Fake, AI-Created Sources. GovTech. Documents the federal court decision to strike Stanford Professor Jeff Hancock’s testimony due to ChatGPT generating fabricated authors for nonexistent studies.

  8. 08

    ABC News coverage of the Deloitte AI report fallout. This news segment provides a firsthand look at the public and governmental response to the errors found in Deloitte’s AI-generated welfare compliance report in Australia.

We show you a new you

Uncovering the insights only top-level consultancies can distill.

Get in touch