Issue #125

EY, KPMG, Deloitte Reports Hit by Fake Citations

Independent audits found many hallucinated or garbled citations in three Big Four AI reports, raising questions about verification.

AI & TechEY, KPMG, Deloitte Reports Hit by Fake Citations

Rechecking the Sources Behind Public Reports

In May 2026, EY, and in June, KPMG, both pulled public reports after citation errors surfaced. GPTZero classified about 60% of the 27 references in EY’s report as hallucinated citations. For KPMG’s report, it found that only 5 of 45 citations pointed to the correct source. That doesn’t mean the other 40 were entirely fabricated documents, though. The mix included cases where a real source’s title or author information was simply wrong, and cases where the underlying material was hard to pin down at all. Add Deloitte’s case to the pile, and three of the Big Four1 now have reports with confirmed citation-verification problems.

GPTZero, which builds AI-detection and citation-verification tools, calls this kind of error “vibe citing”2. The term covers not just sources that don’t exist, but citations that plausibly distort the title or author of a real document.

Looking at these cases, what caught my attention was who checked what, and when, before the AI’s errors made it into the final product. Improving the accuracy of the tools and verifying sources inside the organization are both necessary, together.

The reports GPTZero investigated are EY Canada’s “Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems” and KPMG’s “Total Experience: Redefining Excellence in the Age of Agentic AI.” Let’s walk through what the two investigations confirmed, separating that from the research team’s own inferences.

EY: Citations That Don’t Exist and Numbers That Don’t Match

EY’s report addressed cybersecurity vulnerabilities in loyalty programs like points and mileage. Published in late 2025, the 44-page report listed two partners and one senior manager as authors.

GPTZero checked the citations in the body text and the reference table at the end of the report, publishing its findings on May 14, 2026.

  • Unverifiable sources: Many of the citations had broken links, or no document matching the stated title could be found. That said, a dead link alone doesn’t prove a document never existed in the first place.
  • A fake McKinsey report cited: No source called the “Loyalty Economics Report (2022)” could be verified. The investigation team found that a similar claim, attributed to the same source, had previously appeared in an article from a different fintech outlet.
  • Numbers that don’t add up: Earlier in the report, the total points market is described as $200 billion, with 30-50% unused. Later, the report states that unused points alone amount to $200 billion. If these are meant to describe the same range, they can’t both be true.
  • AI-detection results: GPTZero also presented an AI-detection score. But a detector’s score alone can’t confirm that AI wrote exactly 72% of the text. Directly cross-checking sources against claims is the more important form of evidence.

According to FT’s reporting, EY said it has taken the report down and is reviewing how it was published. The company also stated that the report wasn’t tied to any client project. That explanation doesn’t erase the responsibility to verify a public report, but it shouldn’t be stretched to mean the company had no idea how the report was written or where any of its data came from.

Next, on June 12th, GPTZero released its findings on a KPMG report.

KPMG International published this report on agentic AI3 and customer experience in October 2025. According to GPTZero’s classification, of 45 citations, 5 pointed accurately to their sources, while 28 linked to real material but altered titles or misstated components. The remaining 12 had vague or incorrect information that made it hard to pin down the source.

Whether the cited material actually supported the claims was also an issue. GPTZero assessed that roughly half of the claims attached to these citations appeared to be false or linked to the wrong source. This doesn’t mean GPTZero verified every fact in the report and concluded half were false. One example: the report described Emirates’ “Sara” as a mobile chatbot capable of changing flights, when Sara was actually a check-in assistance robot introduced in 2023.

Some figures didn’t even match the company’s own materials. The report stated that 55% of CEOs rank AI as their top investment priority, citing “KPMG research” as the source. GPTZero pointed out that this differs from the 71% figure in KPMG’s own CEO Outlook, published the same month. Only by confirming exactly which survey and question was being cited could one determine whether this discrepancy is explainable or simply a misquoted number.

KPMG said it took the report down from its website while conducting an internal investigation. The firm also stated that it requires adherence to responsible AI usage guidelines, including human oversight and independent source verification.

The problem also shows up in government reports, court filings, and academic papers

In 2025, Deloitte agreed to refund part of its fee after citation errors were found in a welfare-related report it produced for the Australian government. A Canadian provincial healthcare workforce report was flagged for the same thing—studies that didn’t exist, authors misattributed. This wasn’t a problem confined to public-facing PR material.

Something similar happened in the legal industry. In an April 18, 2026 letter to the U.S. federal bankruptcy court in New York, Sullivan & Cromwell partner Andrew Dietderich acknowledged that a filing contained errors, including inaccurate citations and AI hallucinations. His explanation: the firm had AI-use guidelines and standard citation-checking procedures in place, but they failed to catch the problem this time.

GPTZero examined 4,841 papers accepted at NeurIPS 2025 and, in a table published in January 2026, flagged 100 hallucinated citations spanning 53 papers. The firm noted that a human had verified the citations separately from the AI detection score. A database of AI-hallucination court cases maintained by Damien Charlotin is likewise collecting instances of fabricated citations turning up in court filings. You can’t read these case counts as a comprehensive incidence rate at any given moment—but they do confirm one thing: errors persist even in documents that have supposedly been reviewed by experts.

Three Things to Check When Verifying Deliverables

Public investigations alone can’t reveal everything happening inside each company. Still, I think organizations need to check the following three things.

First, even public-facing promotional reports need a designated person responsible for verification.

These public reports are materials that showcase a company’s expertise to potential clients. Even if a document isn’t delivered directly to a client, there are people reading it who see the company’s name attached. When setting output volume and speed, I think time for source verification needs to be allocated as well. That said, it hasn’t been confirmed that production pressure was the direct cause of this incident, or that internal review was looser than for client deliverables.

Second, AI usage guidelines need to translate into actual verification of outputs.

Deloitte Australia’s report disclosed that Azure OpenAI was used during the revision process. Sullivan & Cromwell also acknowledged having policies and training in place that simply weren’t followed. Beyond disclosing whether AI was used, there needs to be a record of who checked the original text, figures, and citations.

Third, errors that pass through other sources also need to be traced back to their origin.

GPTZero treated the appearance of the same fake McKinsey source in both the EY report and an earlier fintech outlet’s article as a case of “indirect hallucination”4. It didn’t go so far as to prove whether the AI had trained on that article or encountered it through search results. What matters is that a mistaken citation can get reprinted in another document, and that document can then be used as fresh evidence elsewhere. The investigation team presented a case where the EY report’s error also showed up in an AI search answer. This is exactly why source-checking is necessary even for documents from well-known institutions.

Oswarld’s Lens

Watching this incident unfold, I found myself thinking that any organization claiming to manage AI’s errors needs to actually demonstrate that capability through its own output.

There’s a pattern I’ve seen often while building GTM strategies. When a new tool comes out, companies pitch it to customers for adoption while also using it internally. To play both roles at once, a company needs to be able to show, in its actual output, what the risks are and how it checks for them. If a company’s report on AI usage or governance itself contains sourcing errors, customers can’t help but question that company’s verification methods too.

What concerns me most is the time verification takes. Say checking a single citation takes 3 minutes on average — checking 27 of them would take 81 minutes. This isn’t a real measured average; it’s just a way to think through the actual workload involved. Even if generation time gets cut down, you have to factor in how much longer source-checking and error correction take on top of that — only then can you tell whether AI actually saved any time.

Looking at this fragment, the translation is accurate and well-structured. No Hangul characters, no numbers to verify beyond the list numbering (1, 2, 3), and the structure matches the source exactly.

Closing

If your organization publishes reports, you need to set not just a target speed for writing them, but also the conditions that verification must meet before a report goes out.

  1. The fact that similar errors turned up across multiple firms is reason enough to review your own organization’s procedures. That said, a handful of cases can’t tell you the error rate for the whole industry, or the state of any other specific company.

  2. Don’t stop at drafting an AI usage policy — you need to actually assign a person and set aside time to check citations in real documents.

  3. It’s not enough to confirm that a source exists; you also need to check whether it actually supports the claim in the text. Get a title or a figure wrong, and even a citation of a real source ends up delivering a false explanation to the reader.

When you read a consulting report or research material, open the sources behind its key claims yourself. Check whether the title, author, and date are correct, and whether the same figures and explanations actually appear inside.

Looking at the fragment, I compared structure, links, footnotes, numbers, and glossary terms — everything matches well. No issues found.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

Background

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

Footnotes

  1. Big Four: refers to Deloitte, PwC, EY (Ernst & Young), and KPMG, the world’s four largest accounting and consulting firms. They handle accounting audits, strategy consulting, and tax advisory for major corporations and government agencies around the world.

  2. Vibe citing: a coined term for using AI-generated fake sources or citations without verification. It’s a phrase GPTZero uses in its investigations, and it also covers citations that distort the title or author of a real source.

  3. Agentic AI: an AI system that breaks a given goal into tasks and executes multiple steps using the data and tools it needs. How much of the process it handles autonomously depends on its design and permissions.

  4. Secondhand hallucination: refers to the phenomenon where false information that appears to have originated from AI gets cited again in other writing or AI responses. Whether the error spread through training or retrieval needs to be verified separately.