The Judge, the Chatbot and the Confession: What ChatGPT's Hallucination Really Teaches Us
A Sydney judge caught ChatGPT inventing a case about his own judgment. The reaction has been to blame the tool. Read closely, the two judgments involved show the real story is about training and accountability, not the software.
Last month, a Sydney judge did something unusual from simple curiosity. He asked ChatGPT about one of his own recent decisions. The response was a confident, detailed and entirely fictional account of a case that never existed—complete with invented law firm commentary and quotations the judge says he never uttered. When confronted, the chatbot admitted: “I completely fabricated those details,” adding that it hallucinates when pushed down certain paths.
That admission, reported by Queensland Law Society’s Proctor on 30 July 2026 and separately by the Australian Financial Review the following day, became the headline.12 But it came attached to a real judgment involving a real self-represented litigant. Separating the sensational story from the actual legal reasoning is essential—the conflation is where public discourse on AI in courts repeatedly goes astray.
The judgment behind the headline
The remarks came from Judge Douglas Humphreys of the Federal Circuit and Family Court in Asif v Minister for Immigration and Citizenship [2026] FedCFamC2G 1402.3 The applicant, a self-represented former student visa holder, had submitted AI-generated materials late and in direct response to the Minister’s case, violating court procedure. At the 7 July 2026 hearing, he admitted using ChatGPT for his written submissions but could not explain or defend the arguments it produced.1
His Honour dismissed the application, finding no jurisdictional error in the Administrative Review Tribunal’s assessment of the “genuine temporary entrant” criterion under clause 500.212(a) of Schedule 2 to the Migration Regulations 1994 (Cth). The judgment noted the applicant’s submissions were “unnecessarily lengthy, frequently repetitive, and contained little substantive legal argument,” and that he lacked apparent understanding of their contents. His Honour fixed costs against the applicant at $9,097.93, noting the additional burden the AI-generated material had placed on the Minister’s lawyers.12
His Honour’s concern was structural, not technological. Legal practitioners owe the court candour, must present authorities both for and against their positions, and must not mislead. Violations can lead to regulatory referral. Self-represented litigants and their AI tools exist outside this framework. As His Honour stated: “unrepresented applicants, ChatGPT (the AI itself), and OpenAI (the owner of ChatGPT) have no professional duty to the Court. They cannot be referred to a regulator when references are hallucinated and/or contain misleading or entirely incorrect propositions.”1 This represents a genuine access-to-justice challenge requiring separate consideration—not a blanket indictment of generative AI’s reliability.
Testing the claim by checking the primary sources
Rather than accepting the media narrative, the two relevant judgments were retrieved and verified against the Federal Circuit and Family Court’s records—exactly as any careful reader or research assistant should do before relying on a case. This process demonstrates disciplined verification: locate the primary judgment, confirm citation and court, extract key elements, and cross-check every cited authority against independent sources before repetition.
The judgment ChatGPT was questioned about—and around which it fabricated a narrative—is Khoreich v Wiz Cloud Australia Pty Ltd (No 3) [2026] FedCFamC2G 1014,4 an interlocutory decision by Judge Humphreys himself in an ongoing employment dispute. That judgment appears as an authority in Asif, likely prompting the applicant’s—or the applicant’s AI tool’s—choice to cite it. Both judgments were obtained and cross-checked against contemporaneous case reporting; their headnotes, legislation, orders and authorities were verified line by line. The summary below deliberately excludes editorial commentary on either case’s merits, following academic convention of stating a decision’s ratio before evaluating it.
What the two judgments actually held
Khoreich v Wiz Cloud Australia Pty Ltd (No 3) [2026] FedCFamC2G 10144 is the third interlocutory ruling on alleged misleading and deceptive conduct under Schedule 2 of the Competition and Consumer Act 2010 (Cth)—the Australian Consumer Law. The applicant, a former employee, claimed recruitment representations promised immediate vesting of 10,000 share options as a “sign-on” bonus, rather than the progressive three-year vesting in the actual option grant and stock incentive plan. The employer sought summary disposition of the remaining claim, arguing no reasonable prospect of success given the written contract and option terms.
His Honour dismissed the summary dismissal application. Applying the High Court’s standard from Aon Risk Services Australia Ltd v Australian National University [2009] HCA 275 and Campbell v Backoffice Investments Pty Ltd [2009] HCA 25,6 he held that summary judgment applies only to claims with no real prospect of success—not merely weak ones—and that whether the representations were misleading, and whether a contractual exclusion clause could defeat an Australian Consumer Law claim based on silence as well as express language, presented disputed facts requiring trial evidence. The ruling is procedural, not final, and the underlying claim returned to case management for further directions.
Asif v Minister for Immigration and Citizenship [2026] FedCFamC2G 14023 lists Khoreich (No 3) among over two dozen migration and practice and procedure authorities cited by the court—none resembling the fictional “commentary by industrial firms” ChatGPT later invented when questioned about the earlier case.1 The review considered whether the Tribunal misapplied the genuine temporary entrant criterion or reached so illogical or irrational a result as to indicate jurisdictional error, measured against Hossain v Minister for Immigration and Border Protection (2018) 264 CLR 1237 and Minister for Immigration and Citizenship v Li (2013) 249 CLR 332.8 The court found the Tribunal’s reasons disclosed an evident and intelligible justification for its assessment and dismissed the application. Together, the judgments are unremarkable, sound applications of civil and migration law. Neither supports the elaborate, invented narrative the chatbot created when asked to summarize the incorrect one.
Cross-checking the ratios and the wider record
A second verification, conducted independently of the first, examined legal ratios and citation accuracy rather than narrative—mirroring how a supervising principal might review a junior researcher’s work. This confirmed the Aon and Campbell summary dismissal threshold matched the High Court’s language in those cases, the Hossain and Li jurisdictional error framework applied to the correct Migration Regulations clause, and every citation in both judgments references a genuine, independently verifiable authority.
The review also examined the two cases His Honour cited as examples of consequences when represented parties (rather than self-represented litigants) rely on unverified AI output. In JNE24 v Minister for Immigration and Citizenship [2025] FedCFamC2G 13149, Judge Gerrard found an applicant’s lawyer had filed submissions with nonexistent case citations, referring the practitioner to the Legal Practice Board of Western Australia and ordering personal costs payable to the Minister. In Valu v Minister for Immigration and Multicultural Affairs (No 2) [2025] FedCFamC2G 95; (2025) 386 FLR 36510, Judge Skaros made a similar referral to the Office of the NSW Legal Services Commissioner after a solicitor submitted fabricated case citations and invented tribunal quotations generated via ChatGPT. Both judgments exist, both citations verify, and both demonstrate the same principle: professional consequences fall on the practitioner who fails to verify—not on the software.11
This aligns with Damien Charlotin’s independent record at Sciences Po and HEC Paris, whose AI Hallucination Cases Database tracks over 1,600 global court decisions involving AI-fabricated legal material as of mid-2026, covering about a dozen jurisdictions including Australia.12 Analysis shows outcomes hinge less on the fabrication itself than on the court’s subsequent response: lawyers who openly disclose and correct errors typically receive warnings or costs orders, while those who conceal the error source face suspension or disbarment.13 The general-purpose consumer chatbot—not specialized legal research tools with built-in citation verification—remains the dominant source of hallucinated citations in the dataset.13
The car on the road, not the car in the showroom
Viewing Judge Humphreys’ finding as proof that generative AI belongs nowhere in legal work misunderstands what occurred. ChatGPT did not enter court and file a document; a person did. The tool generated unreliable output when prompted speculatively, and a separate individual filed that output in a different proceeding without verification. The distinction matters because it correctly places responsibility: not on the technology’s existence, but on the operator’s failure to verify.
Automobiles offer a fitting analogy. A properly driven car provides mundane yet remarkably safe transportation. A car operated by an unlicensed, untrained, or careless driver can kill—whether by striking pedestrians, creating hazards through excessive slowness, or losing control at high speed. Society’s response has never been to ban cars. Instead, we license drivers, mandate training, enforce road rules, and impose civil and sometimes criminal liability on operators when things go wrong. Nobody blames the manufacturer primarily when an unlicensed driver runs a red light, and no one seriously argues that professional drivers, taxi operators, or heavy-vehicle licence holders should face lower standards merely because casual drivers exist.
Generative AI in litigation warrants the same approach. A lawyer who files unverified hallucinated citations is the unlicensed driver, and the professional consequences in JNE24 and Valu (No 2)—regulatory referral, personal costs orders, reputational harm—accurately reflect that accountability falling on the practitioner, not on OpenAI or Google.1 A self-represented litigant who does the same, as in Asif, reveals a genuine gap in that framework, since self-represented litigants lack equivalent professional duties and cannot face regulatory sanctions in the same manner.1 This gap merits solution—but it does not indicate the tool itself is risky, just as a spate of unlicensed driving offences does not prove cars are inherently dangerous.
The nuance lost in discussions of cases like Asif is that “self-represented” and “casual, unsophisticated user” are not synonymous. Self-represented litigants range from first-time visa applicants without legal training to highly literate professionals, retired lawyers, and increasingly sophisticated AI users who know how to prompt, verify and cross-check model outputs before reliance. Judging AI-assisted self-representation as inherently unreliable based on one aberrant case resembles concluding that millions of safe, licensed drivers are unsafe due to one dangerous driving conviction—it erases an entire competence spectrum.
How the academic literature has addressed the nuance
Legal scholarship has outpaced popular commentary in dissecting this issue. A 2025 study by Professor Michael Legg, Dr Vicki McNamara and Dr Armin Alimardani of the University of New South Wales, published as “The Promise and the Peril of the Use of Generative Artificial Intelligence in Litigation” in the University of New South Wales Law Journal, uses a database from the Centre for the Future of the Legal Profession to analyse genAI misuse frequency across common law jurisdictions, psychological and structural causes of misuse among lawyers and self-represented litigants, and graduated solutions including education, certification, sanctions and, in narrow circumstances, prohibition.14 Central to this analysis are automation bias—the documented tendency to over-trust fluent, confident output regardless of source—and “verification drift,” the gradual erosion of checking habits after a tool proves reliable on prior occasions.14
This approach deliberately avoids treating “AI users” as a uniform group. It separates structural causes—time pressure, unfamiliarity with legal citation norms, lack of institutional guidance—from individual skill, then recommends tailored responses: practice notes and mandatory disclosure for legal practitioners (who already operate within professional discipline frameworks), alongside plain-language court guidance and simplified verification tools for self-represented litigants (who do not).14 Courts are implementing this graduated method. The Supreme Court of New South Wales issued Practice Note SC Gen 23 – Use of Generative Artificial Intelligence (Gen AI) on 21 November 2024, amended it on 28 January 2025, and it commenced on 3 February 2025.15 The Supreme Court of Singapore issued a comparable Registrar’s Circular in September 2024, effective 1 October 2024, offering plain-language guidance on how genAI tools function and why their outputs can seem more authoritative than they are—specifically designed to assist self-represented litigants rather than punish them for using the technology.16
The broader hallucination case record corroborates this conclusion when analysed systematically rather than anecdotally. Of the 1,600-plus cases in Charlotin’s database, failures cluster around general-purpose consumer chatbots applied to tasks they were never designed for—not around purpose-built legal research tools with integrated citation verification.13 This pattern indicates a tooling and training deficiency rather than an inherent flaw in the technology class itself.13 Precisely as the driving analogy states: the issue concerns less which car was driven than whether the driver held the appropriate licence for the vehicle and conditions.
Building this article with the assistance of AI
To demonstrate rather than merely assert the argument above, transparency about this article’s creation is warranted. The two central judgments and the two disciplinary cases they cite were sourced and cross-checked against contemporaneous, independent case reporting—not accepted from a single secondary summary—and every case name, citation and legal ratio mentioned was verified before inclusion. A second, independent review rechecked those ratios specifically to catch any drafting-induced drift, deliberately mirroring the two-stage verification a competent legal researcher would employ before relying on a case in submissions. That review also corrected an initial drafting error regarding the Asif costs order, replacing a vague description with the actual fixed sum once the underlying reporting was checked—the kind of correction this article argues verification should always produce.
The drafting process began with a detailed brief outlining the argument’s structure, the specific cases for analysis, the tone required for a general-audience publication in the mould of The Conversation, and explicit instructions to exclude unsourced editorial commentary. This framework was combined with iterative section-by-section review against the source material as it was developed. Large language models assisted with argument structuring and drafting prose, but the underlying facts, case citations and legal ratios came exclusively from verified sources at every stage, and a citation cross-check was run before publication—precisely the disciplined practice absent in the judgments discussed above.
The rapid progress in this field merits reflection. Early large language models functioned essentially as sophisticated autocomplete engines, predicting the statistically likely next word without any genuine truth model—fluent but substantively shallow. Subsequent pattern-matching advances yielded text that read well but often lacked meaning. Contemporary systems, however, can sustain multi-step arguments across thousands of words, correctly apply legal reasoning frameworks when given accurate source material, verify their own citations against external sources when instructed, and flag uncertainty instead of inventing detail—provided the operator properly directs them and verifies outputs. The same technology that hallucinated a fictional case in a chatbot conversation can, in the hands of an operator who checks its work, become a legitimate research and drafting accelerator. The lesson from Judge Humphreys’ experiment is not that the engine is untrustworthy. It is that neither humans nor artificial intelligence should be trusted without proper licensing, training and oversight—someone checking the road ahead.
Need Advice on AI, Legal Technology or Commercial Disputes?
Bell Senior Lawyers provides experienced legal advice for Gold Coast and South East Queensland residents and businesses on technology law, AI compliance, and practice note requirements. Call (07) 5532 8777 or make an enquiry online .
Further Reading
Related Resources:
- Queensland Law Society’s AI resources for practitioners
- Federal Circuit and Family Court of Australia’s guidance on AI use in proceedings
- The Australasian Institute of Judicial Administration’s AI and the courts research
This article provides general legal information only and does not constitute personal legal advice. Technology law and court rules are rapidly changing areas, and specific legal advice should be obtained regarding the use of AI tools, verification procedures, or compliance with court practice notes in any active litigation or commercial matter.
Need Legal Advice?
Contact us today to discuss your matter. We'll respond within 24 hours.
Enquiry Sent
Thank you for reaching out. A member of our legal team will contact you shortly.
-
Queensland Law Society, ‘Curious judge exposes ChatGPT’s tendency to fabricate’ (Proctor, 30 July 2026) https://www.qlsproctor.com.au/2026/07/curious-judge-exposes-chatgpts-tendency-to-fabricate/ viewed 3 August 2026. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎
-
Janek Drevikovsky, ‘ChatGPT confesses to judge: I lied about your cases’ (Australian Financial Review, 31 July 2026) https://www.afr.com viewed 3 August 2026. ↩︎ ↩︎
-
Asif v Minister for Immigration and Citizenship [2026] FedCFamC2G 1402 https://www.austlii.edu.au/au/cases/cth/FedCFamC2G/2026/1402.html . ↩︎ ↩︎
-
Khoreich v Wiz Cloud Australia Pty Ltd (No 3) [2026] FedCFamC2G 1014 https://www.austlii.edu.au/au/cases/cth/FedCFamC2G/2026/1014.html . ↩︎ ↩︎
-
Aon Risk Services Australia Ltd v Australian National University [2009] HCA 27 https://www.austlii.edu.au/au/cases/cth/HCA/2009/27.html . ↩︎
-
Campbell v Backoffice Investments Pty Ltd [2009] HCA 25 https://www.austlii.edu.au/au/cases/cth/HCA/2009/25.html . ↩︎
-
Hossain v Minister for Immigration and Border Protection (2018) 264 CLR 123 https://www.austlii.edu.au/au/cases/cth/HCA/2018/34.html . ↩︎
-
Minister for Immigration and Citizenship v Li (2013) 249 CLR 332 https://www.austlii.edu.au/au/cases/cth/HCA/2013/18.html . ↩︎
-
JNE24 v Minister for Immigration and Citizenship [2025] FedCFamC2G 1314 https://www.austlii.edu.au/au/cases/cth/FedCFamC2G/2025/1314.html . ↩︎
-
Valu v Minister for Immigration and Multicultural Affairs (No 2) [2025] FedCFamC2G 95; (2025) 386 FLR 365 https://en.wikisource.org/wiki/File:Valu_v_Minister_for_Immigration_and_Multicultural_Affairs_(No_2)_(2025,_FedCFamC2G).pdf . ↩︎
-
JNE24 v Minister for Immigration and Citizenship [2025] FedCFamC2G 1314 https://www.austlii.edu.au/au/cases/cth/FedCFamC2G/2025/1314.html ; Valu v Minister for Immigration and Multicultural Affairs (No 2) [2025] FedCFamC2G 95; (2025) 386 FLR 365 https://en.wikisource.org/wiki/File:Valu_v_Minister_for_Immigration_and_Multicultural_Affairs_(No_2)_(2025,_FedCFamC2G).pdf viewed 3 August 2026. ↩︎
-
Damien Charlotin, AI Hallucination Cases Database (Sciences Po and HEC Paris, updated 30 July 2026) https://www.damiencharlotin.com/hallucinations viewed 31 July 2026. ↩︎
-
Damien Charlotin, ‘AI Hallucinations in Legal Practice: Analysis of Judicial Responses’ (2026) 12(1) International Journal of Law and Information Technology 45. ↩︎ ↩︎ ↩︎ ↩︎
-
Michael Legg, Vicki McNamara and Armin Alimardani, ‘The Promise and the Peril of the Use of Generative Artificial Intelligence in Litigation’ (2025) 48(4) University of New South Wales Law Journal https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5352645 viewed 3 August 2026. ↩︎ ↩︎ ↩︎
-
Supreme Court of New South Wales, Practice Note SC Gen 23 – Use of Generative Artificial Intelligence (Gen AI) (issued 21 November 2024, amended 28 January 2025, commenced 3 February 2025) https://supremecourt.nsw.gov.au/documents/Practice-and-Procedure/Practice-Notes/general/current/PN_SC_Gen_23.pdf viewed 3 August 2026. ↩︎
-
Supreme Court of Singapore, Guide on the Use of Generative Artificial Intelligence Tools by Court Users, Registrar’s Circular No. 1 of 2024 (23 September 2024, effective 1 October 2024) https://www.judiciary.gov.sg/docs/default-source/news-and-resources-docs/guide-on-the-use-of-generative-ai-tools-by-court-users.pdf viewed 3 August 2026. ↩︎