Your References Deserve Better Than a Confident Guess

citation checking without ai

Your References Deserve Better Than a Confident Guess

In 2023, a New York attorney named Steven Schwartz submitted a legal brief citing six federal court decisions. Opposing counsel could not find any of them. Neither could the judge. The cases had names, docket numbers, procedural histories, and quoted passages from real-sounding judges. ChatGPT had generated all of it. When Schwartz asked ChatGPT to confirm the citations were real, it told him they were. He was sanctioned, fined $5,000, and became the cautionary tale every researcher now hears about AI and citations. [1]

That story is from law. But the same failure mode has arrived in academic research, and it is accelerating.

A Columbia University study published in The Lancet in May 2026 scanned 2.5 million biomedical papers and found that fabricated citations surged 12-fold between 2023 and early 2026, tracking almost exactly with the adoption of AI writing assistants. [2] A separate audit of NeurIPS 2025 submissions found hallucinated citations in 53 published papers, each of which had passed review by multiple expert researchers. [3] A Nature analysis from April 2026 concluded that tens of thousands of publications from 2025 likely include invalid references generated by AI. [4] By early 2026, 1 in every 277 PubMed-indexed papers contained at least one reference that did not exist.

This is the context in which we decided, deliberately, not to use AI to check your citations.

A Confident Answer and a Correct Answer Are Not the Same Thing

AI language models do not look things up. They predict what a correct-sounding answer looks like, based on patterns absorbed during training. When asked about a citation, a language model does not query a database. It generates a plausible response. The citation it produces may be real, partially real, or entirely invented, and it will present each outcome with the same confident tone.

This is not a flaw that better models will eventually fix. It is a structural property of how language models work. Prediction and verification are fundamentally different operations.

Think about how a spell-checker works. When it flags a misspelling, it does not estimate or interpret. It compares your word against a fixed reference and returns a binary result. Run the same check a thousand times, you get the same answer. CiteOrbit works the same way: when it verifies a citation, it sends structured metadata (a DOI, an ISSN, an author name, a publication year) to authoritative bibliographic databases, and checks what comes back. Either the reference matches the record, or it does not. No confidence score. No variation between sessions. No hallucination risk.

When CiteOrbit flags an error, that flag is tied to a specific, verifiable discrepancy. When a manuscript clears, that result is stable and repeatable.

What Actually Happens to Your Manuscript

This comes up often, especially from researchers working on unpublished findings or patent-adjacent work, and the answer is straightforward.

  • Your manuscript is never stored: Documents processed through CiteOrbit are handled in memory and discarded immediately after the session ends.
  • Your data is never used for model training: CiteOrbit does not operate a machine learning model that learns from submitted content. There is no model to feed.
  • Queries go outward, not inward: When CiteOrbit checks a reference, it transmits structured metadata to bibliographic databases. It does not send your prose, your abstract, your conclusions, or your unpublished data anywhere.

The Right Tool for a Lookup Problem

We are aware that "we don't use AI for this" sounds, in 2026, like an unusual position. We think it is the correct one.

AI is a powerful tool for the right problems. Citation verification is not a generation problem or a synthesis problem. It is a lookup problem, and lookup problems call for deterministic tools. Applying a probabilistic text engine to a task that requires database queries is not a shortcut. It is a category error.

Research integrity depends on trust in process. A peer reviewer, a journal editor, an institutional compliance officer all need to know that a passing citation reflects a verified fact, not a model's best guess on a given day. That is what CiteOrbit is built to provide.

For drafting, synthesis, and literature exploration, there are excellent AI tools. CiteOrbit is built for the part of the workflow where "probably correct" is not good enough.

See It in Action

If you are evaluating reference integrity tools for your lab, journal, or institution, we would be glad to walk through a real manuscript and show you exactly what gets checked, what gets flagged, and what never leaves your environment.

[Try CiteOrbit for Free] | [Learn more about CiteOrbit]

References

[1] Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023). Available from: https://legalclarity.org/what-happened-in-the-mata-v-avianca-case/

[2] Topaz M, Roguin N, Gupta P, Zhang Z, Peltonen L-M. Fabricated citations: an audit across 2.5 million biomedical papers. Lancet. 2026 May 7;407:1779–1781. doi:10.1016/S0140-6736(26)00603-3

[3] Ansari S. Compound deception in elite peer review: a failure mode taxonomy of 100 fabricated citations at NeurIPS 2025 [Preprint]. arXiv. 2026 Feb 5. Available from: https://arxiv.org/abs/2602.05930

[4] Naddaf M, Quill E. Hallucinated citations are polluting the scientific literature. What can be done? Nature. 2026 Apr 1;652:26–29. doi:10.1038/d41586-026-00969-z

[5] Charlotin D. AI Hallucination Cases Database [Internet]. Paris: Damien Charlotin; 2023 [cited 2026 Jun 1]. Available from: https://www.damiencharlotin.com/hallucinations/