Science
I flagged two research papers for fake authors and both were accepted as orals
Key Points
Q&A from the slop trenches Between the two of us, we reviewed 22 paper submissions this summer, spread across NeurIPS, WACV, and TerraBytes (a geospatial workshop at ECCV). Fifteen of the 22 (68%) contained entirely fabricated citations, fabricated author lists for existing papers, and/or were clearly LLM-generated (e.g. hallucinated technical jargon, nonsensical writing, irrelevant citations). This is called being in the “slop trenches” (i.e. dealing with the output of AI slop cannons).
Q&A from the slop trenches
Between the two of us, we reviewed 22 paper submissions this summer, spread across NeurIPS, WACV, and TerraBytes (a geospatial workshop at ECCV). Fifteen of the 22 (68%) contained entirely fabricated citations, fabricated author lists for existing papers, and/or were clearly LLM-generated (e.g. hallucinated technical jargon, nonsensical writing, irrelevant citations). This is called being in the “slop trenches” (i.e. dealing with the output of AI slop cannons). Here we complain about this being mostly a waste of time, do a small lit review on the state of LLMs in scientific writing and reviewing, get Claude to generate questions for a Q&A, and release a bib-audit
skill that Isaac cooked up.
The table below shows the number of review assignments with fabricated citations, fabricated authors, or writing that was unmistakably LLM-generated, out of the total assignments for that venue.
| Venue | Caleb | Isaac |
|---|---|---|
| NeurIPS (Datasets and Benchmarks track) | 2 of 5 | |
| NeurIPS (Position Paper track) | 2 of 2 | |
| TerraBytes (ECCV workshop) | 1 of 2 | 4 of 5 |
| WACV | 3 of 4 | 3 of 4 |
| Total | 6 of 11 (55%) | 9 of 11 (82%) |
We are definitely not the only ones in the slop trenches. The scale of this problem has been measured in various ways over the past year. A Nature analysis from April found at least tens of thousands of 2025 publications “probably” contain invalid AI-generated references. An audit of arXiv, bioRxiv, SSRN, and PubMed Central by Zhao et al. estimated roughly 146,900 hallucinated citations in 2025 alone, spread thinly across many papers rather than concentrated in a few bad actors, with early-career researchers and small teams the most likely to include them. They also found that reviews are not catching these – they traced bioRxiv preprints that contained hallucinated references to their published versions and found 85.3% of the hallucinations remained. An audit in The Lancet covering 2.5 million biomedical papers found that the share of papers with at least one fabricated reference rose six-fold in two years, from one in 2828 papers in 2023 to one in 458 in 2025, reaching one in 277 in early 2026. Ansari (2026) did an analysis of 100 hallucinated citations drawn from papers that were actually published at NeurIPS 2025. Every one of those citations made it past three to five expert reviewers, and the 53 papers carrying them, about 1% of acceptances, sit in the proceedings today.
The slop also flows both ways. Pangram (a company that has an AI writing detector product) did an analysis of ICLR 2026 reviews and found that 21% of the reviews (15,899 of them!) were fully AI-generated, and over half had some form of AI involvement. Gartenberg et al. (2026) measured a 42% post-ChatGPT surge in submissions to the journal (Organization Science) and found that over 30% of its peer reviews now use some degree of AI. ICML 2026 hid prompt-injection stings in submissions and found “795 reviews (~1% of all reviews) written by 506 unique reviewers who were assigned Policy A (no LLMs) were detected to have used LLMs in their review”. Now that NeurIPS reviews are out and rebuttals are coming in, we’ve seen an obvious AI-generated ethics review, and several obvious AI-generated rebuttals. This is a problem for a whole lot of reasons, one of which is that AI reviewers can be gamed directly – Li et al. (2026) found that adversarially rewritten abstracts improve AI-generated review outcomes “without changing the underlying scientific content and communication of the paper, and even without knowledge of the reviewing model.” Their strongest attack succeeded about 38% of the time, inflating acceptance ratings by +1.31 for Gemini 3 Flash reviewers and +0.88 for GPT 5.4 Mini reviewers on a 10-point scale.
This all is very annoying from inside the review queue. Peer review is unpaid work that we do (often on nights and weekends) because peer review on our own work is so valuable. Spending hours going through a submission and then realizing that there are hallucinated citations is infuriating as it is a waste of our time! If you haven’t spent enough time with your work to even get the references correct, then a.) why should we spend time reviewing it for you, and b.) what are you hoping to accomplish with the submission in the first place? Learning from reviews requires reflecting on your work, and you need to spend time with your work in order to do this.
Q&A
How many papers did you review this summer, and how many would you have desk rejected on citations alone?
Caleb: Eleven (see table above). Five of these had hallucinated authors and/or entire papers in their bibliography. In two of the WACV submissions, reference [1], the literal first entry in the bibliography, listed hallucinated authors for real papers. One of the NeurIPS papers had so much hallucinated/nonsensical jargon (53 pages of it) that there wasn’t any point in going through the bibliography.
Isaac: Eleven as well: two NeurIPS position papers, five for TerraBytes, and four for WACV. Nine of my eleven (!!!). Both position papers were LLM-generated slop, one of which contained the ramblings of an agentic madman. Four of my five workshop assignments had fabricated citations or authors. I’ll be honest, the one human-written paper was incredibly mediocre but didn’t piss me off while reading it so it got my accept.
What is the single worst thing you saw?
Isaac: Two submissions cited papers whose authors I know personally, and swapped them out with imaginary researchers. The paper, venue, and other authors were all real and the invented name was close enough to survive a skim. I recommended reject and flagged them to the organizers directly. Both papers were accepted for oral presentations with the condition that they simply fix the hallucinated references…
Caleb: I mentioned the 53-page NeurIPS paper with hallucinated jargon everywhere. There was also a 40-page one that wasn’t much better (also with made up references). One of these also had embedded notes from the LLM alongside most citations which was pretty funny. I then spent a bunch of time reviewing one of the WACV papers that I thought was pretty good, but then noticed that it listed the SatMAE authors as “Yuyang Cong, Saurabh Khanna, Chen Meng, et al.” The real SatMAE author list starts with Yezhen Cong, Samar Khanna, and Chenlin Meng (arXiv:2207.08051).
Hallucinated authors on a real paper versus fully hallucinated references: which is worse?
Caleb: Both show that the authors haven’t put time into properly citing something and call into question whether they’ve actually read the work they are citing (and whether it’s cited correctly), and what other parts of the paper might be hallucinated.
Isaac: In my opinion, they’re the same thing. If the authors can’t be bothered to read their own references or feed them into ChatGPT with the prompt “make sure my references aren’t fake”, why should I expect the rest of the paper, code, and datasets the authors created to be anything but the same? It’s easier than ever to generate results and write a paper now but reviewing one carefully (even your own work) still takes hours. If the authors did not take the time to read their own paper, why should I? Papers with hallucinated references should just be desk rejected as they are clearly not ready for acceptance.
What are your top three “an LLM wrote this” tells, ranked by reliability?
Caleb: 1.) The usual LLM phrases like “It’s not X, it’s Y”, “The real X is Y”, bold formatting and em-dashes everywhere; 2.) really dense sentences that are hard to parse; 3.) overly hyping results.
Isaac: Typically I’ve found that the text or paragraphs of the paper will likely have numbers or metrics that disagree with the tables. This happens while authors are writing and then rerun experiments and neglect to prompt the agent to update or verify the results match everywhere. Nobody seems to re-read their own paper. Also Claude freaking loves to condense all possible numbers and details into the text, even details that should be referenced in the appendix or code. When you prompt an agent to write a discussion section it will often just restate metrics from tables verbatim without adding any original thought to it. These are more subtle than the swarms of em-dashes, semicolons, and “It’s not X, but Y”, but they differ too much from paper writing from the years before ChatGPT (B.C.).
How has your reviewing workflow changed because of all this?
Caleb: I’m definitely scanning through the bibliography right after reading the abstract now (and will be using the
bib-audit
skill we made).
Isaac: I spend more time looking for LLM tells just to figure out how to desk reject now. I assume most papers were submitted in bad faith. Everybody loses. I’m looking for the LLM stating it achieved SOTA with some bespoke model while the gains are <1% with no mean or standard deviation over multiple runs; fancy math where a citation would do; ablations that belong in the appendix; prose numbers that do not match the tables. Actually, now that I think of it, not much has changed from reviewing human-slop papers other than hallucinated references. However, the number of decent human-written papers or interesting papers which are enjoyable to read has decreased significantly.
How much of your review time goes to evaluating the science versus hunting for LLM artifacts? Is that sustainable?
Caleb: Past the bibliography check, and a general “does it seem like this is straight LLM output” check, I’m not hunting for LLM artifacts. If I can’t tell that an LLM wrote it, then I don’t particularly care. If you use an LLM to help write your paper, and your paper is interesting with valid (reproducible) experiments, then great! If you submit the output of an LLM, then what are we even doing here?
Isaac: Same as above, I’m looking more for: do I think this entire paper is B.S., or is it interesting enough for the conference it’s submitted to? Nothing profound here, it is just a bad use of my time. There’s even more reviewing now because conferences have made it mandatory to review 4-5 papers if you submit to them. I am basically reviewing slop papers at gunpoint. It feels like living in some weird level of hell.
Be honest: you both use Claude for research and writing. What is the difference between how you use it and what you saw in these submissions?
Caleb: Definitely we do, daily, and this post isn’t saying don’t use LLMs. A lot of the content of this blog was initially drafted by an agent of some sort, but we go through all output extensively: verify, edit, rewrite entirely, etc. In fact, I used Claude to start the lit review at the top of this post, but I read through the papers/posts it surfaced (was surprised by the paper on adversarial attacks on LLM reviewers), organized them in a way that made sense to me, pulled out the source studies from blog posts that Claude was initially citing, and trimmed out a bunch of additional details that I didn’t think were immediately relevant. Learning is part of the fun of writing in the first place!
Isaac: I actually read my paper before I submit, thoroughly, so small details like this don’t come up. If something slips through, then that’s on me, but you’ll have to hunt pretty hard. However, this is what I’ve always done even before LLMs. One trick I have found useful is asking Claude to spot anything ambiguous and then interview me with multiple-choice questions to get clarity so we can be on the same page when it’s assisting my writing which is crucial. For references, I run
bib-audit
, but I still check each one manually with Google Scholar and/or whatever the database is IEEE/CVF etc. Most don’t even need a lookup though because I’ve accumulated a Zotero library of real BibTeX from several years of writing that I can copy/paste from. Getting others’ feedback early is pretty critical and helps to ensure the idea and its presentation are fit for a human to digest. Be okay with completely revamping your writing mid-draft. Ideally it will converge to something sensible way before the deadline – I am sure that rush is what causes these last-minute LLM papers. The big one is having developed some form of research taste pre-LLMs. I cannot imagine shooting in the dark as a young researcher now that LLMs are all the rage.
What does “keeping Claude in check” actually look like in practice for a paper? What do you have to repeatedly correct?
Isaac: Caleb and I have talked about going back and reading pre-LLM vision papers, where the discussion actually sounds like a thoughtful human. Agentic writing is too dense and doesn’t flow well even if you prompt it 100 times to “make my paper flow well and make no mistakes”. Use em-dashes and semicolons sparingly. Fable 5 spams colons rather than em-dashes now and tends to produce marketing B2B SaaS speak for whatever reason, likely because of its generalist training. I think there’s a balance between not being too punchy or bland, because we want people outside the field to actually read our papers, and sounding overly smart isn’t fun to read. One thing that’s worked for me recently is to use a skill that is somewhere in between scientific writing and ASD-STE100 Simplified Technical English.
Caleb: Not much to add on Isaac’s response here. When you’re writing a paper, or blog post, or anything (with or without LLM assistance), then you need to know both what you are trying to say and who your audience is. Once you know those things you need to format your message for that audience. By default Claude and friends have a specific writing style that isn’t well suited for a scientific audience. Have you ever tried drafting an email with an LLM then thought “wait, I don’t want to send that, it doesn’t sound like me” (because the person receiving it will question it just because of the style of the writing)? Pasting LLM output into a scientific manuscript is basically the same thing. If I see a ton of bold formatting, enumerated lists, “marketing speak” like Isaac says, etc. then I question the source of the content just like our hypothetical email recipient would.
Should venues require disclosure of LLM use? Would it change anything?
Caleb: Require disclosures – maybe. Change anything – probably not. Again, if you use an LLM to write an awesome research paper, and the results check out, and it is reproducible, then I’m all for it! What I don’t like is being asked to spend my time providing feedback on something that someone hasn’t spent time on themselves. We need to figure out how to filter that out of the submission pool.
Isaac: I am a heavy skeptic of disclosure requirements. They are useful for good-faith people, but most people will not use them honestly anyway. We added an AI policy to TorchGeo with tiers like “no LLMs,” “LLM-assisted,” and “an LLM did the whole thing,” and nobody is ever going to check that last box and dox themselves. So is it useful? Sure. Will it change anything? Not really. What we actually need is a desk-reject mechanism, or at least a flagging tool that catches the common LLM errors and says whether a paper is even ready for review. Right now reviewers are crowdsourcing the desk rejects, and that is not a fair use of the good-faith reviewers we have left. As a side note, I noticed recently that ICLR is now making author names public on OpenReview which is definitely a choice on how to reel the slop back. I suspect other conferences will follow suit.
Who is most responsible: the authors, their advisors, the ACs and organizers, or the conferences?
Isaac: The boring answer is everyone. The fun answer is the authors, because there is a real discrepancy between good-faith submitters and people just yeeting papers into the OpenReview portal to pad a resume. The incentive structure behind publishing was under scrutiny before LLMs ever showed up (Lipton and Steinhardt were cataloguing mathiness, misuse of language, and misaligned incentives back in 2018). I recall some debate if it’s appropriate for this Kevin Zhu guy and his Algoverse company to scam high-schoolers’ parents to submit something like 500 papers to NeurIPS. The slop needs to stop. We need to be intentional about which ideas are actually valuable and which could have been a blog post. You can turn any class project into a full NeurIPS submission now. The real question is, should you?
Caleb: Authors, don’t submit low effort papers. When Opus 4.6 (I think) and Karpathy’s autoresearch repo first came out I had an autoresearch loop get “state of the art” on the LandCoverAI dataset then had agents write a “paper” about it out of curiosity. It had the look and feel of a research paper, but I definitely wouldn’t submit that anywhere for review as it wasn’t really a research contribution. Advisors/mentors, don’t let your students/mentees submit low effort papers. Organizers, figure out how we can reject low effort papers. This is wayyy easier said than done as the scale of submissions to AI/ML conferences has been truly enormous for a few years already (and continuing the trend, I just saw that AAAI 2027 got 48k abstract submissions).
A skill for auditing bibliographies
Checking a bibliography properly means resolving every entry against the publisher’s deposited record and diffing what the paper printed against what the registrar has – title, year, and every author name, forty-odd times per paper. That work is mechanical and automatable, and it is the part of reviewing we most resent doing for free, so we automated it. The audit we now run ships as bib-audit
, an MIT-licensed Claude Code skill in our new skills marketplace:
Install it from the CLI:
claude plugin marketplace add isaaccorley/skills claude plugin install bib-audit@isaaccorley-skills
Then ask Claude to audit your references and the skill triggers on its own.
Feed it a .bib
, a .bbl
, or the PDF itself; for the latter two, Claude first lifts each rendered reference back into structured fields (un-rendering a bibliography is a language task, not a regex one), and the skill’s scripts do the rest. Every reference is resolved against Crossref, arXiv, DataCite, and Semantic Scholar, diffed field by field against the registrar’s record, and reported worst-first: cited works that do not exist, fabricated identifiers and invented authors, wrong metadata on real papers (truncated author lists, preprint-vs-published year drift), and formatting last. The diff covers full author lists, which matters here because title-only existence checks are exactly what author-swap fabrications like the SatMAE one sail through. The formatting pass is distilled from John Owens’s Common Errors in Bibliographies (worth reading in full once), so the same run that hunts fabrications also flags single-hyphen page ranges, DOIs stored as URLs, and J.D.
-style initials.
It is careful about accusations, too. A reference that prints no DOI or arXiv ID can only be bound by title search, so those findings come back as advisory checks to verify by hand, and the right move in a review is to report the observation – “the first author on the arXiv record is Yezhen Cong, not Yuyang Cong” – and leave motive to the editors. Run it on your own draft before you submit and nobody has to run it on you during review; the .bib
path is read-only and dependency-free and fails only on identifier-pinned mismatches, so it drops into CI as a pre-submission gate. (Not a Claude user? npx skills add isaaccorley/skills
installs the same skill into Codex, Copilot, Cursor, and most other agents.)
Submissions are confidential, bibliographies included, and the audit works by sending pieces of one to a hosted LLM – even though the LLM never writes a word of your review. ECCV 2026’s reviewing policies state that LLMs “are NOT allowed to be used to write reviews or meta-reviews, whether it is run locally or via an API,” and separately bar reviewers from sharing substantial excerpts of a submission with an LLM. WACV’s reviewer guidelines call LLM-generated reviews “highly irresponsible behavior,” sanctionable by desk rejection of the reviewer’s own papers, and their confidentiality rules forbid showing a submission’s material to anyone who is not a reviewer – which a hosted LLM is not. NeurIPS’s LLM policy restricts what reviewers can share with LLM services; its AI-assisted reviewing experiment is the sanctioned route.
References
Verified by humans :)
- Miryam Naddaf and Elizabeth Quill (2026). “Hallucinated citations are polluting the scientific literature. What can be done?” Nature 652:26–29.
- Zhenyue Zhao, Yihe Wang, Toby Stuart, Mathijs De Vaan, Paul Ginsparg, and Yian Yin (2026). “LLM hallucinations in the wild: Large-scale evidence from non-existent citations.” arXiv:2605.07723.
- Maxim Topaz, Nir Roguin, Pallavi Gupta, Zhihong Zhang, and Laura-Maria Peltonen (2026). “Fabricated citations: an audit across 2.5 million biomedical papers.” The Lancet 407(10541):1779–1781.
- Samar Ansari (2026). “Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025.” arXiv:2602.05930.
- Bradley Emi (2025). “Pangram Predicts 21% of ICLR Reviews are AI-Generated.” Pangram Labs blog, November 18, 2025.
- Lin Li, Qi Zhang, Xander Davies, Jianing Qiu, and Yarin Gal (2026). “Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community.” arXiv:2606.10159.
- Claudine Gartenberg, Sharique Hasan, Alex Murray, and Lamar Pierce (2026). “More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review.” Organization Science 37(3):795–812.
- Gautam Kamath (2026). “On Violations of LLM Review Policies.” ICML blog, March 18, 2026.
- Zachary C. Lipton and Jacob Steinhardt (2019). “Troubling Trends in Machine Learning Scholarship: Some ML papers suffer from flaws that could mislead the public and stymie future research.” Queue 17(1):45–77.
Venue policies:
- NeurIPS: 2025 LLM policy, 2026 Main Track Handbook, workshop guidance, AI-assisted reviewing experiment, and the blog post on desk-rejecting AI-generated position papers
- WACV: author and reviewer guides for 2026 and 2027, and the 2026 reviewer guidelines
- ECCV 2026: What’s New
Tools:
bib-audit
, the MIT-licensed reference-audit skill from this post, in the isaaccorley/skills marketplace- John Owens’s Common Errors in Bibliographies, the source of the audit’s formatting rules
Citation
@online{robinson2026,
author = {Robinson, Caleb and Corley, Isaac},
title = {Q\&A from the Slop Trenches},
date = {2026-07-30},
url = {https://geospatialml.com/posts/reviewing-ai-slop/},
langid = {en}
}