Skip to main content

Publishing Rigor: What Engineering Fields Owe Mathematics

Picture

Member for

1 year 2 months
Real name
SIAI Editor
Bio
SIAI Editor

Modified

Fabricated citations rose twelvefold in three years
Math enforces rigor; engineering never matched it
Reproducibility checks could close the gap

A biomedical audit released in 2026 found that one in every 277 papers indexed on PubMed in the first weeks of the year cited a source that does not actually exist. Three years earlier, the same audit had found roughly one in 2,828. The fabrication rate rose more than twelvefold in under three years, far beyond a simple doubling or tripling and the acceleration began right around when large language models got cheap enough for any lab to lean on under deadline pressure. Mathematics departments would not have absorbed a shift like that quietly. A false citation in a formal proof gets caught the moment somebody tries to build on it because the next step simply will not close. Engineering and applied computing fields carry no equivalent trip wire and publishing rigor, the standard that decides what counts as verified before it reaches print, has been drifting apart across disciplines for years now. The gap is no longer subtle enough to ignore.

A Discipline Built to Catch Its Own Errors

Mathematics enforces its standard structurally rather than through good intentions alone. A proof either holds under a formal system or it does not and reviewers in the field are trained specifically to hunt for the step that does not follow, not merely to judge whether a result sounds plausible. That structure carries real costs: mathematical papers move slowly through review and a controversial or unusually long proof can sit under scrutiny for years before it's accepted. But the slowness buys something real. Once a paper claims a new theorem, the claim can in principle be checked by anyone with the training to follow the argument line by line and that checkability is largely why a math result, once accepted, tends to stay accepted.

Engineering and computer science operate under a different economy entirely. A growing share of machine learning results depend on benchmark numbers that only the original authors can fully reproduce, since the exact training data, hyperparameters and random seeds behind a reported score often aren't disclosed in full. NeurIPS submissions grew from roughly 3,240 in 2017 to more than 21,000 by 2025 and reviewer capacity was never expanded anywhere close to that pace.

Figure 1: Submissions climbed from 3,240 in 2017 to 21,575 in 2025.

In a widely circulated 2021 experiment, the same batch of accepted NeurIPS papers was resubmitted through a second, independent review process and close to half of the papers accepted the first time were rejected on the second pass. That gap describes a review system where the outcome for any single paper carries a large element of chance, not a story about a few careless reviewers and chance is close to the opposite of what a genuinely rigorous field is supposed to guarantee.

Figure 2: 50.6% of accepted papers would have been rejected on a second, independent review.

Where the Fabrication Actually Lives

The contrast sharpens further once the citation-fabrication numbers are examined directly. A Lancet audit that scanned more than 2.5 million biomedical papers and 125 million references found the fabrication rate held roughly flat, near four per 10,000 papers, through 2023, then climbed to 51 per 10,000 by the final quarter of 2025 and 57 by early 2026. Ninety-eight percent of the flagged papers were still uncorrected in the literature at the time of the audit. A separate analysis out of Tübingen scanned 1.19 million PubMed full texts and found that by December 2025, 89 percent of papers showed excess use of vocabulary strongly associated with large language model writing, concentrated most heavily in Discussion sections, where authors interpret and extend their own results rather than simply report raw measurements.

Fabricated citations and machine-generated prose are not unique to engineering and cancer research in particular has faced its own reckoning with templated, paper-mill-style submissions. But the two fields differ sharply in what a fabrication actually costs the reader. A biomedical paper with a fake citation can still report a real experiment underneath it; the false reference is a flaw layered onto findings that exist independently. A machine learning paper's central claim, though, is frequently the benchmark number itself and when that number cannot be reproduced or the citation trail behind the method cannot be verified, there is often no independent finding left to salvage underneath it. Engineering fields inherited computer science's culture of rapid iteration and light-touch review right at the moment tools for generating plausible, ungrounded text became free and instant and that combination has done more damage there than the raw fabrication numbers alone would suggest.

What Adopting Mathematics' Standard Would Actually Require

Borrowing the math department's model does not mean asking engineering conferences to slow every submission down to a multi-year proof review. It means separating two things current review tends to conflate: whether a result is interesting and whether a result is verified. Mathematics keeps those judgments distinct almost by habit, since a proof's correctness gets assessed on its own terms before anyone weighs in on significance. Engineering venues could adopt a version of that split by requiring a reproducibility check, actual code and data sufficient for an independent party to regenerate the reported numbers, as a precondition for acceptance instead of an optional appendix most reviewers never open. Several ICML and NeurIPS working groups have already piloted reproducibility badges along these lines but the badges remain voluntary and voluntary standards tend to be the first ones eroded under submission-volume pressure.

A related requirement follows from the citation data specifically. Verification tools capable of flagging a fabricated or fabricated-looking reference already exist, since much the same audit techniques behind the Lancet numbers could, in principle, be run automatically at the point of submission. Building that check into conference and journal pipelines would not eliminate fabrication outright but it would catch the crudest and most common form of it before publication, not years afterward through outside detective work. Mathematics does not really face this problem in the same form since a fabricated citation in a proof doesn't save the author any real labor. Engineering fields, where a citation can stand in for background work an author never actually verified, lack that same natural disincentive, so the check has to be built rather than assumed.

The most common objection to all this is that mathematics can afford its slowness because a wrong proof rarely costs anyone much beyond wasted reading time, whereas engineering fields solve problems where speed itself carries value; a faster training method or a better diagnostic model helps people sooner if it reaches the field sooner. There's real weight to that. But the current alternative is not actually faster in any way that matters, since a benchmark result that cannot be reproduced does not help downstream users who build on it and only later discover the foundation was never solid. Speed that produces retractable work only looks fast. The real cost gets deferred, paid later by whoever tried to build on the unverified claim.

A second worry is that mathematics benefits from a smaller, more tightly networked community where informal social accountability substitutes for formal process, a mechanism that presumably will not scale to a field publishing tens of thousands of papers a year. That's a fair structural point and it argues for automation rather than for giving up on the goal. The reproducibility checks and citation-verification tools described above exist specifically because informal accountability stopped scaling once submission volume passed a certain point. Mathematics never needed automated verification because its community size kept informal accountability functional on its own; engineering fields lost that option years ago, which is exactly what makes automated verification the substitute rather than some optional extra.

One more objection raised most often by editors at applied and industry-facing venues holds that requiring full reproducibility would penalize legitimate work built on proprietary data or systems that can't be shared for competitive reasons. This concern is real, though narrower than it sounds. A reproducibility requirement can apply to the verifiable claims within a paper, the parts of a method that do not depend on proprietary inputs, without demanding disclosure of trade secrets themselves. Journals in other applied fields already draw this line routinely, requiring enough methodological detail for independent verification of an approach even when the underlying dataset stays private.

The Standard the Data Now Demands

A twelvefold rise in fabricated citations within three years signals something well beyond statistical noise in the scientific record. It points to engineering and applied computing fields adopting the speed of modern research without adopting the verification habits that make that speed survivable elsewhere in science. Mathematics arrived at its standard less through caution for its own sake than through necessity, since a field built entirely on chained logical steps cannot tolerate an unverified link anywhere in the chain and it built a review culture to match that constraint over time. Engineering fields now face a comparable moment, not because every result needs a formal proof but because the tools generating plausible, ungrounded claims have gotten too fast and too cheap for informal trust to keep functioning as a stand-in for verification. The fix does not require reinventing peer review, only a shift in what counts as a precondition: engineering venues need to treat reproducibility and citation integrity the way mathematics has always treated a valid proof, as requirements rather than optional virtues. The data are already public and the twelvefold curve is still climbing.


This article reflects the analytical judgment of The SIAI Editorial Board and does not constitute policy advice or the official position of any affiliated institution.


References

Beygelzimer, A., Dauphin, Y., Liang, P. and Wortman Vaughan, J. (2021) 'The NeurIPS 2021 consistency experiment', Neural Information Processing Systems Blog.
Holzwarth et al. (2026) 'Most biomedical publications show signs of LLM-assisted writing', arXiv preprint arXiv:2608.10715.
Kobak, D., González-Márquez, R., Horvát, E-Á. and Lause, J. (2025) 'Delving into LLM-assisted writing in biomedical publications through excess vocabulary', Science Advances, 11(27), eadt3813.
Mascáto Fontaína, N. et al. (2025) 'Identifying common patterns in journals that retracted papers from paper mills', Research Integrity and Peer Review, 10(21).
Retraction Watch (2026) 'One in 277 PubMed-indexed papers in 2026 shows fabricated references, says analysis'. New York: The Center for Scientific Integrity.
Su, B., Zhang, J., Collina, N., Yan, Y., Li, D., Cho, K., Fan, J. and Su, W. (2026) 'Recommending best paper awards for ML/AI conferences via the isotonic mechanism', arXiv preprint arXiv:2601.15249.
Topaz, M., Roguin, N., Gupta, P., Zhang, Z. and Peltonen, L-M. (2026) 'Fabricated citations: an audit across 2.5 million biomedical papers', The Lancet, 407, pp. 1779–1781.

Picture

Member for

1 year 2 months
Real name
SIAI Editor
Bio
SIAI Editor