Bastian Rieck has had enough.
The mathematician and machine learning researcher, currently leading a research group at the University of Fribourg in Switzerland, published a blistering essay this month titled simply “BS” — a deliberate provocation aimed at what he describes as the systematic rot eating through the machine learning research community from the inside. His argument is not that the field lacks talent or ambition. It’s that the incentive structures governing academic publishing have created a self-reinforcing cycle of mediocrity, where volume is rewarded over substance and genuine scientific progress is buried under an avalanche of incremental, poorly motivated papers.
The essay, published on Rieck’s personal blog, doesn’t mince words. He describes a research culture in which the primary objective is no longer discovery but production — of papers, of benchmarks, of marginal improvements on existing methods that are packaged as breakthroughs. “The field,” Rieck writes, “is drowning in papers that no one reads, addressing problems that no one has, using methods that no one needs.”
That’s a provocative claim. But Rieck is far from alone in making it.
The sheer scale of machine learning publication has become its own kind of problem. The NeurIPS 2024 conference received over 15,000 submissions. ICML and ICLR aren’t far behind. These numbers have roughly doubled or tripled in just five years. The peer review system, already strained before the deep learning boom, is now operating well past its breaking point. Reviewers are overworked, under-qualified for the specific papers they’re assigned, and often working on their own submissions to the same venue simultaneously. The result, as Rieck and others have argued, is a review process that functions more like a lottery than a quality filter.
Rieck’s critique goes deeper than the logistics of peer review, though. His core contention is about intellectual honesty — or the lack of it. He points to a pattern he’s observed repeatedly: researchers identify a benchmark, engineer a method that achieves a marginal state-of-the-art result on that benchmark, then write a paper claiming the improvement validates their approach. The theoretical motivation is often thin. The experimental comparisons are often selective. And the broader question — does this actually matter? — is almost never asked, let alone answered.
This is the BS in the title. Not outright fraud, though that exists too. Something more insidious: a culture where everyone tacitly agrees to pretend that small numbers going up on a leaderboard constitutes scientific progress.
The Benchmark Trap and the Incentives That Feed It
The benchmark problem is well-documented at this point. Researchers have noted for years that common ML benchmarks — ImageNet, GLUE, SuperGLUE, various graph classification datasets — have become targets rather than measures. When a benchmark becomes the goal, Goodhart’s Law kicks in: the measure ceases to be a good measure. Models get optimized for the specific idiosyncrasies of a dataset rather than for the underlying problem the dataset was supposed to represent.
Rieck highlights this dynamic with particular frustration. In his area of topological machine learning, he’s watched papers claim advances based on tiny datasets with known statistical issues, where the variance between runs often exceeds the claimed improvement. The papers get accepted anyway. The cycle continues.
And the incentives are clear. Junior researchers need publications to get jobs. Senior researchers need publications to get grants. Universities need publications to climb rankings. Conferences need submissions to justify their existence and their registration fees. Everyone benefits from more papers — except, arguably, science itself.
This isn’t a new observation. In 2020, the “Troubling Trends in Machine Learning Scholarship” paper by Zachary Lipton and Jacob Steinhardt cataloged many of the same pathologies: failure to distinguish between explanation and speculation, mathiness (the use of formal notation to obscure rather than clarify), and inadequate consideration of alternative hypotheses. That paper was widely discussed. Not much changed.
More recently, the conversation has intensified as generative AI tools have made it even easier to produce paper-shaped objects. Large language models can now generate plausible-sounding related work sections, smooth over logical gaps in argumentation, and produce grammatically perfect prose that says remarkably little. Several conferences have begun implementing AI-detection tools, but the arms race is already lopsided. The tools for generating text are improving faster than the tools for detecting it.
Rieck’s essay touches on this too, though his primary concern isn’t AI-generated papers per se. It’s the broader intellectual climate that makes such papers possible — a climate where the form of a paper matters more than its content, where hitting the right stylistic notes and citing the right prior work can substitute for genuine insight.
The problem extends beyond academia into industry. As companies race to publish at top venues to attract talent and signal technical credibility, the volume of corporate-authored ML papers has surged. Some of this work is excellent. Much of it is not. And the presence of well-resourced industry labs in the submission pool creates its own distortions: they can run experiments at scales that academic labs cannot, making it harder for university researchers to compete on benchmarks even when their ideas are more interesting.
So what’s the solution? Rieck doesn’t claim to have a comprehensive one, and he’s honest about that. But he gestures toward several directions. Slower, more careful review processes. A cultural shift away from paper counting as the primary metric of academic success. More emphasis on replication and negative results. Greater willingness to say, publicly and without embarrassment, that a line of research isn’t working.
These are not new proposals either. The question is whether the field has the collective will to implement them when every individual actor benefits from the status quo.
There are some signs of movement. The Machine Learning Reproducibility Challenge, now in its sixth year, encourages researchers to verify the claims made in accepted papers. Some venues have experimented with longer review cycles, author response periods, and ethics reviews. A handful of prominent researchers have publicly committed to publishing less and thinking more — though whether this survives the next grant cycle remains to be seen.
The broader context matters here too. Machine learning is not just an academic discipline anymore. It’s a multi-trillion-dollar industry driver. The pressure to publish is entangled with the pressure to commercialize, to hype, to attract venture capital and government funding. When a field sits at the intersection of science and money on this scale, the incentives for BS — in Rieck’s sense of the word — become enormous.
Rieck’s essay is ultimately a plea for seriousness. Not solemnity. Not gatekeeping. Seriousness — the kind that asks whether a piece of work actually advances understanding, whether an experiment actually tests what it claims to test, whether a result actually means what the authors say it means. These are basic scientific questions. The fact that they need to be reasserted so forcefully suggests just how far the field has drifted from its foundations.
Will the essay change anything? Probably not on its own. Essays rarely do. But it joins a growing chorus of voices — from researchers, reviewers, and even some conference organizers — arguing that the current trajectory is unsustainable. The machine learning research community is producing more than it can consume, review, or build upon. At some point, that debt comes due.
The question isn’t whether the system is broken. It’s whether anyone with the power to fix it has sufficient incentive to try.


WebProNews is an iEntry Publication