Developers Now Refuse to Code Without AI. The Data Says It Slows Them Down

Developers now decline assignments without AI coding tools, yet METR's controlled study found experienced programmers took 19% longer with them. Companies like Amazon and Uber see costs soar without productivity gains. The perception-reality gap widens as maintenance burdens grow. This paradox challenges engineering leaders to measure true impact beyond self-reported speed.
Developers Now Refuse to Code Without AI. The Data Says It Slows Them Down
Written by Eric Hastings

Software engineers have drawn a line. Many simply will not accept assignments that bar them from using AI coding assistants. The stance emerged clearly this spring when researchers at METR attempted to rerun a controlled productivity trial. Participants declined. They would not work without AI, even for a limited set of tasks in a study setting.

The refusal, reported first by TechCrunch on May 29, 2026, underscores a profound shift. Developers have grown so accustomed to tools like Cursor, GitHub Copilot and Claude-powered editors that stepping back feels intolerable. Yet the very study they helped make impossible delivered uncomfortable news in 2025: experienced programmers took 19 percent longer to finish real issues from their own open-source repositories when allowed to use frontier AI models.

METR’s randomized controlled trial recruited 16 seasoned contributors to large projects. These developers averaged years of familiarity with their codebases. Each supplied genuine tasks — bug fixes, small features, refactors. Researchers randomly assigned conditions. One group worked as usual. The other could call on AI. Screen recordings and self-reported times told the story. AI users spent extra minutes steering prompts, reviewing output, fixing hallucinations and waiting for suggestions to generate. The net effect was slower completion. METR published the findings in July 2025.

The perception gap stunned observers. Before starting, participants predicted AI would accelerate them by 24 percent. Afterward, even though objective data showed the opposite, they still estimated a 20 percent speedup. That disconnect has only widened. A May 2026 METR survey found developers now report feeling twice as valuable with AI. Self-reported sentiment and hard measurement continue to diverge.

Companies have noticed. Amazon created an internal leaderboard called KiroRank to track employee AI token consumption. The goal was to encourage adoption. Instead it produced tokenmaxxing — engineers assigning trivial work to AI agents simply to climb the ranks. Costs ballooned with no corresponding lift in output. Amazon shut the experiment down. “Don’t use AI just to use AI,” a spokesperson told reporters. The episode, covered yesterday by Yahoo Finance, illustrates how proxy metrics can mislead.

Uber burned through its entire 2026 AI budget in four months. Productivity metrics barely budged. Salesforce reportedly anticipates spending $300 million on Anthropic tokens this year alone. These figures come from The Next Web‘s reporting published hours ago. The pattern repeats across enterprises. Heavy investment meets flat or declining throughput once downstream effects surface.

Review times tell part of the tale. One analysis found pull requests increased 98 percent with AI assistance. Yet review duration jumped 91 percent. More code arrives, but much of it demands extra scrutiny for subtle errors, duplicated logic or missing edge cases. Faros AI documented the dynamic in its 2025 engineering report. Teams generate volume. They do not necessarily accelerate delivery.

Maintenance costs compound the problem. AI often produces code that functions in the moment but accrues technical debt. James Shore, a longtime programmer and author, offered a blunt assessment in The Next Web article. “You write code twice as quick now? Better hope you’ve halved your maintenance costs… Otherwise, you’re screwed. You’re trading a temporary speed boost for permanent indenture.” The warning lands harder as organizations discover that AI-generated snippets require sustained human oversight.

Specialized tools reveal the same pattern. Entelligence AI, a reliability engineering startup, found that 44 percent of AI tokens in some environments went toward bug fixes. CodeRabbit, itself an AI code-review product, determined that AI-produced code created 1.7 times more problems than human-written contributions. Singapore Management University researchers reached similar conclusions in independent work cited by The Next Web. The tools accelerate initial drafting. They shift effort downstream.

But developers refuse to relinquish them. The dependency has become cultural. New graduates expect AI autocomplete as table stakes. Mid-career engineers describe flow states that evaporate without intelligent suggestions. Even researchers cannot recruit control groups anymore. This reality, highlighted in both The Next Web and TechCrunch coverage, complicates future measurement. How do you quantify impact when the baseline — coding without AI — has vanished from willing participants?

Some organizations respond by tightening guardrails. They route prompts through intermediary layers to control cost and quality. Others invest in better evaluation frameworks that track not just lines written but defects introduced, review cycles consumed and long-term ownership burden. A few have begun experimenting with AI usage caps on critical paths. The moves acknowledge a central tension: individual velocity feels higher. System velocity does not.

Recent coverage reinforces the complexity. A HackerRank analysis published in December 2025 applied Jevons’ paradox to software. Efficiency gains expand demand for developers rather than reduce headcount. Gradle’s Trisha Gee, writing in November 2025, described the “developer productivity paradox” at the DPE Summit. Engineers crank out more code. Organizational metrics — deployment frequency, lead time, change failure rate — refuse to improve in lockstep. The flood of output overwhelms testing, integration and operations teams.

Stack Overflow’s 2025 developer survey captured the sentiment split. Eighty-four percent of professionals use or plan to use AI tools. Only 16 percent say the assistance makes them substantially more productive. The largest cohort reports modest or negligible gains. Those numbers align with METR’s controlled data. Feeling productive and delivering faster are not the same.

So what explains the stubborn attachment? Part of it is genuine help on rote tasks. AI excels at boilerplate, test stubs and simple translations. Juniors gain confidence drafting initial implementations. Even veterans appreciate the sounding board for exploring approaches. But the same mechanism that aids exploration can seduce users into over-reliance. Prompt engineering becomes its own cognitive load. Context switching between human intent and machine output consumes attention that might otherwise go to architecture or edge-case thinking.

And the economic signals confuse the picture. Venture funding flows toward AI coding startups. Job postings tout “AI-first” workflows. Compensation for prompt-savvy engineers rises in some markets. These incentives push adoption even when rigorous evidence remains mixed. Amazon’s decision to kill its leaderboard and Uber’s rapid budget exhaustion suggest finance teams are beginning to ask harder questions. Engineering leaders must answer them with something better than self-reported velocity.

The industry now sits at an inflection. AI coding assistants are ubiquitous. Their presence reshapes hiring, training and daily practice. Yet the productivity case stays unsettled. Controlled trials show slowdowns for experienced hands on familiar code. Enterprise rollouts reveal hidden costs in review, maintenance and cloud spend. Developers, meanwhile, vote with their keyboards. They will not go back.

Resolving the paradox will demand sharper measurement. Teams need telemetry that captures full-cycle outcomes — not just completion time but defect density, onboarding speed for new contributors, and total cost of ownership over quarters. They must distinguish between tasks where AI shines and those where human judgment remains irreplaceable. Most of all, they must resist the temptation to optimize for tokens consumed or pull requests merged. Those proxies failed Amazon. They will fail others.

Until better data arrives, the refusal to work without AI will likely harden. The tools have become part of the craft. Whether they ultimately raise the profession’s output or simply redistribute effort remains the open, expensive question facing technology leaders in 2026.

Subscribe for Updates

DevNews Newsletter

The DevNews Email Newsletter is essential for software developers, web developers, programmers, and tech decision-makers. Perfect for professionals driving innovation and building the future of tech.

By signing up for our newsletter you agree to receive content related to ientry.com / webpronews.com and our affiliate partners. For additional information refer to our terms of service.

Notice an error?

Help us improve our content by reporting any issues you find.

Get the WebProNews newsletter delivered to your inbox

Get the free daily newsletter read by decision makers

Subscribe
Advertise with Us

Ready to get started?

Get our media kit

Advertise with Us