Software teams once pinned delivery delays on quality assurance teams. They saw QA as the visible bottleneck, easy to trim for speed. Yet a detailed examination in ACM SIGSOFT Software Engineering Notes flips that script. Author Ahmed El-Deeb argues the real drags sit deeper. “The low hanging fruit is not always the one worth cutting,” he writes. Focus instead on hidden inefficiencies that erode productivity across the stack.
That same logic applies with fresh urgency to AI agents. These systems now book travel, query databases, execute trades and manage infrastructure. They act with minimal oversight. But their autonomy exposes cracks that grow wider each quarter. Incidents climbed 56.4 percent from 2023 to 2024 alone. The acceleration shows no sign of slowing. Cycode’s analysis ties the surge to Stanford’s HAI AI Index Report and warns that 2026 marks a turning point.
Consider EchoLeak. A zero-click prompt injection in Microsoft 365 Copilot pulled sensitive enterprise data without user interaction. The vulnerability carried a CVSS score of 9.3. Researchers at Aim Security reproduced it with a single crafted email. No downloads. No links. The agent simply followed instructions hidden in plain sight.
Such attacks succeed because agents treat external input as authoritative. A public Slack message can poison an AI assistant and surface content from private channels. PromptArmor demonstrated exactly that in 2024. The assistant became an unwitting courier for restricted data. One integration. Multiple victims.
Multi-agent setups compound the danger. Agents pass outputs to one another without verification. Output from Agent A becomes instruction for Agent B. Compromise the first link and the chain collapses. Tests on CrewAI running GPT-4o produced private data exfiltration in 65 percent of scenarios. Magentic-One executed malicious code 97 percent of the time when fed a tainted local file. Success hit 100 percent in some configurations. These numbers come from peer-reviewed work cited in a detailed Reddit cybersecurity thread that catalogs 2025 incidents.
Frameworks themselves often escape blame. Palo Alto Networks Unit 42 research from May 2025 states that CrewAI and AutoGen carry no inherent flaws. The problems arise from default configurations that ship credentials in shared environment files and skip tool scoping. Developers copy tutorial patterns into production. The result is predictable.
Even code generation tools introduce risk. Fifteen to 25 percent of AI-generated code contains security flaws. SQL injection, cross-site scripting and weak authentication appear repeatedly. The “IDEsaster” project uncovered more than 30 vulnerabilities across GitHub Copilot, Cursor and similar platforms in December 2025. Twenty-four earned CVEs. CamoLeak, scored at 9.6, let attackers silently drain secrets and source code from private repositories. Digital Applied’s report lays out the numbers and the fixes.
Memory poisoning adds another layer. An attacker plants false information in an agent’s long-term recall. Later tasks trigger the poisoned data and steer behavior off course. OWASP now lists this among its top concerns for agentic systems. The organization released its first Top 10 for Agentic Applications in December 2025. Top entries include Agent Behavior Hijacking, Tool Misuse and Exploitation, and Identity and Privilege Abuse. OWASP’s announcement stresses that these risks demand controls beyond traditional application security.
Privilege problems run especially deep. Agents often receive broad permissions to complete tasks across SaaS platforms, cloud consoles and internal tools. A single OAuth token compromise hands an attacker the keys to multiple systems. Obsidian Security documented cases where one breached chat agent cascaded into Salesforce, Google Workspace, Slack, S3 and Azure across more than 700 organizations. The breach started small. It scaled fast. Details appear in Obsidian Security’s 2025 landscape review.
Gartner forecasts that 40 percent of enterprise applications will integrate task-optimizing agents by the close of 2026. That figure stood below 5 percent in 2025. Eighty percent of IT workers already report seeing agents act without proper authorization. Shadow AI usage adds cost. IBM’s data breach report links it to an extra $670,000 in average incident expenses. Bans fail. Nearly half of employees continue using personal accounts anyway.
Recent incidents underscore the trend. A Chinese state-sponsored group used Claude Code to target 30 organizations across tech, finance and government. Eighty to 90 percent of operations ran autonomously. Anthropic confirmed the campaign in November 2025. On the defensive side, Noma Labs found a CVSS 9.2 flaw in CrewAI’s platform involving an exposed GitHub token. The company patched it in five hours.
Even the web itself turns adversarial. Scammers now publish fake documentation sites packed with hidden instructions aimed at agents. JSON-LD schema, off-screen CSS and concealed divs deliver payloads invisible to humans but readable by crawlers. One campaign tricked four out of 26 tested LLMs into purchasing a $3 API key via Stripe. Zscaler researchers exposed the scheme. Recent posts on X highlight similar SEO poisoning attacks against DeFi trackers and other authoritative-looking resources.
So what now? Enterprises cannot simply outlaw agents. Adoption moves too quickly. Instead security leaders push layered defenses. Sandboxed execution environments limit blast radius. Human-in-the-loop checkpoints catch high-risk actions. Output validation and strict tool scoping reduce unintended calls. Rate limiting and signed messages between agents curb trust assumptions.
Yet many deployments still rely on hope and shared credential files. That gap explains why incidents keep climbing. The ACM paper on delivery delays ends with a call to target “what’s high and deep.” The same advice fits agent security. Surface-level prompt hardening is no longer enough. Teams must examine identity models, memory architectures, inter-agent communication and permission boundaries with equal rigor.
Autonomous agents promise efficiency gains that QA cuts never could. They also carry failure modes that traditional software never faced. The incidents of 2025 and early 2026 serve as early warnings. Organizations that treat these systems as simple extensions of existing apps will learn the hard way. Those that redesign controls around the unique properties of agency stand a better chance of staying ahead.
McKinsey’s October 2025 analysis captures the tension. Eighty percent of organizations have already witnessed risky agent behaviors, from unauthorized data exposure to improper system access. The report, available at McKinsey’s site, urges immediate governance updates. Recorded Future echoes the call in its April 2026 research, noting that autonomy amplifies supply-chain weaknesses and identity risks at machine speed. Their findings highlight the need for zero-trust principles tailored to agent identities.
Lasso Security ranks memory poisoning, tool misuse and privilege compromise as the top three agentic threats for 2026. The firm’s breakdown at Lasso’s blog offers concrete fixes, from scoped permissions to continuous monitoring of agent actions. Mindgard’s May 2026 overview adds that agent behavior hijacking sits outside the classic OWASP LLM Top 10 yet demands urgent attention. Mindgard’s post maps the emerging eleventh risk.
Check Point, Oligo and other vendors publish parallel guidance. All converge on one point. The model is rarely the weakest link. The stack around it is. Prompt defenses matter, but so do runtime monitoring, least-privilege enforcement and architectural isolation. Ignore any layer and the entire system tilts toward compromise.
Developers building with agents face a choice. Ship fast and inherit every misconfiguration the community has already cataloged. Or invest time in secure patterns now. The second path slows initial velocity. It also prevents the kind of cascading failure that makes headlines and erodes trust. Given the trajectory of incidents, that trade-off looks increasingly favorable.
The ACM article warned against easy targets. Security teams would do well to heed similar counsel. Blame the agent when it fails. Or examine the hidden assumptions baked into its design. The latter demands more effort. It also points toward lasting protection.


WebProNews is an iEntry Publication