An autonomous car barrels down a simulated city street. Suddenly, an elephant lumbers into the crosswalk. The vehicle doesn’t brake. It veers right into the beast, mowing it down.
That’s no nightmare from a sci-fi flick. It’s footage from Futurism, showcasing Fail2Drive, a new testing regime that throws outlandish—but telling—obstacles at self-driving models. Researchers at the University of Tübingen’s Autonomous Vision Group unveiled the benchmark this week, and it lays bare a core weakness in the field: AI systems ace familiar roads but crumble when chaos intrudes.
Andreas Geiger, head of the group, put it bluntly in a LinkedIn post. “Why did the elephant cross the road? To expose how fragile your model is.” He argues that most models don’t memorize data points exactly. No. They memorize entire scenarios. “What looks like strong benchmark performance may just be strong memorization.”
Fail2Drive plugs into CARLA, the open-source simulator that powers much of the industry’s development work. It pairs normal routes with twisted versions—same path, but now with unseen surprises. Think playground slides blocking lanes. Firetrucks parked dead center. Walls painted to mimic open road, Looney Tunes style.
The results? Seven state-of-the-art models, including SimLingo, HiP-AD, and Orion, saw success rates plunge an average 22.8%. Some dropped more. Way more. That’s from the benchmark’s GitHub repo, built by Simon Gerstenecker, Geiger, and PhD student Katrin Renz. Their preprint paper, posted to arXiv, details 100 route pairs across varied townscapes, 17 fresh scenarios, and 30 novel assets like stray animals and visual tricks. No training on these allowed. Just pure reaction.
Renz shared the damning video on X, where it racked up thousands of views. A car halts before a slide—then accelerates straight into it. Another plows a firetruck at full tilt. Elephants get no mercy. “Performance drops by 22.8% on average,” she noted in a follow-up. “Instead of learning robust driving principles, the models rely heavily on shortcut learning tied to their training sets.”
This hits close to home for companies like Tesla. Its Full Self-Driving software, now in version 14 and eyeing robotaxi deployment, trains on billions of fleet miles. Yet real-world clips show persistent blind spots. Tesla cars have misread trains as trucks, veered into rail yards, as captured in older Futurism reports. AVs have struck animals repeatedly, from ducks to deer. Fail2Drive mimics that unpredictability in code, asking: Can your neural net generalize, or does it fold at the first pachyderm?
Industry insiders know the drill. CARLA Leaderboard scores dazzle—90% plus on standard Leaderboard runs. But swap in out-of-distribution noise? Crashes galore. Geiger’s team built a toolbox for custom tests, too. Design your own elephant stampede online. Watch models fail spectacularly.
Critics in comments on Geiger’s post push back. Raul Sena Ferreira called outlandish objects a distraction. Focus on realistic perturbations inside the design domain, he said, citing his 2022 IEEE work on genetic algorithms for vuln hunting. Fair point. Elephants rarely roam freeways. But rare events kill. A parked firetruck? That’s daily heroism gone wrong.
Tesla sidesteps simulators like CARLA, betting on end-to-end vision from real data. Elon Musk touts FSD’s safety stats—one crash per 5.3 million miles, nine times better than U.S. averages. Fleet learning should conquer tails, right? Yet Tesla’s history whispers otherwise. In 2022, critics like Dan O’Dowd released videos of FSD Beta smashing child mannequins, prompting takedown demands from the company. Recent X clips show supervised FSD dodging debris in Austin ride-hails, but unsupervised? Robotaxi unveilings loom, with promises of no drivers by year’s end.
But. Fail2Drive doesn’t test Tesla directly. CARLA models are research prototypes, often vision-plus-planning hybrids. Tesla’s black-box neural net might fare differently. Or not. The benchmark screams for proprietary runs. Wayve, where Renz interned, fuses video with maps—could that blunt the drop?
Broader implications ripple. Regulators watch. NHTSA probes Tesla crashes. IIHS rates Model 3, Y tops for passive safety, but active systems? Fail2Drive quantifies the gap between lab polish and street grit. Models overfit benchmarks. They shortcut. Humans improvise.
Short punch: Fix this, or elephants win.
Researchers offer hope. Paired routes isolate failures precisely—visual shifts tank perception; behavioral ones hit planning. Toolbox lets anyone probe. Open-source ethos could harden the field. Geiger’s group, creators of nuScenes dataset, keeps pushing.
As AVs chase Level 4, from Tesla robotaxis to Waymo vans, Fail2Drive stands as warning. Streets teem with the weird. Ducks. Slides. Painted traps. Elephants, metaphorically. Current AIs memorize highways. They don’t drive them. Not yet.


WebProNews is an iEntry Publication