Three researchers just answered a question nobody asked them to: what happens when you hand a real car over to the chatbots living on your phone. DrivingBench, built by Aditya Ramabadran, Simon Mahns, and Tobias Gessler, put Claude, GPT, Grok, and GPT-6 Astra behind the wheel of a 2022 Toyota Corolla and let them drive a 134.7-meter parking-lot course marked out with cones. No simulator, no safety net beyond a human hovering a foot over the brake pedal.
The results are rough. Across 11 total runs split between four models, eight failed to cover more than 11% of the course before crashing out or stalling near the first turn. Only one model, OpenAI’s GPT-6 Astra, finished the route start to finish, and it needed a second attempt to do it.
Grok 4.6 ran through the course three times and never got past 11% of it. On its first attempt, the model spotted a gap between the mini-cones marking the course boundary and decided it was a gate meant to be driven through. It drove straight off the intended path within two commands. Grok’s other runs weren’t much better, landing at 8% and 10% completion, never clearing more than about 22 meters before stopping.
One GPT model, running as GPT-5.6 Sol, had its own breakdown in logic. It invented a rule that cones on one side of the lane were all the same color, something the original command had explicitly told it was not true. It flatlined at 6% completion across all three attempts. Claude’s Fable 5.1 fared better on paper, climbing from 9% to 10% to 45% completion by its third try, but it still never finished. Across the board, the agents struggled to calculate how much steering angle a turn actually required, so they either understeered into the cones or overcorrected and ran wide.
The car moved in short bursts, capped at 3.5 meters per second, about 8 mph, with steering limited to 100 degrees of wheel angle per second. Each model observed the course through camera frames, then issued a motion command with a direction, steering percentage, speed, and duration, and had to send a new command before the old one expired or the car would start braking automatically.
That loop is where things fell apart. A model that takes ten seconds to think has let the car travel roughly ten meters with no input, at walking pace. Fable spent only 31 seconds in motion out of a roughly 190-second run, parked and deliberating the rest of the time. These are models built to generate text, not to run a continuous read-react loop against real-time momentum and geometry, and the gap showed up immediately at the first corner, where most agents couldn’t work out which side of a diagonal line of cones the lane was actually on.
Astra’s first run got it to 49% of the course. Its second attempt, five minutes and 22 seconds of crawling at under a walking pace, finished the job, hugging the long right-hander and nearly drifting off the course near the finish box before parking in the marked zone. It used full steering lock on 20 of its 24 commands, a habit it seemingly picked up after reflecting on its first failed run.
The win cost $7.74 in token usage, close to four times what its failed first attempt had burned through. Grok’s third, equally unsuccessful attempt reportedly used more than 516,000 tokens just to cover about 22 meters before giving up. None of this reads like a threat to Waymo or any purpose-built autonomous system. What it does show is that letting a general-purpose chatbot steer a real car is, at best, a slow and expensive way to find a cone to hit.
No Comments