Levels of trust in AI chatbots differs pretty significantly in the population. While many are skeptical, calling large language models (LLMs) slop or plagiarism machines and worrying about problems such as "AI psychosis," others are more than happy to ask a bot for mental health advice (spoiler alert; not a great idea).
Then, there are those who have no qualms putting it in control of a vehicle, to see how it handles things (spoiler alert; also not a great idea, at least if you want to arrive at your destination).
Autonomous vehicles have been touted for quite some time now. There's even a handy Wikipedia list of Elon Musk's predictions for autonomous Tesla vehicles, documenting the numerous times he erroneously suggested truly autonomous vehicles were just around the corner.
"I think we will be feature complete — full self-driving — this year, meaning the car will be able to find you in a parking lot, pick you up and take you all the way to your destination without an intervention, this year," he said in 2019, a year of particularly optimistic statements about autonomous driving by Musk.
"I would say I am certain of that. That is not a question mark," he added for good measure.
Of course, that did not happen. But progress in the area, while slower than Musk may like, is nonetheless impressive. There are cars on the roads that can drive themselves under supervision, and even taxi services like Waymo, which can ferry people around without "autonomous specialists" to take over if things go wrong.
So why let chatbots get involved? Well, though specialized, highly engineered self-driving cars are awesome, it would be far cooler if a general-purpose algorithm could figure out driving all on its lonesome.
Generative AI, though definitely not there yet, has shown promise in picking up tasks it wasn't necessarily trained to do (for instance, representing board game positions in order to play them). So why not let it have a go in a car?
In new tests by the group DrivingBench, AI chatbots GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol were asked to do just this, operating a Toyota Corolla using cameras to feed the chatbots with information about the course they were attempting to navigate.
The first problem, it turns out, was convincing them to drive at all.
"Some models (especially GPT-6 Astra) would refuse to drive the physical car sometimes, citing safety reasons (even in a completely empty lot, after prompting it with all the safety measures we had including the very low speed limit caps and human ready the [Sic] brake)," the team said in a statement.
"We tried many prompt changes to get them to consistently drive, for example attempting to call it a 'simulation' (but then in some trials they would see the real images and realize it's real, and start freaking out.)"
The team eventually found that appending the name of one of the components with "sandbox" was enough to convince them to start driving consistently. But when they did, they weren't particularly good at it.
GPT-5.6 Sol showed the most consistency, in that it was consistently the worst, only making its way around 6 percent of the 130-meter course in its three attempts to navigate it. Total distance traveled was a touch higher, though, because the LLM backtracked and wiggled around a bit relative to the course line.
The chatbots, following each attempt, were asked to "reflect" upon the previous attempt to see if they could "learn" from their mistakes, but given that GPT-5.6 Sol only traveled 17.1 meters (56 feet) by its third attempt (only 2 meters longer than its first attempt) this didn't have much of an impact.
Grok didn't fare much better, with its best attempt being a fairly unimpressive 22.6 meters (74 feet), with which it managed to complete around 11 percent of the course. That leaves Fable and Astra as the clear winners.
"GPT-6 Astra was the only model to fully complete the course," the team said.
"Claude Fable 5.1's third attempt got around halfway through the course, as did Astra's first attempt. All other attempts didn't make it past the first corner."
While at low speeds and with a low success rate, the team was impressed that the chatbots were able to drive at all, and "learn" to use the steering commands, with Astra and Fable showing improvements between runs.
That said, they acknowledge they were only able evaluate each model's set of attempts once, whereas a more thorough experiment would include more repeats and potentially try the models out on different courses.
"We expect progress to continue to improve in this area as further related data is added to the training mix, model inference speeds / latency improve, and model planning and in-context learning abilities sharpen over time," they conclude.
"While the success of the models on this benchmark was shocking and exciting to see, this also calls for additional pressing work on safety/alignment/evaluation."





