In short
Three Axiom engineers used OpenAI’s GPT-6 Astra to steer a Toyota Corolla to an In-N-Out drive-thru, showing early signs that general-purpose AI can handle limited physical tasks. The experiment is promising but also highlights major safety and reliability gaps before AI driving is anywhere near real-world deployment.
- Axiom engineers linked GPT-6 Astra to a Toyota Corolla and used it to navigate an In-N-Out drive-thru.
- The test suggests multimodal AI models may be developing basic physical reasoning and spatial awareness.
- A new benchmark, DrivingBench, shows the models are still far from reliable driving performance.
- Researchers say the next AI frontier may be embodied systems such as robots and vehicles.
- Safety remains the biggest issue because a mistake in a car can have serious consequences.
Three AI engineers at startup Axiom recently used OpenAI’s GPT-6 Astra to steer a 2024 Toyota Corolla through an In-N-Out drive-thru in the Bay Area, an experiment that underscores how quickly general-purpose AI is beginning to handle physical-world tasks. The stunt was low-speed and closely supervised, but it also showed why the race to build models with real-world reasoning skills now matters beyond chatbots and coding assistants.
The trip to the burger chain’s take-out window was not a product launch or a polished demo. It was an informal test that nevertheless offered a glimpse of a new phase in AI development: models that were originally designed for text generation now being pushed, in real time, to interpret camera feeds, understand a vehicle’s environment and make driving decisions on the fly.
How did a language model end up driving to lunch?
It happened because the engineers connected a conversational AI model to the hardware and sensor stack of a car and asked it to control the vehicle directly. The system used a chat interface linked to a server, which in turn received input from windshield-mounted cameras and sent commands to the Corolla’s power steering. A human safety driver remained in the vehicle with a foot ready over the brake.
That setup matters because GPT-6 Astra was not built as a purpose-trained self-driving system. In principle, it is the kind of model most people associate with writing, coding, and image generation. Yet in this case, it was asked to navigate a narrow, messy, highly human environment: a fast-food drive-thru lane with painted lines, other vehicles, curbs, signage, and a transactional handoff at a window.
One of the engineers joked that the experience made it feel as if artificial general intelligence might already be here, reflecting the surprise of seeing a text-first model handle a physical task in the real world.
That reaction captures the broader significance of the experiment. Self-driving cars have existed in various forms for years, but most rely on specialized systems trained specifically for vehicle control. What made this test unusual was not merely that a car drove itself. It was that a general-purpose model, without prior task-specific coaching for this route, appeared capable of making enough sense of the environment to complete a mundane errand.
Why this experiment matters for AI development
The drive-thru test highlights a major frontier in AI: physical reasoning. Today’s models are strong at operating in the digital world, where text, code, images and structured data can be processed at scale. But they often struggle when the same intelligence has to be applied to real spaces that are incomplete, noisy and unpredictable.
That gap is exactly what researchers and startups are now trying to close. If AI systems can better understand depth, motion, object placement, spatial relationships and the consequences of actions in the physical world, they could become useful in robotics, manufacturing, retail, transportation and home automation.
In practical terms, that would mean more than novelty. A model that can reason about a sidewalk, a kitchen, a warehouse aisle or a delivery route could eventually power robots, inspection systems, logistics tools or assistive devices. The same advances could also raise safety concerns if models become capable of controlling machines faster than humans can react to mistakes.
From text generation to spatial judgment
The key shift is that the latest AI systems are increasingly multimodal, meaning they can process and combine images, video, audio and 3D information alongside language. Researchers say that kind of training may create unexpected abilities, including basic spatial judgment that was not explicitly taught as a driving skill.
That possibility is one reason the In-N-Out test got attention. It suggests that competence in one area may spill into another. A model trained broadly on multimodal data might not just describe a scene; it may begin to infer how to move through that scene, even if imperfectly.
What did the researchers build?
The Axiom engineers built a system that connected a chat-style interface to the car’s control environment. Their vehicle had cameras mounted on the windshield, and the AI output fed into steering-related functions. A human backup driver stayed ready throughout the process, making this a supervised experiment rather than autonomous public road driving.
The team said the goal was not to prove they had solved self-driving. Instead, they wanted to understand how well current models could operate in an unstructured physical setting. The drive-thru route was deliberately ordinary, which made it a useful stress test: everyday driving includes lane positioning, speed adjustments, tight turns, and continuous interpretation of the environment.
The engineers reportedly tried several models, including systems from OpenAI, Anthropic and Grok’s maker before settling on a result worth testing further. Early responses from the models were cautious. Some initially refused to issue commands to a physical car, saying they could interpret road images but not control motion. With more careful prompting, however, they were coaxed into participating.
Why the models’ resistance matters
That initial refusal is revealing. It shows that safety and capability are not the same thing. A model may be able to understand a scenario at a high level while still having guardrails that stop it from acting. Once those constraints are relaxed or circumvented, the underlying ability becomes visible — along with the risk.
The test also suggests that AI systems may be developing forms of in-context adaptation. The engineers said the models appeared to improve while interacting with the controls, learning from mistakes as the task unfolded. If accurate, that would mean the systems were not just executing prewritten behavior but adjusting in response to feedback.
What is Humanity’s Sixth Sense?
Humanity’s Sixth Sense is a new benchmark designed to measure how well AI models understand physical scenes, not just what is visible in them. It was developed by Elorian AI and Scale AI as part of a broader effort to test intuitive scene understanding, a skill that humans use constantly without conscious thought.
According to the researchers involved, much of visual AI has historically focused on perception: recognizing objects, identifying labels, and spotting patterns. The benchmark aims to push beyond that by asking a more difficult question — can a model understand what is happening in a scene in a way that supports intelligent action?
| Benchmark / Test | Purpose | Outcome | What it suggests |
|---|---|---|---|
| In-N-Out drive-thru experiment | Real-world vehicle control using GPT-6 Astra | Corolla successfully reached the window | General-purpose models may have emerging physical reasoning |
| Humanity’s Sixth Sense | Measure intuitive understanding of physical scenes | Used to evaluate scene comprehension | Researchers are moving beyond perception toward action-aware vision |
| DrivingBench | Test basic driving around a parking-lot course | GPT-6 Astra completed the course slowly; other models fell short | Current models are capable but still far from reliable driving |
How far can today’s models actually drive?
Not very far, at least not yet. The team’s DrivingBench benchmark, which uses a simple parking-lot course, showed that current frontier models remain limited. GPT-6 Astra was the only one to finish the loop, and it did so slowly. Anthropic’s Claude Fable 5.1 reached about 45 percent of the course, while Grok managed only 11 percent.
Those results are important because they separate spectacle from readiness. A car creeping through a parking-lot course or making a quick stop at a drive-thru is not the same as safe, consistent driving in traffic. The benchmark results suggest the models are beginning to grasp the task, but they are not close to mastering it.
That distinction is critical for anyone thinking about commercialization. A system that can appear to drive under controlled conditions may still fail unpredictably when exposed to pedestrians, weather, road closures, aggressive drivers or unexpected objects. In other words, the gulf between a memorable demo and a safe product remains wide.
What the results say about multimodal scaling
The researchers believe the vehicular capability may not have been explicitly trained as driving skill. Instead, it could be an emergent property of scaling up multimodal training — the growing practice of teaching models using images, video, 3D representations and language together.
If that theory holds, it would mean AI systems can acquire useful real-world abilities indirectly, by learning general spatial patterns rather than memorizing driving examples. That would be a powerful development. It would also make these systems harder to predict, because new behaviors might arise without being directly engineered.
Who is building physical-reasoning AI?
A growing number of startups and labs are now focusing on the same problem: how to make models understand the physical world as intuitively as humans do. One example is Elorian AI, founded by former Google DeepMind researcher Andrew Dai.
Dai argues that stronger visual and spatial reasoning could unlock practical uses ranging from restaurant-monitoring systems that can tell whether diners are enjoying a meal to domestic robots that can navigate homes and complete tasks. He sees robotics as a natural proving ground for this work, since robots must constantly interpret real spaces and respond safely.
Elorian’s founder has said that home robotics is difficult to imagine without better physical reasoning, because systems need to understand objects, distances, obstacles and changing environments in real time.
That perspective reflects a broader industry shift. For several years, the AI boom has centered on chatbots, code assistants and creative tools. The next wave may be less about conversation and more about embodiment — giving software the ability to perceive, plan and act in the material world.
Why are cars becoming AI test beds?
Cars are attractive test beds because they combine perception, decision-making and motion in a single system. A vehicle has to interpret lanes, signs, obstacles and human behavior while continuously adjusting its path. That makes driving a useful proxy for broader physical intelligence.
The Bay Area setting also matters. The region is saturated with autonomous-vehicle culture, from Tesla’s driver-assistance features to Waymo’s robotaxi presence. That environment has likely normalized the idea that software can increasingly be asked to handle transportation tasks, even when the underlying technology is still imperfect.
But the experiment also reveals a tension. Consumer exposure to driver-assist systems can make AI driving seem routine, when in reality the technical and safety thresholds for reliable autonomy remain very high. A slow, supervised drive to a burger window is a far cry from unsupervised navigation across a city.
What risks come with giving AI physical control?
The risks are obvious: a mistake in a digital interface is inconvenient, but a mistake in a moving car can be catastrophic. That is why the engineers maintained a human safety driver and why the experiment took place in a tightly controlled setting.
As models become more capable, the challenge will be deciding how much authority they should have over real-world machines. Physical systems do not offer the same forgiveness as text generation. A hallucinated sentence is embarrassing; a hallucinated maneuver could be dangerous.
That means researchers, regulators and companies will need to think carefully about testing, safeguards and deployment. The most advanced models may be able to infer a route or understand a scene, but proving reliability across millions of edge cases is a much harder problem.
Key safety questions to answer
- How should a model’s confidence be measured before it is allowed to move hardware?
- What failsafe should override the AI when it behaves unexpectedly?
- How can researchers validate performance across rare but dangerous edge cases?
- Where should supervised experiments stop and public-road deployment begin?
How should readers interpret the In-N-Out stunt?
It should be read as a signpost, not a finish line. The experiment does not prove that general-purpose AI can safely replace dedicated driving systems, nor does it mean the arrival of artificial general intelligence. But it does show that current models are beginning to cross an important boundary: from describing the world to interacting with it.
That transition is likely to reshape the next stage of AI competition. Companies that once focused mainly on benchmark scores, chat quality and coding ability now have an added challenge: proving that their models can reason about space, motion and consequences outside the screen.
For now, the most striking part of the story is its ordinariness. A lunch run to In-N-Out would normally be forgettable. In this case, it became a demonstration of how quickly AI is moving from a virtual assistant to something closer to a physical actor — imperfect, tentative and supervised, but undeniably more embodied than before.
Timeline of the experiment and the surrounding research
The progression from lunch idea to benchmark development shows how quickly interest in physical reasoning is accelerating.
| Stage | What happened | Why it matters |
|---|---|---|
| Weekend discussion | Three Axiom engineers wondered whether multimodal models could transfer to real-world driving | Shows curiosity around emergent spatial reasoning |
| System setup | They connected a chat interface, cameras and steering control in a Toyota Corolla | Turned a language model into a real-time vehicle controller |
| Drive-thru test | The car navigated to an In-N-Out window with a safety driver onboard | Demonstrated an unexpected ability to handle a physical task |
| Benchmarking | The team created DrivingBench and connected it to broader scene-understanding research | Provides a more systematic way to evaluate capability |
| Results | Astra outperformed Claude Fable 5.1 and Grok on the benchmark, but all models remained limited | Indicates progress, but not readiness for general driving |
What comes next for AI and the physical world?
The next phase of AI development is likely to be defined by embodiment: robots, vehicles, cameras, industrial systems and assistive devices that require models to make sense of space, motion and timing. The core challenge will be turning promising demonstrations into reliable systems.
That will require better training data, stronger benchmarks, more robust fail-safes and probably new ways of measuring generalization. It will also require a more sober public conversation. The idea that a chatbot can drive a car sounds like sci-fi, but the real story is more nuanced: AI is learning enough about the physical world to be useful in some settings, while still being far from safe in many others.
For now, the image of GPT-6 Astra inching a Corolla up to a drive-thru window serves as a vivid snapshot of where the field stands. AI is no longer confined to the screen. It is reaching, awkwardly but undeniably, into the physical world.
Frequently asked questions
Did GPT-6 Astra really drive a car to In-N-Out?
Yes, GPT-6 Astra was used in a supervised experiment to help steer a 2024 Toyota Corolla through an In-N-Out drive-thru. A human safety driver stayed in the car and was ready to intervene, so this was not unsupervised autonomous driving.
Why is this AI driving test important?
It is important because it suggests that a general-purpose AI model may be developing enough physical reasoning to handle limited real-world tasks. That could matter for robotics, vehicles and other embodied systems, even though the technology is still far from safe public deployment.
How well did the models perform on the DrivingBench benchmark?
GPT-6 Astra completed the simple parking-lot course, but only slowly. Anthropic’s Claude Fable 5.1 finished about 45 percent of the route, while Grok managed just 11 percent, showing that current models still struggle with reliable vehicle control.
What is Humanity’s Sixth Sense?
Humanity’s Sixth Sense is a benchmark created by Elorian AI and Scale AI to measure how well models understand physical scenes. It is meant to test intuitive spatial understanding rather than simple object recognition, which is a harder step toward useful real-world action.
Are AI models close to replacing self-driving systems?
No, current AI models are not close to replacing purpose-built self-driving systems. The experiment shows progress in physical reasoning, but the results are still too limited and inconsistent for safe, general driving in real traffic.









