Drag Play
Video

Planetary Health

The Holy Grail of Autonomy is Coming into View

Story by Mo Islam
09/04/2026

In December of 2014, when I worked at the venture firm DFJ, we made a seed investment in a company called Zoox that wanted to build a car with no steering wheel, brake pedals or driver’s seat. At the time, it was a radical idea rooted in technologies that did not fully exist yet. But the bet was that software would learn to drive, and that autonomy would eventually change the vehicle itself.

Twelve years later, Zoox began commercial service this month. What’s more, the broader self-driving industry has largely converged on Zoox’s approach.

Dedicated hardware is pivotal to commercial Level 4 operations today, and coupled with frontier AI models in autonomy, I believe it will help us finally achieve Level 5, the north star in autonomous driving.

Zoox’s purpose-built prescience

When Zoox raised its seed round, robotaxis did not exist, regulators had no framework for a vehicle without human controls, and most companies were betting that the way to autonomous vehicles was to mount sensors and cameras on existing cars. Zoox made a broader and more ambitious bet. If a driver was no longer needed, then the car didn’t need to be designed around a steering wheel, pedals, and a panel of signals and buttons.

Source: Wikipedia

That idea led to a new design with a bidirectional cabin and four inward-facing seats. More importantly, it forced Zoox to dramatically transform how a car was built. It would need to vertically integrate the car, its autonomy stack, its manufacturing factory, and its regulatory compliance all in one stack. Working on existing cars would have been faster. But Zoox’s new design ultimately led to a better, more functional, and more efficient autonomous vehicle.

The robotaxi horse race

Following Zoox’s example, Waymo and Tesla have also moved toward dedicated autonomous vehicles with their latest hardware architectures. As the technology has developed, the industry has concluded that driverless mobility is a new product category, not a feature added to a conventional vehicle.

The robotaxi race now depends on a new series of decisions. Should companies deploying  autonomous vehicles own the chassis or buy it? Should they use LIDAR or rely on cameras to navigate? Should they seek a regulatory exemption or self-certify? These design and technology bets are the next frontier. Each company is trying something different.

Fig. 1 — Three approaches to Level 4 autonomous vehicles.

The three vehicles embody distinct strategic bets. The differences reveal how each company plans to compete on manufacturing, regulatory strategy, and scaling.

Zoox’s bidirectional symmetry is removing the distinction between front and back. Sensor pods at all four upper corners maximize coverage. The design produces the cleanest driverless form factor, but it also requires the most vertically integrated manufacturing and the longest regulatory path.

The Waymo Ojai takes a hybrid approach, integrating its sixth-generation Driver—including a roof-mounted Lidar dome and perimeter sensor pods—into a Zeekr-built chassis. By retaining manual controls, Waymo can follow the standard FMVSS compliance path and avoid exemption-based volume limits, trading a fully driverless interior for a faster route to fleet scale.

The Tesla Cybercab pairs a low, two-seat body, butterfly doors, and no manual controls with a camera-only system, meaning no Lidar and no roof dome. This approach lowers hardware costs while placing a greater burden on safety validation.

The bridge from Level 4 to Level 5

Autonomous driving is measured across five levels, defined by how much responsibility the system assumes versus how much remains with a human driver. In Levels 1 and 2, the system controls elements of steering and breaking, but a human still remains responsible and available to intervene. In Level 3, the system takes over all driving, but may request a human to take over when requested. 

Level 4 expands to full autonomous driving without human intervention, but only in a defined operating domain, such as within city limits, on surface streets, or in good weather conditions. Level 5 removes those constraints. In Level 5 automation, the system can drive anywhere and under all conditions a human driver could. It does not require a driver, steering wheel, or pedals.

Commercial Level 4 autonomy is already operating at meaningful, but varying, scales. Waymo currently leads, operating roughly 3,700 vehicles across seven public cities and completing about 500,000 rides per week. Zoox has launched paid service with the first purpose-built vehicle without manual controls to clear NHTSA review. Tesla, meanwhile, has begun unsupervised operations in the Austin metro area with a fleet of roughly 300 vehicles.

All of them will continue to expand their operational domains to achieve Level 5 autonomy, and all are boosted in their progress by frontier AI. The next round of innovations that will narrow the gap are integration of dedicated hardware, innovation in pretrained models, and simulation inside world models.

Pretrained model proficiency

Over the past decade while I’ve surveyed this market, the state of the art in autonomy has moved through many generations, each collapsing the amount of proprietary infrastructure a new startup must build before it can deploy in the physical world.

A vision-language-action (VLA) model starts with a multimodal foundation model that already links images and language, then adds an output head for trajectories or control commands. That architecture changes the task. Instead of teaching a model what the world is and how to act in it at the same time, a team can focus its effort on the final mapping from general understanding to domain-specific action.

Pretraining that once required tens of millions of dollars is available as a public checkpoint, and open-weight foundation models such as Alibaba’s Qwen-VL already contain useful spatial reasoning. A team no longer needs nine figures and four years to reach a working baseline. As a result, capital moves  away from recreating perception and toward data quality, validation, and deployment.

Under prior architectures, a new city or environment meant months of data collection, retraining, and validation. With a post-trained VLA, a human can drive the route for hours, a lightweight adapter can train overnight on a single GPU, and the system can begin testing within its operational domain the next day. The geofence doesn’t disappear, but its marginal cost falls dramatically. Geographic expansion is finally a repeatable operating process.

Wonder of world models

While the autonomy stack is now easier to build than in prior architectures, challenging edge cases still exist. Even the best innovation must confront the reality of hundreds of thousands of miles of different road conditions, high-stakes safety demands, and proofs against events that rarely occur. Tesla’s small unsupervised fleet and Waymo’s temporary pause on freeway driving both point to the same conclusion: without sufficient information, scaling stops.

Reality is a slow generator of rare events. A fleet cannot schedule tire debris in the fast lane at dusk or a forklift crossing behind a reversing trailer. It can only accumulate miles and wait. That makes validation expensive, time-consuming, and biased toward whatever the fleet has already encountered.

Generative world models change that equation. In 2026, they moved from research artifacts to production tools. Earlier simulators were trained on a single fleet’s data and could therefore vary only what that fleet had already seen. A world model that has broader physical world knowledge (i.e. post-trained on multimodal models such as video models) can generate arbitrary scenarios far beyond what any AV companies have ever collected.

Waymo is already demonstrating this shift. Its world model, post-trained from Google DeepMind’s Genie 3, generates both camera and lidar outputs and allows engineers to request scenarios in natural language, from snow on the Golden Gate Bridge to an elephant in the road, rather than wait for those events to occur. NVIDIA’s open-weight Cosmos models make similar capabilities available beyond the largest fleets.

That converts the scarcest input in autonomy—data from low probability events—from something a fleet passively and slowly accumulates into something an engineering team can actively generate. It does not eliminate real-world validation, but it makes the long tail testable before reality happens to deliver it. Companies like Moonlake AI and AMI Labs—an Obvious portfolio company—are working to make this simulation infrastructure available for autonomous driving developers.

What the future holds

Level 5 asks a system to perform in places and conditions it has not encountered, with no human fallback. Under prior paradigms, that was nearly intractable because edge cases were addressed one at a time or required massive amounts of fleet data. 

Pretrained backbones help invert that relationship. General knowledge is the starting point: the model can recognize a flooded underpass, infer what a person carrying a mattress may do, or interpret a yard worker’s hand signal because it has learned from more than fleet miles. 

World models then provide a way to rehearse the cases that public data and real-world driving underrepresent. Tightly coupled with dedicated hardware, these new autonomous systems are finally equipped to drive in almost any condition.

For investors, the opportunity is not limited to identifying which robotaxi platform wins. As pretrained models make baseline autonomy cheaper and faster to build, value will shift toward the remaining bottlenecks,including proprietary data and post-training loops, world-model simulation, safety validation and compliance, and the vertical integration of autonomy with purpose-built hardware. 

We also expect the new wave of autonomy companies to commercially operate initially in valuable, constrained environments, such as ports, yards, warehouses, construction sites and agricultural farms. These are the sites where labor is scarce or precious, the economics are immediate, and Level 4 can deliver value today, as we move toward Level 5. 

Twelve years after investing in Zoox, the lesson is not simply that autonomy takes time. Zoox had to build every layer itself, operate in a state that had banned the category, fit a driverless vehicle into safety rules written for drivers, and validate rare events without pretrained models or generative simulation. Its success reveals three triumphs at once: a transformation of hardware infrastructure, regulatory infrastructure, and model infrastructure. All three now exist and I believe will soon be tightly integrated across the major players. Level 4 is already commercially viable today. I believe Level 5 is finally feasible in a venture timeframe, and that the companies building vertically-integrated autonomy with purpose-built hardware will define the next era of autonomy.

Author
Mo Islam

Mo is a Partner at Obvious, where he leads new investments in deep tech and frontier AI companies.

Obvious Ideas