All posts

Waabi: a simulation-first bet on autonomous trucking

What Waabi has actually demonstrated, what remains a published claim, and how its neural simulator argues for replacing road miles as validation.

By Robo2u Editorial · 14 min read

Waabi is a Canadian company founded in 2021 by Raquel Urtasun, building an autonomous driving system for trucks and, since January 2026, robotaxis. Its company page lists teams in Toronto, San Francisco, Dallas, Phoenix and Pittsburgh, and designates no headquarters. Its public argument is that the industry validates self-driving the wrong way. Real-world mileage accumulates too slowly to encounter rare events, and closed-course proving grounds cannot stage them on demand. Waabi's answer is a neural simulator it says is faithful enough that results inside it count as safety evidence.

This profile reads only what Waabi publishes on waabi.ai: the technology, research, safety and company pages, and the dated insights posts. Everything below is vendor-stated unless flagged otherwise. No page fetched carries a third-party measurement, an arXiv link, a code repository, or a leaderboard result. Some individual research pages do host a paper PDF, and those PDFs are the only primary documents on offer. That framing matters for the reader question this series asks, which is what a lab has demonstrated versus what it has asserted.

Companion reading: Sim-to-real transfer, robot simulation and digital twins, self-driving cars and autonomous vehicles, edge AI and robot compute.

Table of contents

The thesis and the stack

Waabi calls its approach "AV2.0": one end-to-end model that interprets sensing, reasons over possible actions, and selects a maneuver in milliseconds, in place of a pipeline of hand-assembled modules. The company states three requirements for what it calls Physical AI: generalization to unseen scenarios, efficient low-power edge compute alongside training that costs "a fraction of today's capital-intensive brute force approaches," and safety validated "through robust scientific evidence to build comprehensive safety cases."

Three named components carry that thesis.

Waabi Driver is the vehicle-side product: the autonomy software plus sensors and onboard compute, described as vertically integrated into OEM platforms with factory-built redundancy. Listed subsystems include embedded compute and real-time decisioning, sensor fusion and calibration, emergency vehicle audio detection, sensor cleaning, redundancy systems, and external safety status lighting.

Waabi World is the simulator, described as high-fidelity and closed-loop. Its published pipeline builds data-driven digital twins of real scenarios, capturing actor behaviors, weather and infrastructure, then runs a loop that simulates sensor data, controls surrounding agents, models vehicle dynamics, and updates world state each fraction of a second. The autonomy stack consumes simulated sensor data and emits steering and acceleration. One detail is worth calling out because it is easy to skip and hard to fake. The technology page states that "Latency injection simulates computation time on the truck." Simulators that ignore compute delay flatter the policy under test. That the feature exists is Waabi's own description of its pipeline, with no external check available.

Mixed Reality Testing puts a real truck on a real track and injects synthetic actors into the live sensor stream, with virtual elements living inside what Waabi describes as a 4D neural digital replica. This depends on Onboard Waabi World, a version of the simulator that runs in a few milliseconds on vehicle compute and modifies live LiDAR, camera and RADAR readings in real time. Waabi disclosed MRT on 2025-07-14 and said it had been a central testing approach for more than two years by then.

No version numbers are published for any of these components.

How Waabi defines simulator realism

This is the strongest technical writing on the site, and it deserves credit before the gaps get discussed. It appears in a post dated 2025-03-11.

Waabi defines realism by behavioral outcome. The test is whether the autonomy system drives in the virtual environment the way it drove in reality when presented with the same situation. Visual appearance is explicitly set aside. The measurement method is pair-setting. Take a real log, recreate the scenario in simulation by matching actors, their appearance, their precise behavior, the weather, illumination and road conditions, seed the simulator with the real log as preroll, then unroll closed-loop for a fixed duration, typically tens of seconds, with the same autonomy software release running in both. Then measure how far the trajectories diverge.

The published formula is Realism Score = (1 - average relative distance) x100.

Waabi explicitly rejects the common alternative of distribution-matching, comparing speed, acceleration and hard-brake distributions between sim and real. The stated reason is sharp: two systems can share aggregate distributions while reacting differently to every individual situation. Their illustration is a hard brake triggered in simulation by a false-positive pedestrian and in reality by a tire shred. Same histogram bucket, different causes, different system.

They also publish a three-tier ladder they position themselves at the top of. Standard simulators catch bugs and regressions. Advanced simulators reveal where the system fails. The stated holy grail is a simulator admissible as evidence in a safety case. And they name their own conditional: a simulator that misses the realism bar "would be rendered useless as a solution for assessing a system's performance."

Waabi reports an average realism score of 99.7 percent for Waabi World. That figure is self-reported by the lab, measured by the lab, on its own simulator, using its own definition. No independent party has measured it and there is no benchmark it can be compared against.

What is demonstrated and what is claimed

Take the realism figure first. The method is well specified and the result is unverifiable. The site gives no sample count, describing "extensive paired tests" and "statistically significant evidence" without a number. It reports an average with no variance, no percentiles, and no worst case, which is the awkward part, because a safety case turns on the tail and the tail is exactly what is withheld. The term "relative distance" is not defined: relative to what baseline, over what horizon, in what frame. The scenario mix is unstated, so a reader cannot tell whether the pairs were routine highway cruising or the safety-critical long tail the entire argument rests on. Realism on nominal driving is cheap. There is no software release, measurement date, ODD, or highway versus surface-street split. There is no released data, protocol document, or third-party review.

The same article calls on the industry to publicly demonstrate the quantified realism of their simulators and to adopt this as a standard. Waabi has published the ruler and withheld the measurements.

Second, the transfer claim. On 2026-06-26 Waabi stated that a Waabi Driver trained on a Peterbilt 579 generalized zero-shot to the Volvo VNL Autonomous, "without requiring new real-world data, simulation data, or fine-tuning," and was "fully performant from the very first mile" across highways and surface streets. If true, that is the most consequential result on the site, because cross-embodiment transfer is the thing that makes a fleet business scale. It is supported by prose, self-reported by Waabi. No intervention rate, no mileage, no comparison metric, no evaluation protocol.

Third, driverless status. In a post dated 2025-12-19 Waabi reported completing autonomy missions with no human on board in October, at its Phoenix test track, on a closed course. That is a real, dated milestone, and it is Waabi's own account of it. It is also closed-course, and Waabi says so plainly, adding that OEM redundant platforms "have not received all necessary approvals and validation for driverless deployment yet," so the runs used a small development fleet with redundancies built in-house and from suppliers. No driverless operation on public roads appears anywhere on the site.

Fourth, operational metrics. Autonomous miles driven, disengagement or intervention rates, and safety incidents are absent entirely. So is an ODD definition: no geographic scope, weather envelope, or speed limits. The only safety PDF linked from the safety page has a filename dating it to 2023-10-30, predating surface streets, driver-out testing, the Volvo integration and the robotaxi pivot, and the page presents it as "Our publicly shared VSSA" without flagging its age.

One more claim deserves flagging. The safety page states the end-to-end trained system "is fully interpretable, and can see exactly why a decision was made." That is a strong claim about a class of model where interpretability is contested, and nothing on the site defines or evidences it.

Credit where it is earned: Waabi does name real limitations. Remote assistance is deliberately bounded to high-level instructions such as new routes, because "remote systems cannot be relied upon for real-time control of the vehicle or any safety critical operations due to potential network latency or outages." And the realism post acknowledges that simulation error compounds, citing actor behavior, sensor simulation, latency modeling and vehicle dynamics, including a specific admission about "not accurately modeling gear shifts."

Research output and openness

The research page lists venue-accepted work with Urtasun on every paper: SaLF, a language-to-scenario orchestration paper, GenRe and a Conditional Flow-VAE for safety-critical scenario generation at ICRA 2026; DriveGATr at CVPR 2026; Flux4D at NeurIPS 2025; FOMO-3D at CoRL 2025; DIO, MAD and GenAssets at CVPR 2025; UniSim at CVPR 2023; Copilot4D at ICLR 2024. The list is paginated behind a "Show More" control, so it may not be exhaustive.

Venue acceptance is the only external validation signal present anywhere on the site, and it is genuine. The pattern is heavy on simulation, reconstruction and scenario generation, which is consistent with the stated thesis.

Openness is partial. Each paper gets its own page with an abstract, and some of those pages host the paper PDF. Of the pages checked on 2026-08-24, UniSim, SaLF, GenRe, DriveGATr and MAD carry a PDF link; Copilot4D, Flux4D, DIO, FOMO-3D and GenAssets do not. The UniSim page carries the CVPR 2023 abstract, a PDF hosted on Waabi's asset domain, and qualitative video and GIF results covering actor removal, actor insertion, trajectory manipulation, viewpoint shifts and closed-loop sensor simulation. What is absent across every page fetched is an arXiv link, a GitHub repository, model weights, datasets, and any licence terms. Waabi is a closed commercial lab that publishes at conferences, hosts some of its own papers, and ships no artifacts.

Partners, capital, and commitments

Volvo Autonomous Solutions is the truck OEM partner, with the Waabi Driver integrated into the Volvo VNL Autonomous, which Waabi describes as purpose built for autonomy with redundancies or back-up systems for safety critical functions such as braking, steering and communication (2025-10-28). NVIDIA supplies compute, with the joint solution integrating NVIDIA DRIVE AGX Thor and the NVIDIA DRIVE AGX Hyperion 10 architecture, and NVentures among the investors. Uber Freight is a freight channel partner, and Waabi names Samsung as a shipper now moving loads with Waabi on that network (2025-11-25). Uber holds an exclusive robotaxi partnership including a stated commitment to deploy "25,000 or more Waabi Driver-powered robotaxis over time."

The Series C closed at 750 million USD, announced 2026-01-28, co-led by Khosla Ventures and G2 Venture Partners, described by Waabi as the largest fundraise in Canadian history, with the 1 billion USD headline including Uber's milestone-based commitment. Investors named on that announcement are Uber, NVentures, Volvo Group Venture Capital, Porsche Automobil Holding SE, funds and accounts managed by BlackRock, Radical Ventures, HarbourVest Partners, a wholly owned subsidiary of the Abu Dhabi Investment Authority, Linse Capital, Incharge Capital, BDC Capital's Thrive Venture Fund, Export Development Canada, TELUS Global Ventures and BMO Global Asset Management. Both halves of that list close with "and others," so it is not complete. Waabi's company page carries a separate investor logo wall that also includes Scania and Ingka, named there with no corporate suffix and with no funding round attached. Lior Ron joined as COO on 2025-08-12.

The robotaxi commitment is the loudest gap in this section. Tens of thousands of units are committed with no OEM, no vehicle platform, no launch city, and no date published.

What to watch

  • A dated public-road driverless run. The December 2025 post ends with "Stay tuned." That milestone converts the closed-course result into an operational one.
  • The distribution behind the realism average. N, variance, worst case, and scenario mix. Publishing the tail would turn a self-reported figure into a measurement.
  • A protocol document defining relative distance. Without it, no one outside Waabi can reproduce or contest the score.
  • An updated safety self-assessment. The current one predates most of what the company now does.
  • Any transfer evidence with numbers. An intervention rate on the Volvo platform, compared against the Peterbilt baseline, would settle the zero-shot claim.
  • A named robotaxi vehicle and city. The Uber commitment stays abstract until then.
  • Whether anyone adopts the realism standard. Waabi asked the industry to quantify simulator fidelity. Uptake by a competitor or a regulator would be the strongest possible validation of the method.

Frequently asked questions

What does Waabi actually build? The Waabi Driver, an autonomy system combining software with sensors and onboard compute, integrated into OEM truck platforms, plus Waabi World, the neural simulator used to train and test it.

Is Waabi operating driverless trucks on public roads? Not according to anything published on its site. The dated driver-out milestone was on a closed course at Waabi's Phoenix test track in October, reported on 2025-12-19.

What is the realism score and can it be checked? It is defined as (1 - average relative distance) x100, measured by recreating real scenarios in simulation and comparing trajectories with the same software release in both. The reported average is self-reported by Waabi and cannot be checked externally. No sample count, variance, scenario mix, data, or protocol document is published.

Does Waabi release code or model weights? No. Some research pages host the paper PDF, and that is the extent of it. There are no model weights, no code, no datasets, no licence terms, no arXiv links and no repositories on any page fetched. Papers appear at ICRA, CVPR, NeurIPS, CoRL and ICLR.

What is the zero-shot cross-embodiment claim? That a driver trained on a Peterbilt 579 ran on a Volvo VNL Autonomous with no new real or simulated data and no fine-tuning, performant from the first mile. Waabi published it as narrative with no supporting metrics.

Who are the confirmed partners? Volvo Autonomous Solutions for the truck platform, NVIDIA for compute, Uber for robotaxis, Uber Freight for freight, with Samsung named by Waabi as a shipper.

How much has Waabi raised? A 750 million USD Series C announced 2026-01-28, co-led by Khosla Ventures and G2 Venture Partners. The 1 billion USD headline figure includes Uber's milestone-based commitment.

Related guides