Label Science for Autonomous Vehicles
COSIMO
An autonomous vehicle learns to see the world by analyzing and learning from labels affixed to the recorded driving footage. Labels are notes that identify tricky, safety-related scenes that need clarification.
However, it is incredibly expensive to label a fleet's recorded driving footage:
COSIMO's WITNESS platform reads a fleet's entire archive and ranks every scene by how much it can teach the model. The scenes that matter most, the ones with pedestrians, roadside workers, and anyone else on foot, move to the front of the queue. The empty scenes move to the back.
WITNESS does this with neither AI nor human review. It measures the geometry of each scene directly.
The result is a scored, ranked, and sorted queue instead of an unsorted pile of footage. We move every label-relevant scene to the front of the line, where the footage is densely packed with people on foot, 8.4 times as many as the unsorted footage.
Physical AI learns faster when it studies the scenes that matter.
The Problem
Human Hours
Human labeling is expensive in hours.
It can take up to 800 human hours to label a single hour of driving footage.
AI Compute Costs
AI-based auto-labeling is expensive in compute.
It can take up to $560 in compute to auto-label a single hour of driving footage.
Time and Money
Inference-based labeling and human labeling both cost a great deal.
When these expensive processes spend time labeling empty driving footage, the unhelpful footage that carries no label-relevant material, it is an unfortunate waste of time and money.
Model-Free | Human-Free
WITNESS does this with neither AI nor human review. It measures the geometry of each scene directly.
WITNESS solves this problem by relying on neither AI nor humans to analyze, score, and rank a fleet’s recorded driving footage. With a much more cost-effective solution, WITNESS hands your labeling pipeline a ranked and ordered queue of scenes, densely packed with label-relevant material.
Accelerated discovery of label-relevant scenes
The Top 5% of our ranked scenes holds 8.4 times as many people as the first 5% of the unranked, unsorted footage.
The result is that physical AI companies have decided to throw away their greatest asset. They are throwing away valuable footage generated by an expensive fleet of cars equipped with expensive sensors to capture every hour of driving footage. Having invested so much up-front capital to purchase an expensive fleet of autonomous vehicles, it would make sense to generate as much value as possible from it.
WITNESS scores every scene, then moves the label-relevant scenes to the front of the labeling queue. The empty scenes are shuffled to the end of the line. No need to waste your expensive human hours or AI compute cycles on them.
As recorded
WITNESS Ranking
We cut the ranked archive into ten equal bands. Each band is 10,000 scenes. The first band is the 10,000 scenes we scored highest. The last band is the 10,000 we scored lowest. Then we counted, band by band, how many of those scenes have a person in them.
The dark bars do the same thing to the same archive, left in the order it came off the vehicle.
Read it left to right. Our bars start high and fall away to almost nothing. The recorded order stays flat, between 708 and 1,205 a band, because nothing sorted it and the people are spread evenly through it.
| Ranking band, 10% each | PAI Ranking | Temporal Order | Multiple |
|---|---|---|---|
| Top band | 2,983 | 708 | 4.21x |
| 2nd band | 2,391 | 870 | 2.75x |
| 3rd band | 1,761 | 988 | 1.78x |
| 4th band | 1,323 | 1,104 | 1.20x |
| 5th band | 905 | 1,140 | 0.79x |
| 6th band | 557 | 1,205 | 0.46x |
| 7th band | 291 | 1,175 | 0.25x |
| 8th band | 122 | 1,027 | 0.12x |
| 9th band | 30 | 1,009 | 0.03x |
| Last band | 4 | 1,141 | 0.00x |
Inference-based labeling and human labeling both cost a great deal. WITNESS saves you real money and real time. What is it worth to get to market more quickly with a safer product?
The Archives
We did not test our solution on our own data or against our own performance standards. All three archives below are published by other organizations, carry labels made by those organizations, and can be downloaded by anyone who wants to run the same measurement we ran.
NVIDIA PhysicalAI
Published by NVIDIA
An autonomous vehicle archive of twenty-second driving clips carrying radar, lidar and camera. Its labels are placed by machine rather than by people, so it shows how WITNESS behaves when the answer sheet is itself automated.
nuScenes
Published by Motional
A city driving archive carrying radar, lidar and camera. Its labels are placed by human annotators, which makes it the strictest test we run: our order is graded against people rather than against another machine.
Zenseact ZOD
Published by Zenseact
The largest archive we have run, carrying camera and lidar. It has no radar, which is why our first pass on it is a lidar pass. It is released under a licence that permits commercial use, so a customer can download it and repeat our measurement without asking anyone's permission.
The result: reviewing one scene in twenty, a team working in PAI Ranking order reaches 4.9 times as many scenes with people in them as a team reviewing the same driving footage in the order it was recorded. The figure is measured on the Zenseact Open Dataset.
NVIDIA and nuScenes ask that numbers measured on their datasets not be published. So the figure we publish is measured on ZOD, which is released under a licence that permits it.
Zenseact Open Dataset (ZOD), © 2022 Zenseact AB, licensed under CC BY-SA 4.0.
Pedestrians
The PAI Ranking prioritizes scenes that are relevant for labeling. One category that deserves special attention is pedestrians.
Finding those scenes is vital to the safety of a driving model.
Accelerated discovery, measured
From the same amount of review, the PAI Ranking reached 4.9 times as many scenes with people in them as Temporal Order.
The result: WITNESS finds the scenes with people in them 4.9 times faster than reviewing the same driving footage in the order it was recorded.
82%
This chart keeps a running total. Walk down the list and at every point it asks the same question: of all 10,367 scenes in the archive that have a person in them, how many have you met so far?
Our line climbs steeply and then flattens out, because most of the people are near the front. By the time you are 40% of the way down the list, 81.6 of every 100 scenes with a person are already behind you. Go 40% of the way through the footage in the order it was recorded and 35.4 of every 100 are.
You do not need an expensive model to get this. The ranking reads the vehicle's own lidar and nothing else.
| Depth of the ranking | PAI Ranking | Temporal Order |
|---|---|---|
| Top 10% | 28.8% | 6.8% |
| Top 20% | 51.8% | 15.2% |
| Top 30% | 68.8% | 24.8% |
| Top 40% | 81.6% | 35.4% |
| Top 50% | 90.3% | 46.4% |
| Top 60% | 95.7% | 58.0% |
| Top 70% | 98.5% | 69.4% |
| Top 80% | 99.7% | 79.3% |
| Top 90% | 100.0% | 89.0% |
| Top 100% | 100.0% | 100.0% |
Four bands in, you have reached most of the people in the archive. The rest of your footage is still there, in order, whenever you want it.
The Process
Send us your archive
Send
Send us your entire multi-sensor driving footage archive, every sensor stream intact.
Rank
We process it and rank every scene by its label-relevance.
Return
You get the archive back, sequenced and organized in order of priority.
The result accelerates discovery by 2.5x to 2.7x on all label-relevant material, and by 4.2x to 4.9x on the scenes with people in them. Faster through the same archive, and no more setting aside scenes that were label-relevant all along.