WITNESS

Label Science for Autonomous Vehicles

Scroll

COSIMO

Teaching Machines to See the Physical World

An autonomous vehicle learns to see the world by analyzing and learning from labels affixed to the recorded driving footage. Labels are notes that identify tricky, safety-related scenes that need clarification.

However, it is incredibly expensive to label a fleet's recorded driving footage:

COSIMO's WITNESS platform reads a fleet's entire archive and ranks every scene by how much it can teach the model. The scenes that matter most, the ones with pedestrians, roadside workers, and anyone else on foot, move to the front of the queue. The empty scenes move to the back.

WITNESS does this with neither AI nor human review. It measures the geometry of each scene directly.

The result is a scored, ranked, and sorted queue instead of an unsorted pile of footage. We move every label-relevant scene to the front of the line, where the footage is densely packed with people on foot, 8.4 times as many as the unsorted footage.

Physical AI learns faster when it studies the scenes that matter.

The Problem

Label the meaningful scenes first.

8x

Accelerated discovery of label-relevant scenes

The Top 5% of our ranked scenes holds 8.4 times as many people as the first 5% of the unranked, unsorted footage.

The result is that physical AI companies have decided to throw away their greatest asset. They are throwing away valuable footage generated by an expensive fleet of cars equipped with expensive sensors to capture every hour of driving footage. Having invested so much up-front capital to purchase an expensive fleet of autonomous vehicles, it would make sense to generate as much value as possible from it.

Ranked and Sorted

An unsorted archive becomes a ranked queue.

WITNESS scores every scene, then moves the label-relevant scenes to the front of the labeling queue. The empty scenes are shuffled to the end of the line. No need to waste your expensive human hours or AI compute cycles on them.

Label-relevant scenes Empty scenes

As recorded

Label the meaningful scenes first.
The empty scenes won’t waste time clogging your expensive labeling pipeline.

Schematic. The bar counts illustrate our methodology and are not a measure of exact distribution. Measured results vary from fleet to fleet and from one condition to the next.

WITNESS Ranking

Bringing Order to Chaos

We cut the ranked archive into ten equal bands. Each band is 10,000 scenes. The first band is the 10,000 scenes we scored highest. The last band is the 10,000 we scored lowest. Then we counted, band by band, how many of those scenes have a person in them.

The dark bars do the same thing to the same archive, left in the order it came off the vehicle.

Read it left to right. Our bars start high and fall away to almost nothing. The recorded order stays flat, between 708 and 1,205 a band, because nothing sorted it and the people are spread evenly through it.

PAI RankingTemporal Order, the order it came off the vehicle
08001,6002,4003,200PAI Ranking, Top tenth: 2,983 scenes with peopleTemporal Order, Top tenth: 708 scenes with peoplePAI Ranking, 2nd tenth: 2,391 scenes with peopleTemporal Order, 2nd tenth: 870 scenes with peoplePAI Ranking, 3rd tenth: 1,761 scenes with peopleTemporal Order, 3rd tenth: 988 scenes with peoplePAI Ranking, 4th tenth: 1,323 scenes with peopleTemporal Order, 4th tenth: 1,104 scenes with peoplePAI Ranking, 5th tenth: 905 scenes with peopleTemporal Order, 5th tenth: 1,140 scenes with peoplePAI Ranking, 6th tenth: 557 scenes with peopleTemporal Order, 6th tenth: 1,205 scenes with peoplePAI Ranking, 7th tenth: 291 scenes with peopleTemporal Order, 7th tenth: 1,175 scenes with peoplePAI Ranking, 8th tenth: 122 scenes with peopleTemporal Order, 8th tenth: 1,027 scenes with peoplePAI Ranking, 9th tenth: 30 scenes with peopleTemporal Order, 9th tenth: 1,009 scenes with peoplePAI Ranking, Last tenth: 4 scenes with peopleTemporal Order, Last tenth: 1,141 scenes with peopleTop2nd3rd4th5th6th7th8th9thLastTenth of the ranking, best first
Each band is 10,000 scenes, one tenth of the archive. The counts are scenes with a person in them.
Table view
Ranking band, 10% eachPAI RankingTemporal OrderMultiple
Top band2,9837084.21x
2nd band2,3918702.75x
3rd band1,7619881.78x
4th band1,3231,1041.20x
5th band9051,1400.79x
6th band5571,2050.46x
7th band2911,1750.25x
8th band1221,0270.12x
9th band301,0090.03x
Last band41,1410.00x
The takeaway

Inference-based labeling and human labeling both cost a great deal. WITNESS saves you real money and real time. What is it worth to get to market more quickly with a safer product?

The Archives

We tested WITNESS on three public driving archives.

We did not test our solution on our own data or against our own performance standards. All three archives below are published by other organizations, carry labels made by those organizations, and can be downloaded by anyone who wants to run the same measurement we ran.

The result: reviewing one scene in twenty, a team working in PAI Ranking order reaches 4.9 times as many scenes with people in them as a team reviewing the same driving footage in the order it was recorded. The figure is measured on the Zenseact Open Dataset.

NVIDIA and nuScenes ask that numbers measured on their datasets not be published. So the figure we publish is measured on ZOD, which is released under a licence that permits it.

Zenseact Open Dataset (ZOD), © 2022 Zenseact AB, licensed under CC BY-SA 4.0.

Pedestrians

A Physical AI model must “get it right” when it comes to pedestrians, roadside workers, and anyone else on foot in the geometry of a scene.

The PAI Ranking prioritizes scenes that are relevant for labeling. One category that deserves special attention is pedestrians.

Finding those scenes is vital to the safety of a driving model.

5x

Accelerated discovery, measured

From the same amount of review, the PAI Ranking reached 4.9 times as many scenes with people in them as Temporal Order.

The result: WITNESS finds the scenes with people in them 4.9 times faster than reviewing the same driving footage in the order it was recorded.

82%

Where the people sit in the ranking

This chart keeps a running total. Walk down the list and at every point it asks the same question: of all 10,367 scenes in the archive that have a person in them, how many have you met so far?

Our line climbs steeply and then flattens out, because most of the people are near the front. By the time you are 40% of the way down the list, 81.6 of every 100 scenes with a person are already behind you. Go 40% of the way through the footage in the order it was recorded and 35.4 of every 100 are.

You do not need an expensive model to get this. The ranking reads the vehicle's own lidar and nothing else.

PAI RankingTemporal Order, the order it came off the vehicle
0%25%50%75%100%Temporal Order: the top 10% holds 6.8% of every scene with a personTemporal Order: the top 20% holds 15.2% of every scene with a personTemporal Order: the top 30% holds 24.8% of every scene with a personTemporal Order: the top 40% holds 35.4% of every scene with a personTemporal Order: the top 50% holds 46.4% of every scene with a personTemporal Order: the top 60% holds 58.0% of every scene with a personTemporal Order: the top 70% holds 69.4% of every scene with a personTemporal Order: the top 80% holds 79.3% of every scene with a personTemporal Order: the top 90% holds 89.0% of every scene with a personTemporal Order: the top 100% holds 100.0% of every scene with a personPAI Ranking: the top 10% holds 28.8% of every scene with a personPAI Ranking: the top 20% holds 51.8% of every scene with a personPAI Ranking: the top 30% holds 68.8% of every scene with a personPAI Ranking: the top 40% holds 81.6% of every scene with a personPAI Ranking: the top 50% holds 90.3% of every scene with a personPAI Ranking: the top 60% holds 95.7% of every scene with a personPAI Ranking: the top 70% holds 98.5% of every scene with a personPAI Ranking: the top 80% holds 99.7% of every scene with a personPAI Ranking: the top 90% holds 100.0% of every scene with a personPAI Ranking: the top 100% holds 100.0% of every scene with a person82%35%10%20%30%40%50%60%70%80%90%100%How far down the ranking you go
Both lines reach 100%, because the ranking reorders your archive rather than shrinking it.
Table view
Depth of the rankingPAI RankingTemporal Order
Top 10%28.8%6.8%
Top 20%51.8%15.2%
Top 30%68.8%24.8%
Top 40%81.6%35.4%
Top 50%90.3%46.4%
Top 60%95.7%58.0%
Top 70%98.5%69.4%
Top 80%99.7%79.3%
Top 90%100.0%89.0%
Top 100%100.0%100.0%
The takeaway

Four bands in, you have reached most of the people in the archive. The rest of your footage is still there, in order, whenever you want it.

The Process

WITNESS measures the scene your vehicles already recorded

Send us your archive

We rank it. You stop throwing away label-relevant scenes.

  1. 01

    Send

    Send us your entire multi-sensor driving footage archive, every sensor stream intact.

  2. 02

    Rank

    We process it and rank every scene by its label-relevance.

  3. 03

    Return

    You get the archive back, sequenced and organized in order of priority.

The result accelerates discovery by 2.5x to 2.7x on all label-relevant material, and by 4.2x to 4.9x on the scenes with people in them. Faster through the same archive, and no more setting aside scenes that were label-relevant all along.

Connect

Physical AI learns faster when it studies the scenes that matter.