Tesla's already solved this - their vision model does this phenomenally well.
And they've demonstrated adding a sidecar LLM to it as well, mostly for these kinds of "read these 3 street signs, what should i do next?" sort of situations.
The same Tesla that pulled radar to go vision only and a person was killed because the vision model didn't recognize a truck? https://www.bbc.com/news/technology-36680043
Not sure that counts as phenomenally well.
Totally solved: https://electrek.co/2026/09/15/tesla-four-fatal-driver-assis...