logoalt Hacker News

jxcole • yesterday at 3:59 PM • 1 reply • view on HN

I think the point here is that LeCun was arguing that training on pure text would not grant spatial understanding. I believe most models are trained on spatial data as well, so you are both right.


Replies

brainwad • yesterday at 4:18 PM

Is that what he meant? He works on models with an explicitly spatial internal representation, whereas I was using a standard LLM that edited the provided plan by using a bajillion python calls to inspect small regions of the image at a time and then generate edits.