Okay, but there are multi-modal models. So do those experience qualia of sound and images? Do the text based ones experience the qualia of reading, or of finally comprehending a difficult concept after struggling with it?
For that matter, what exactly is the difference between "seeing" that a car is red versus an LLM reading an array of pixel data? Your eyes "just" translate photons into analog signals after all.
I've already answered this. There's no way someone deaf from birth could know what it's like to hear sounds, by looking at the pixels in a sound recording they know the format of, and if you don't believe me you can ask them. Or by looking at a speech spectrogram (which, with practice, are more easily readable).
Or you can ask yourself what ultraviolet looks like, or what it's like to sense a magnetic field. Lots of animals can do these, and would know what it's like, but people can't.