Are we assuming that this is the companies themselves scraping data from training or is this "agents" acting on behalf of users? Nowadays every major chat UI (ChatGPT, Claude etc) has a "tool" that allows LLM to load web pages, so it must generate some traffic.
No, most of the load comes from armies of residential IP addresses that look like Google Chrome on the wire. The major chat UIs properly identify themselves. These waves of attacks do not.