Maybe LLM's should be integrated with the software, so they could optimize to the real hardware and workload ?