How do we know the models are writing memory-safe code? How will people who haven’t written it audit the output?