Interesting. Have you repeated these experiments with recent models? I'm thinking frontier models APIs have tools/MCPs for math stuff but curious about recent Qwen models, etc.