logoalt Hacker News

xpct • today at 9:02 PM • 1 reply • view on HN

You can definitely try to regularize against ruleset changes by generating a bunch of cards and making the agent play in randomized subsets of those cards.

I didn't look for prior work on this, but my estimate is that it's probably within 2-3 orders of magnitude of additional training compared to a static game. (Still a lot!)


Replies

hnedeotes • today at 9:22 PM

But wouldn't (couldn't) the model then hallucinate play patterns and get itself into problems when playing against a real opponent?

➕ show 1 reply