Browser demo of our Clash Royale RL environment: a 5.6k-parameter REINFORCE policy learns defensive placement against a brute-force optimum [P]
Yesterday I shared our open-source Clash Royale simulator and its recurrent PPO agent here. A training loop is easier to understand when you can watch it, so we put a small interactive version online: https://itzik123.gi
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/MachineLearning
Originally published by r/MachineLearning. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.