Adding memory to search instead of sampling in reward maximization tasks [R]
I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run. I find it rather funny that
š
This source provides headlines only. Use the button below to read the complete article on the original site.
š° Read the original article on r/MachineLearning
Originally published by r/MachineLearning. Aggregated on AIWithGhost for educational purposes ā full credit and traffic to the original publisher.