Update: Yandex/AliceAI 80B-A3B fine tune progress
loss curve (taken from the last micro of every step, to explain the variation) some help from gemini 3.8 flash high About 40% of the way done with the initial fine tune. The loss is so spiky because I accidentally used t
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/LocalLLaMA
Originally published by r/LocalLLaMA. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.