Dear Authors,
I followed the instructions in your README, including the fine-tuning step for the LLM, followed by inference and self-improvement using the 1B model. However, I obtained a performance of around 56% to 60%, rather than the reported 73.5%.
I also followed the hyperparameters and settings provided in the paper as closely as possible. Could you please let me know if there are any additional settings, checkpoints, or implementation details that are necessary to reproduce the reported 73.5% result?
I would greatly appreciate any guidance on how to reproduce the reported results.
Dear Authors,
I followed the instructions in your README, including the fine-tuning step for the LLM, followed by inference and self-improvement using the 1B model. However, I obtained a performance of around 56% to 60%, rather than the reported 73.5%.
I also followed the hyperparameters and settings provided in the paper as closely as possible. Could you please let me know if there are any additional settings, checkpoints, or implementation details that are necessary to reproduce the reported 73.5% result?
I would greatly appreciate any guidance on how to reproduce the reported results.