Currently, the open-source version requires configuring APIs for the judge and search functions.
The estimated API cost for a single RL training run is around 1w RMB, which makes this work impractical for academic researchers to replicate.
Would you consider releasing an academic-friendly variant with locally deployed judge models and on-premise search libraries?
Currently, the open-source version requires configuring APIs for the judge and search functions.
The estimated API cost for a single RL training run is around 1w RMB, which makes this work impractical for academic researchers to replicate.
Would you consider releasing an academic-friendly variant with locally deployed judge models and on-premise search libraries?