You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Separate model (same architecture) that's fine tuned to output "true" rewards (or maybe distribution of probabilities, e.g. beta distribution), using supervised learning from labelled datasets.
Adjust the quantilizer function to allow for this (inc. do maths)