routing llm - problem 1 : static finetuning problem
Can a small model be finetuned or altered to increase the frequency of queries that are sent to it relative to a large model? The notion of ‘easy’ queries (queries that have nearly the same performance for both large and small models) was introduced in [Wan+24]. Is it possible for small models to get more share of routing?