Currently, models are filtered based on the results from OpenLLM Benchmark. However this is already archived and do not support new models. TODOs: - [ ] Extend to other recent benchmarks For example: https://artificialanalysis.ai/leaderboards/models - [ ] For each benchmark, define a mapper - from the platform tasks to the different datasets in the benchmark - [ ] More filters including: - inference time - language - co2 emissions, etc.
Currently, models are filtered based on the results from OpenLLM Benchmark. However this is already archived and do not support new models.
TODOs:
For example: https://artificialanalysis.ai/leaderboards/models
- inference time
- language
- co2 emissions, etc.