After determining sections we might benefit from semantically segmenting the section into say publications or presentations or whatever. This does not need to happen for everything but for sections maybe that is bound to have a lot of different items.
These can be:
- publications
- grants
- presentations
- mentorship
- maybe education
For this we can start with a very strong model that is locally runnable. I'm thinking Qwen/Qwen3-VL-Embedding-8B and use model2vec to make it faster.
Model2vec and native options need to be tested for performance. Since ingestion is a one time thing we can probably afford to use something expensive.
After determining sections we might benefit from semantically segmenting the section into say publications or presentations or whatever. This does not need to happen for everything but for sections maybe that is bound to have a lot of different items.
These can be:
For this we can start with a very strong model that is locally runnable. I'm thinking
Qwen/Qwen3-VL-Embedding-8Band use model2vec to make it faster.Model2vec and native options need to be tested for performance. Since ingestion is a one time thing we can probably afford to use something expensive.