There are many Firefox features that use various models (mostly language models). I would like to be able to redirect the processing to model servers (local or otherwise) of my choice for all of these.
There are a number of reasons why I suggest this:
- Sometimes inbuilt models fall short: For example, the history semantic search relies on the MiniLM model. While this is ok for basic cases (and probably good for something that is built in to not be enormous), I have often found it to be subpar for more complex cases. I would love to be able to generate the embeddings with a locally hosted 0.6b or 2b embedding model or something and see if I get better results for more technical topics that fail miserably for the default model.
- Efficiency: if you use a bunch of programs and they all decide to run their own models directly, this is not the best. I would rather have one instance of models I use running and on a central server than use a bunch of different ones that are each used by only one program.
I'm sure other people have reasons too.
Implementing this would probably need to have some of the internal stuff redone in a generic way (e.g. if you expected to store vector embeddings of a specific size that won't work), but I think this is probably more maintainable in the long run anyway.