You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Dec 2, 2025. It is now read-only.
Hello
I'm running the model locally with a llama 7b model using llama-cpp-python compiled with cublas (gpu is working).
Whenever two messages are sent by a user before the AI sends a response, the whole program crashes.
This might be fixed by queueing messages and generating messages one after the other, or by simply ignoring new messages while creating.