There are issues when running the service backend on low RAM servers (1GB and below). Consider removing tiktoken to help with this. We are currently using it to evenly split requests. Check if using a simpler metric such as simply counting words wouldn't be enough.
There are issues when running the service backend on low RAM servers (1GB and below). Consider removing tiktoken to help with this. We are currently using it to evenly split requests. Check if using a simpler metric such as simply counting words wouldn't be enough.