Replies: 1 comment
|
That error is the gRPC connection between LocalAI core and the model backend breaking — "Unavailable / error reading from server: EOF" means the llama.cpp backend process died or wedged mid-connection. Slow degradation over days on a box hosting several models is the classic pattern: a backend that's either leaking memory or silently hung, and everything routed to it fails until the process is recycled. Before you schedule nightly reboots, there's a built-in answer for exactly this: the watchdog.
(both also settable via If you want to actually find the leak instead of managing it: when it next happens, check With the idle watchdog enabled you may find the problem simply stops occurring, since nothing sits around long enough to wedge. If it turns out to be a real leak in a specific model's backend, the RSS/dmesg evidence is what the maintainers will need. If this gets it stable, marking the answer helps the next person whose server quietly rots after a few days. |
Uh oh!
There was an error while loading. Please reload this page.
I'm running a localai server hosting a few models via llama.cpp (primarily qwen3.8-27b-q4, gemma4-31b-it, and mistral-nemo-instruct-2407) on a dgx spark (actually a lenovo pgx but same cpu/gpu and specs.) It works fine for a couple days in a row, but eventually degrades to a state where I cant get responses and just get this error in my agents. Not really sure how to troubleshoot or prevent it (beyond a nightly reboot or something, which I guess I could do but feels like kicking the can). I didn't see anything obvious in the error logs, but will capture them the next time it occurs if it would be helpful
All reactions