LocalAI

mirror of https://github.com/mudler/LocalAI synced 2026-04-21 13:27:21 +00:00

History

Ettore Di Giacinto 44e7d9806b fix(distributed): stop queue loops on agent nodes + dead-letter cap pending_backend_ops rows targeting agent-type workers looped forever: the reconciler fan-out hit a NATS subject the worker doesn't subscribe to, returned ErrNoResponders, we marked the node unhealthy, and the health monitor flipped it back to healthy on the next heartbeat. Next tick, same row, same failure. Three related fixes: 1. enqueueAndDrainBackendOp skips nodes whose NodeType != backend. Agent workers handle agent NATS subjects, not backend.install / delete / list, so enqueueing for them guarantees an infinite retry loop. Silent skip is correct — they aren't consumers of these ops. 2. Reconciler drain mirrors enqueueAndDrainBackendOp's behavior on nats.ErrNoResponders: mark the node unhealthy before recording the failure, so subsequent ListDuePendingBackendOps (filters by status=healthy) stops picking the row until the node actually recovers. Matches the synchronous fan-out path. 3. Dead-letter cap at maxPendingBackendOpAttempts (10). After ~1h of exponential backoff the row is a poison message; further retries just thrash NATS. Row is deleted and logged at ERROR so it stays visible without staying infinite. Plus a one-shot startup cleanup in NewNodeRegistry: drop queue rows that target agent-type nodes, non-existent nodes, or carry an empty backend name. Guarded by the same schema-migration advisory lock so only one instance performs it. The guards above prevent new rows of this shape; this closes the migration gap for existing ones. Tests: the prune migration (valid row stays, agent + empty-name rows drop) on top of existing upsert / backoff coverage.		2026-04-19 21:27:05 +00:00
..
advisorylock	feat(distributed): durable backend fan-out + state reconciliation	2026-04-19 08:34:57 +00:00
agentpool	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
agents	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
dbutil	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
distributed	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
finetune	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
galleryop	feat(ui): surface backend upgrades in the System page	2026-04-19 08:14:49 +00:00
jobs	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
mcp	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
messaging	fix(distributed): detect backend upgrades across worker nodes	2026-04-19 08:03:20 +00:00
monitoring	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
nodes	fix(distributed): stop queue loops on agent nodes + dead-letter cap	2026-04-19 21:27:05 +00:00
quantization	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
skills	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
storage	feat: track files being staged (#9275 )	2026-04-08 14:33:58 +02:00
testutil	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00