LocalAI

mirror of https://github.com/mudler/LocalAI synced 2026-05-24 09:28:23 +00:00

History

LocalAI [bot] 22ff86d64f fix(distributed): round-robin replicas of the same model (#9695 ) FindAndLockNodeWithModel previously ordered candidate replicas by in_flight ASC, available_vram DESC. The primary key is correct, but the tiebreaker meant that whenever in_flight tied — the common case at low to moderate concurrency where requests don't overlap — the node with the largest available VRAM won every pick. With autoscaling placing replicas of the same model on multiple nodes, the fattest GPU node ended up taking nearly all the load while the others sat idle. Insert last_used ASC between the two existing tiers. last_used is already refreshed inside the same transaction that increments in_flight (and by TouchNodeModel on cache hits in the router), so the "oldest-used" replica naturally rotates through the candidate set — strict round-robin without a schema change. available_vram DESC is demoted to a final tiebreaker for cold starts where last_used is identical across replicas. Placement queries (FindNodeWithVRAM, FindLeastLoadedNode, and the *FromSet variants) have the same fattest-GPU bias on tiebreakers but are higher-cost to fix consistently. Deferred to a follow-up so the routing fix can land first — for the user-observed symptom routing was the dominant cause anyway. Test: registry_test.go adds a focused spec that loads three replicas on three nodes with 24/16/8 GB VRAM and asserts each is picked at least twice across 9 in_flight-tied calls. Assisted-by: claude-code:claude-opus-4-7 [Read] [Edit] [Bash] [Grep] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>		2026-05-06 19:40:54 +02:00
..
application	feat(gallery): Speed up load times and clean gallery entries (#9211 )	2026-05-06 14:51:38 +02:00
backend	fix(backend): resolve relative draft_model paths against the models dir (#9680 )	2026-05-06 00:58:38 +02:00
cli	feat: support word-level timestamps for faster-whisper (#9621 )	2026-05-06 00:32:52 +02:00
clients	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
config	feat(gallery): Speed up load times and clean gallery entries (#9211 )	2026-05-06 14:51:38 +02:00
dependencies_manager	feat(ui): move to React for frontend (#8772 )	2026-03-05 21:47:12 +01:00
explorer	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
gallery	feat(gallery): Speed up load times and clean gallery entries (#9211 )	2026-05-06 14:51:38 +02:00
http	feat(gallery): Speed up load times and clean gallery entries (#9211 )	2026-05-06 14:51:38 +02:00
p2p	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
schema	feat: support word-level timestamps for faster-whisper (#9621 )	2026-05-06 00:32:52 +02:00
services	fix(distributed): round-robin replicas of the same model (#9695 )	2026-05-06 19:40:54 +02:00
startup	feat: add distributed mode (#9124 )	2026-03-30 00:47:27 +02:00
templates	fix(vision): propagate mtmd media marker from backend via ModelMetadata (#9412 )	2026-04-18 20:30:13 +02:00
trace	feat: add LocalVQE backend and audio transformations UI (#9640 )	2026-05-04 22:07:11 +02:00