Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion kubernetes/ai/llmkube/models/embed.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ spec:
# CPU embedding service for the ToolHive gateway tool search.
# Bound concurrent embedding work: eight slots and 2048-token batches
# exceeded the 4Gi limit during ToolHive indexing.
image: ghcr.io/ggml-org/llama.cpp:server@sha256:af8c29600d1de84945de13fa7f126d06fdae16c4b19dc7db1c69c74287eae257
image: ghcr.io/ggml-org/llama.cpp:server@sha256:33868c035b21dc63f7c60b7438774283fd99215bc319114eb03de5df4ce7cd6b
replicas: 1
contextSize: 8192
parallelSlots: 2
Expand Down