mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
258820e4ce
The self-hosted `run-demo.sh` path launches a native Metal `text-embeddings-router` on :7067 for the durable-memory demo. TEI's default `--max-batch-tokens` (16384) can fault the Metal backend during its warmup forward pass on some Apple Silicon machines. The process then either deadlocks (every thread parked in a pthread cond wait at 0% CPU) or dies silently with no panic — a GPU-level abort — so it never binds :7067 and the 300s health wait times out. The demo appears to "crash" with no actionable error. Pass `--max-batch-tokens 512` so warmup uses a small forward pass, which clears reliably. This only bounds per-request tokens (memory texts are short), not the embedding vectors, so recall stays byte-identical to the docker/CI embedder. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>