mirror of
https://github.com/vectorize-io/hindsight.git
synced 2026-09-14 19:31:49 +08:00
20bd4f3618
* fix(ci): compare OpenAPI against the merge-base; build benchmark role configs whole Two unrelated CI failures, both of which fail without anything being wrong with the code under test. OpenAPI compatibility diffed the branch's spec against the LIVE tip of the base branch, so every endpoint main gained after a branch was cut is reported as "Endpoint removed (breaks old clients)" by that branch. Three open PRs failed this way today on /health/live and /health/ready (added by #3329), none of which touch the spec at all; the only cure was an unrelated rebase. Compare against `git merge-base origin/$BASE_BRANCH HEAD` instead, which asks the question the check means to ask: did *this branch* remove something. Genuine removals still fail — verified both directions against the real specs. The scheduled LoComo benchmark has failed every night since at least Aug 8, before its first question: `LoComoAnswerGenerator()` raised "HINDSIGHT_API_LLM_VERTEXAI_PROJECT_ID is required". The workflow does export that variable — but each benchmark role built its LLMConfig from exactly four env vars (provider, api_key, base_url, model), and LLMConfig deliberately does not read the environment for provider-specific settings (from_env is where the API resolves them). Every Vertex AI value was therefore dropped. The same four-var construction was copy-pasted at three sites — locomo, longmemeval and the shared judge — so fixing only the crashing one would have moved the failure down a line. Replace all three with a shared builder that carries the Vertex AI project/region/service-account through. hindsight-dev/tests had no CI job, which is why a plain construction bug was left for a nightly benchmark to find hours later. Add one, so those tests (and the new regression test) actually run on PRs. * fix(ci): make the benchmark role-config tests hermetic They passed locally off the developer's HINDSIGHT_API_LLM_API_KEY and failed in the new test-dev job, where no key exists, with "API key is required for openai" — the tests were reading ambient environment instead of declaring what they need. Clear every HINDSIGHT_API_*LLM* var before each test and set the ones under test explicitly.