Files
Michael Ramos 6ec1a66c9b feat(review): large-PR pipeline, instant-open checkout, scroll perf, and worker-pool highlighting (#893)
* feat(review): large GitHub PR fallback + non-blocking PR checkout

Two PR-mode improvements:

1. Large GitHub PRs no longer fail to load. When `gh pr diff` is refused
   (HTTP 406 for oversized diffs), fetchGhPR pages through the pulls files
   API and stitches the per-file patches into a unified diff — mirroring
   the existing GitLab raw_diffs fallback. Path quoting matches git's
   exact rules (bare spaces unquoted) so downstream parsers round-trip;
   truncation at the API's 3000-file cap is surfaced, never silent.

2. The --local worktree/clone no longer blocks startup. The review server
   opens as soon as the platform diff arrives; the checkout warms in the
   background as a seeded not-ready pool entry. Consumers that need real
   files (agent jobs, full-stack diff, code-nav, semantic diff, AI
   sessions) await pool.ensure(), with creations serialized so concurrent
   fetches can't clobber the shared FETCH_HEAD. Cross-repo clone steps
   converted from spawnSync to async spawns; warmup children are killed
   on exit (plus `git worktree prune`) so aborted sessions can't leak
   stale registrations; failed checkouts degrade honestly (no agent runs
   in the wrong directory claiming local access) with a 30s retry
   cooldown.

* fix(review): survive long PR checkout warmups + classify reconstructed renames

Stress-testing against oven-sh/bun#30412 (2,188 files) surfaced three bugs:

- Bun.serve's default 10s idleTimeout killed /api/semantic-diff while it
  parked on the background checkout warmup (a clone that can take minutes).
  Disable the idle timeout on all servers — AI SSE streams can also stall
  >10s between bytes while a permission prompt waits.
- The file-badge hook memoized that failed fetch in a module-level cache
  keyed by patch, pinning every badge to empty until a hard refresh. Never
  cache failures; retry with backoff (5s/15s/30s).
- reconstructGhPatch/reconstructPatch omitted the `similarity index` line,
  which Pierre's parser keys rename classification off — pure renames
  rendered as blank plain changes with no old path. Emit 100% for
  patch-less renames/copies (exactly accurate) and a synthetic 99% for
  patched ones (consumers only branch on 100% vs not).

* feat(review): local full-diff upgrade for PRs whose API diff is truncated

On oversized PRs the platform APIs withhold per-file patch content entirely
(bun#30412: 1,066 of 2,188 files came back with status added/modified, zeroed
counts, and no patch). Those files rendered as empty stubs with no diff.

- fetchGhPR/fetchGlMR flag the result `patchIncomplete` when patch-less
  non-rename entries exist or the 3000-file cap truncates the listing.
- New runPRLayerLocalDiff (pr-stack.ts) recomputes the exact layer diff in
  the local checkout: platform merge-base + head SHA two-dot diff (three-dot
  vs baseSha fallback), fetch-by-SHA for objects missing from shallow clones,
  -l0 so rename detection doesn't silently degrade on huge PRs.
- The review UI shows a "Partial diff · Load full diff" notice in layer
  scope; clicking re-requests the layer scope and the server swaps in the
  recomputed full diff (waiting out the background clone if needed).
- PR scope/switch state writes are epoch-guarded: a request parked on the
  checkout warmup can no longer overwrite a newer scope select or pr-switch.
- draftKey follows the upgraded patch so annotation drafts survive pr-switch
  round-trips; recompute failures surface in the response error field.
- Pi server mirrors all of it, including an agentCwd fallback so the upgrade
  works for PRs switched-to under a cross-repo clone pool.

* fix(review): use GitLab's too_large/collapsed flags for withheld-diff detection

External review caught a false negative: a too-large ADDED file comes back
new_file:true with an empty diff — indistinguishable from a legitimately
empty new file under the old heuristic, so the partial-diff upgrade was
never offered for exactly the files that matter most on big MRs.

The REST /diffs endpoint marks withheld content explicitly per entry
(verified against gitlab.com): too_large/collapsed are now authoritative in
both directions — withheld adds/deletes are flagged, binaries and empty
files are never misflagged. Older GitLab without the fields keeps the
empty-diff-on-modification heuristic.

* feat(prompts): unify review-denied suffix — triage first, no coding off raw feedback

The per-runtime defaults map (#627) gave OpenCode and Pi a different
review-denied suffix than every other runtime; updating one meant the
others silently kept "you must address all of them" — an instruction to
start coding immediately. Claude Code, Amp, Droid, Codex, Copilot, Gemini,
and Kiro were all still on it.

One default for every runtime now: triage the feedback, verify it against
the code, discuss before changing anything. Per-runtime customization
remains available via config (prompts.review.runtimes.<rt>.denied), which
resolves above the built-in default as before.

* fix(prompts): generalize review-denied suffix — 'from review', not 'external AI reviewers'

Review feedback isn't always from AI reviewers or agent jobs; often it's
the human reviewer's own annotations. Neutral wording covers both.

* fix(review): non-blocking 'Load full diff' + flag-handling hardenings

Self-review findings:

- The partial-diff upgrade reused the scope-switch handler, so clicking
  "Load full diff" raised the full-screen PRSwitchOverlay — blocking the
  entire UI, potentially for minutes behind a cold clone, with no text and
  no cancel. The upgrade now has its own loading state: the notice shows a
  spinner ("Loading full diff…") and the reviewer keeps working with the
  partial diff while the request parks. Server-side epoch guards already
  handle scope/PR changes made during the wait.
- GitLab too_large/collapsed: treat explicit null like absent (flags
  inconclusive → legacy heuristic decides) instead of silently exonerating.
- Rename-limit lift uses -l100000 instead of -l0 ("0 = unlimited" only
  holds on git >= 2.29; on older git it could disable detection outright).

* fix(review): stop scroll-driven sem stampede when semantic diff is failing

The badge retry change (a2d19a4e) cleared the client-side sem cache on
failure so transient errors could recover. But file-header badges mount and
unmount on every scroll in the virtualized all-files view, and each mount
re-requests /api/semantic-diff — and the server only cached SUCCESSFUL runs.
With sem erroring, scrolling spawned a continuous stream of sem processes,
pegging the CPU and making scrolling severely choppy.

Bound retry rate by time, not by mount events:
- client: keep the failed result memoized and expire it after a 60s
  cooldown instead of clearing immediately
- server (Bun + Pi): memoize failed sem runs for 30s in
  SemanticDiffResponseCache — request rate can no longer drive execution
  rate

* fix(review): eliminate all-files scroll jank (pre-existing on main, from #885)

The CodeView migration introduced severe scroll chop; scrolling UP could
freeze the viewport entirely ("scrolling but nothing changes"). Three
compounding causes, diagnosed against Pierre 1.2.8 source:

1. Lazy full-content augmentation landed updateItem() mid-scroll-gesture:
   the full-content parse counts collapsed-context regions the raw-patch
   parse doesn't, so the item GROWS — re-render + re-tokenize hitches both
   directions, and when the grown item sat above CodeView's scroll anchor,
   its corrective scrollTo() killed wheel momentum (the up-scroll freeze).
   Fetches still start as items enter the window; the item mutation now
   waits for 150ms of scroll quiet (staleness re-checked at apply time).

2. reportVisibleFile read container.scrollTop/clientHeight/scrollHeight on
   EVERY scroll event — a forced synchronous layout right after each
   frame's DOM writes. Replaced with CodeView's cached accessors and
   coalesced the handler to once per animation frame.

3. Missing containment CSS: Pierre's own production wrapper uses
   contain:strict + will-change:scroll-position so forced layouts stay
   scoped to the scroller instead of the whole document. Adopted.

Also: __devOnlyValidateItemHeights now requires explicit opt-in
(VITE_PIERRE_VALIDATE_HEIGHTS=1) — it runs getBoundingClientRect() per
rendered item per frame and made dev-server scrolling choppy by itself.

* feat(review): change-type status in headers + tree, diffshub CSS parity

Adopts two diffshub practices identified in the architecture comparison:

- DiffFile now carries a derived status (added/deleted/renamed/modified)
  from the chunk's git metadata lines. FileHeader shows a status icon and
  renders renames as "old/path → new/path" (dimmed old, arrow — diffshub's
  treatment, including its rename blue); the file tree shows A/D/R markers.
  'modified' is deliberately undecorated so the others pop. Works in both
  the all-files surface and the single-file panel, including header-only
  pure renames from the large-PR reconstruction.

- CodeView container gains diffshub's remaining perf CSS: overflow-anchor:
  none (native scroll anchoring fights CodeView's own anchor resolution
  whenever item heights change — exactly our augmentation applies),
  overflow-x-clip, and overflow-clip containment on item elements.

* feat(review): worker-pool syntax highlighting (diffshub parity)

A performance trace of scrolling a small local diff attributed 2.2s of
2.6s main-thread CPU to findNextMatchSync — shiki's TextMate regex
scanner tokenizing on the main thread. diffshub avoids this entirely by
running tokenization in Pierre's worker pool; we never opted in.

Wires WorkerPoolContextProvider around the review app (pool size
min(cores-1, 3), 100-entry AST LRU, common languages preloaded), gates
the all-files surface on pool readiness with a 5s escape hatch (a dead
pool degrades to plaintext-then-highlight, never a blank view), and
syncs the UI theme pair into the long-lived pool.

Single-file build constraint solved with Vite's ?worker&inline (base64
blob worker) + worker.format 'es' with inlineDynamicImports — the
worker's lazy import("shiki/wasm") branch collapses into the bundle and
is never taken (shiki-js engine: the win is moving work off the main
thread, with no .wasm asset to smuggle into one HTML file). Bundle
+850KB.

* fix(review): un-poison worker-pool theme dedup on failed setRenderOptions

A failed round-trip recorded the theme as synced and never retried,
pinning the pool to the wrong palette for the session.

* fix(review): report partial diffs without a checkout; fail fast on missing checkout

Dogfood review of this PR (via plannotator itself) caught two valid issues:

- prPatchIncomplete was gated on the worktree pool, so a --no-local session
  showed a truncated diff with no indication at all. Partiality is
  information; upgradability is a capability. The flag is now always
  reported, with a separate prPatchUpgradeAvailable — the UI shows the
  amber notice either way, with the "Load full diff" button only when a
  checkout can exist (otherwise a "re-run with --local" hint).

- After a FAILED checkout warmup, Ask AI sessions and agent jobs fell back
  to process.cwd() (or a wrong revision on Pi) — running in the wrong tree
  instead of failing. Both launch points now refuse with a clear "Local
  PR checkout unavailable — retry shortly" error (503); the job handlers
  surface buildCommand refusals instead of mislabeling them "Invalid
  JSON". Bun and Pi mirrored.

A third finding (sem availability stuck after warmup) was triaged invalid:
the availability probe detects the sem binary, which is cwd-independent.

* fix(review): runtime-neutral copy for the no-checkout partial-diff hint

--local is a CLI remedy; OpenCode sessions have no such flag. Visible
text states the fact, the tooltip carries the CLI guidance.
2026-06-12 13:50:09 -07:00

350 lines
11 KiB
TypeScript

import { describe, expect, test } from "bun:test";
import {
getManagedSemBinaryPath,
getSemanticDiffAvailability,
parseSemVersion,
runSemanticDiff,
semanticDiffCacheKey,
semanticDiffFileExtsFromSearchParams,
SemanticDiffResponseCache,
type SemanticDiffRuntime,
} from "./semantic-diff";
import type { SemanticDiffResponse } from "./semantic-diff-types";
interface MockCommand {
version?: string;
diff?: string;
stderr?: string;
exitCode?: number;
}
function makeRuntime(options: {
cwd?: string;
env?: Record<string, string | undefined>;
files?: string[];
commands?: Record<string, MockCommand>;
pathDelimiter?: string;
platform?: NodeJS.Platform;
} = {}): SemanticDiffRuntime & { calls: Array<{ command: string; args: string[]; input?: string }> } {
const calls: Array<{ command: string; args: string[]; input?: string }> = [];
const files = new Set(options.files ?? []);
const commands = options.commands ?? {};
return {
calls,
env: options.env ?? {},
cwd: options.cwd ?? "/repo",
dataDir: "/home/user/.plannotator",
pathDelimiter: options.pathDelimiter ?? ":",
platform: options.platform ?? "linux",
fileExists(path) {
return files.has(path);
},
async runCommand(command, args, runOptions) {
calls.push({ command, args, input: runOptions?.input });
const mock = commands[command];
if (!mock) return { stdout: "", stderr: "not found", exitCode: 1, error: "not found" };
if (args.includes("--version")) {
return { stdout: mock.version ?? "", stderr: "", exitCode: mock.version ? 0 : 1 };
}
return {
stdout: mock.diff ?? "",
stderr: mock.stderr ?? "",
exitCode: mock.exitCode ?? 0,
};
},
};
}
describe("semantic diff runner", () => {
test("parses real sem version output and rejects other sem commands", () => {
expect(parseSemVersion("sem 0.8.0\n")).toBe("0.8.0");
expect(parseSemVersion("sem 0.8.0-dev+abc\n")).toBe("0.8.0-dev+abc");
expect(parseSemVersion("GNU parallel 20250122\n")).toBeNull();
});
test("reports unavailable when sem cannot be resolved", async () => {
const runtime = makeRuntime();
await expect(getSemanticDiffAvailability(runtime)).resolves.toMatchObject({
available: false,
reason: "sem-not-found",
});
});
test("validates PLANNOTATOR_SEM_PATH before using it", async () => {
const runtime = makeRuntime({
env: { PLANNOTATOR_SEM_PATH: "/missing/sem" },
});
await expect(runSemanticDiff({ rawPatch: "diff --git a/a.ts b/a.ts\n" }, runtime)).resolves.toMatchObject({
status: "unavailable",
reason: "sem-path-missing",
});
});
test("runs sem with patch input and normalized file extensions", async () => {
const runtime = makeRuntime({
env: { PLANNOTATOR_SEM_PATH: "mock-sem" },
commands: {
"mock-sem": {
version: "sem 0.8.0",
diff: JSON.stringify({
summary: { fileCount: 1, added: 1, modified: 0, deleted: 0, total: 1 },
changes: [
{
entityId: "src/a.ts::function::hello",
changeType: "added",
entityType: "function",
entityName: "hello",
filePath: "src/a.ts",
startLine: 3,
endLine: 5,
},
],
}),
},
},
});
const result = await runSemanticDiff({
rawPatch: "diff --git a/src/a.ts b/src/a.ts\n@@ -0,0 +1 @@\n+export function hello() {}\n",
fileExts: ["ts", ".tsx", "ts"],
}, runtime);
expect(result).toMatchObject({
status: "ok",
summary: { fileCount: 1, added: 1, total: 1 },
changes: [{ entityType: "function", entityName: "hello", filePath: "src/a.ts" }],
semVersion: "0.8.0",
semSource: "env",
});
expect(runtime.calls[1]).toMatchObject({
command: "mock-sem",
args: ["diff", "--patch", "--format", "json", "--file-exts", ".ts", ".tsx"],
input: expect.stringContaining("diff --git"),
});
});
test("does not pass a file extension filter unless one is requested", async () => {
const runtime = makeRuntime({
env: { PLANNOTATOR_SEM_PATH: "mock-sem" },
commands: {
"mock-sem": {
version: "sem 0.8.0",
diff: JSON.stringify({
summary: { fileCount: 0, added: 0, modified: 0, deleted: 0, total: 0 },
changes: [],
binaryChanges: [],
}),
},
},
});
await runSemanticDiff({
rawPatch: "diff --git a/src/a.py b/src/a.py\n@@ -1 +1 @@\n-a\n+b\n",
}, runtime);
expect(runtime.calls[1]).toMatchObject({
command: "mock-sem",
args: ["diff", "--patch", "--format", "json"],
});
});
test("does not run a sem package from the reviewed cwd", async () => {
const repoSem = "/repo/node_modules/@ataraxy-labs/sem/vendor/sem";
const runtime = makeRuntime({
cwd: "/server",
files: [repoSem],
commands: {
[repoSem]: {
version: "sem 0.8.0",
diff: JSON.stringify({
summary: { fileCount: 1, added: 1, modified: 0, deleted: 0, total: 1 },
changes: [
{
changeType: "added",
entityType: "function",
entityName: "fromRepoPackage",
filePath: "src/a.ts",
},
],
binaryChanges: [],
}),
},
},
});
await expect(runSemanticDiff({
rawPatch: "diff --git a/src/a.ts b/src/a.ts\n@@ -0,0 +1 @@\n+export function fromRepoPackage() {}\n",
cwd: "/repo",
}, runtime)).resolves.toMatchObject({
status: "unavailable",
reason: "sem-not-found",
});
expect(runtime.calls.map(call => call.command)).not.toContain(repoSem);
});
test("returns error instead of throwing when sem exits nonzero", async () => {
const runtime = makeRuntime({
env: { PLANNOTATOR_SEM_PATH: "mock-sem" },
commands: {
"mock-sem": {
version: "sem 0.8.0",
stderr: "parse failed",
exitCode: 2,
},
},
});
await expect(runSemanticDiff({ rawPatch: "diff --git a/a.ts b/a.ts\n" }, runtime)).resolves.toMatchObject({
status: "error",
reason: "sem-exit",
exitCode: 2,
message: "parse failed",
});
});
test("returns error instead of throwing when sem returns invalid JSON", async () => {
const runtime = makeRuntime({
env: { PLANNOTATOR_SEM_PATH: "mock-sem" },
commands: {
"mock-sem": {
version: "sem 0.8.0",
diff: "not json",
},
},
});
await expect(runSemanticDiff({ rawPatch: "diff --git a/a.ts b/a.ts\n" }, runtime)).resolves.toMatchObject({
status: "error",
reason: "invalid-json",
});
});
test("uses managed sidecar before PATH fallback", async () => {
const managed = getManagedSemBinaryPath("/home/user/.plannotator", "linux");
const runtime = makeRuntime({
files: [managed],
commands: {
[managed]: { version: "sem 0.8.0" },
sem: { version: "sem 0.9.0" },
},
});
await expect(getSemanticDiffAvailability(runtime)).resolves.toMatchObject({
available: true,
semVersion: "0.8.0",
semSource: "managed",
});
expect(runtime.calls[0].command).toBe(managed);
});
test("does not fall back to bare sem on Windows when PATH resolution misses", async () => {
const runtime = makeRuntime({
platform: "win32",
pathDelimiter: ";",
env: {
PATH: "C:/repo;C:/tools",
PATHEXT: ".EXE",
},
commands: {
sem: { version: "sem 0.8.0" },
},
});
await expect(getSemanticDiffAvailability(runtime)).resolves.toMatchObject({
available: false,
reason: "sem-not-found",
});
expect(runtime.calls).toEqual([]);
});
test("resolves an absolute sem.exe from PATH on Windows", async () => {
const semPath = "C:/tools/sem.exe";
const runtime = makeRuntime({
platform: "win32",
pathDelimiter: ";",
env: {
PATH: "C:/tools",
PATHEXT: ".EXE",
},
files: [semPath],
commands: {
[semPath]: { version: "sem 0.8.0" },
},
});
await expect(getSemanticDiffAvailability(runtime)).resolves.toMatchObject({
available: true,
semVersion: "0.8.0",
semSource: "path",
});
expect(runtime.calls[0].command).toBe(semPath);
});
test("cache key accounts for patch, cwd, and file extensions", () => {
const a = semanticDiffCacheKey({ rawPatch: "a", cwd: "/repo", fileExts: ["ts"] });
const b = semanticDiffCacheKey({ rawPatch: "a", cwd: "/repo", fileExts: [".ts"] });
const c = semanticDiffCacheKey({ rawPatch: "a", cwd: "/other", fileExts: [".ts"] });
expect(a).toBe(b);
expect(a).not.toBe(c);
});
test("parses requested file extensions without applying a default filter", () => {
expect(semanticDiffFileExtsFromSearchParams(new URLSearchParams())).toEqual([]);
expect(semanticDiffFileExtsFromSearchParams(new URLSearchParams("fileExt=ts&fileExts=.tsx,jsx"))).toEqual([
".ts",
".tsx",
".jsx",
]);
});
test("response cache clears when the patch changes and evicts oldest entries", () => {
const cache = new SemanticDiffResponseCache(1);
const first: SemanticDiffResponse = {
status: "unavailable",
reason: "sem-not-found",
message: "missing",
};
const second: SemanticDiffResponse = {
status: "error",
reason: "sem-exit",
message: "failed",
};
cache.set("a", "patch-a", first);
expect(cache.get("a", "patch-a")).toBe(first);
cache.set("b", "patch-a", second);
expect(cache.get("a", "patch-a")).toBeUndefined();
expect(cache.get("b", "patch-a")).toBe(second);
expect(cache.get("b", "patch-b")).toBeUndefined();
});
test("failures are memoized within their TTL and retryable after it", () => {
const cache = new SemanticDiffResponseCache();
const failure: SemanticDiffResponse = {
status: "error",
reason: "sem-exit",
message: "failed",
};
// Within TTL: served from memo — repeated requests must not re-run sem.
cache.setFailure("k", "patch-a", failure, 60_000);
expect(cache.get("k", "patch-a")).toBe(failure);
// Expired TTL: gone — the next request may retry.
cache.setFailure("k2", "patch-a", failure, -1);
expect(cache.get("k2", "patch-a")).toBeUndefined();
// A success overwrites and outlives the failure memo.
const ok = { status: "ok", changes: [], binaryChanges: [] } as unknown as SemanticDiffResponse;
cache.setFailure("k3", "patch-a", failure, 60_000);
cache.set("k3", "patch-a", ok);
expect(cache.get("k3", "patch-a")).toBe(ok);
// Patch change clears failure memos too.
cache.setFailure("k4", "patch-a", failure, 60_000);
expect(cache.get("k4", "patch-b")).toBeUndefined();
});
});