mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
5e48486ad8
* test: daemon leak oracle around the real-subprocess daemon lanes (#1781 B1) Adds test/integration/support/daemon-leak-oracle.ts and calls it at the end of smoke-daemon-clean, smoke-daemon-http and daemon-replace-exit-flush. After shutdown the oracle asserts that no daemon-owned process survives (ownership: PPID descendant, PGID = daemon pid, AGENT_DEVICE_STATE_DIR env, state-dir argv — never global counts) and that the isolated state dir holds only classified artifacts (no *.tmp, no daemon.json/lock without a live daemon, no open capture descriptor). Red-proofs: pre-fix #1324 (a84caa818) leaves simctl recordVideo with ppid 1 / pgid = dead daemon; pre-fix #1109 (be4bd092b) leaves the agent-browser daemon plus its Chrome fleet. Both are clean on main. Refs #1781 #1431 * test: fail the daemon leak oracle on a surviving daemon and pin its rules Review of #1859 found the oracle reported "clean" when the daemon itself outlived shutdown: daemon pids were excluded from ownership, liveDaemonPids never reached hasDaemonLeaks, and a live daemon even flipped daemon.json/lock from stray to expected. stopProcessForTakeover is best-effort void, so only smoke-daemon-clean independently asserted death. - a live daemon pid at phase 'after-shutdown' is now itself a leak, and its metadata files stay stray; a live daemon remains legitimate at 'after-close' - split the pure ownership/residue rules into daemon-leak-model.ts and pin them with daemon-leak-model.test.ts, using the real ps rows captured during the #1109 and #1324 red-proofs (the lanes' daemons own no children, so the fixture test is what guards those shapes in CI) - exempt the managed tools/ install tree before the .tmp rule, so agent-browser's own download temporaries are no longer a false LEAK - report empty directories as residue (an unswept session scaffold leaves no file) - reuse src/utils/host-process.ts (expandProcessTree, uniquePositivePids) and its /bin/ps convention instead of re-deriving them - assert in smoke-daemon-http on the success path, not in finally, so the settle window cannot replace a primary assertion's diagnostic Refs #1781 #1431 * test: make the leak oracle's after-close phase name the session that closed Re-review of #1859: `after-close` accepted every capture descriptor and legacy marker, so the closed session's unfinalized handle was indistinguishable from another session's legitimately live one — the phase could not fail, the same shape as the surviving-daemon blocker, and it left half of B1's stated scope undelivered. - the observation carries the session directories closed at the checkpoint; a capture handle must be finalized once its owning session is gone (every session after shutdown, only the closed ones after close), while another session's live handle stays expected and a legacy pid marker is never a finish record - the oracle accepts `closedSessions` and a `sessionsDir` override, normalizing entries to the canonical `sessions/<name>/…` shape so an in-process harness rooted directly at its own sessions dir is classified the same way - add a real route regression: a provider-backed session with a live screen recording is closed through the daemon route, and the oracle refuses a descriptor left `lifecycle: "open"`. Reverting the close-route finalization (session-close-lifecycle-teardown.ts, the #1325 fix) turns it red naming sessions/default/screen-recording.resource.json; with the pre-fix model it stays green, which is the P1 in one line Refs #1781 #1431 * test: require the closed-session identity at the leak oracle's owning interface Re-review of #1859: `closedSessions` was optional and defaulted to empty, so any caller — or the standalone CLI invoked with `--phase after-close` and no `--closed-session` — silently restored the vacuous checkpoint that accepts every unfinalized capture handle and reports clean. - phase and the identity that phase needs are now one discriminated shape: { phase: 'after-close'; closedSessions: NonEmpty<string> } | { phase: 'after-shutdown' }, so the empty case is not expressible and 'after-shutdown' is unchanged - the CLI refuses the same invocation with a typed INVALID_ARGS error and a hint naming what the missing identity would have cost, instead of degrading - daemon-leak-oracle-cli.test.ts drives the real CLI: the unnamed after-close invocation must fail without reporting, the named one finds the planted unfinalized handle, and after-shutdown is unaffected. Reverting the guard makes it print 'clean (after-close)' over that same handle and turns the test red Refs #1781 #1431 * test: split the leak oracle's unrecorded process arm out of the lane path Maintainer asked whether the harness can be trimmed. The two arms differ in kind: the daemon already records captures and state-dir artifacts, so those rules are short and fire on every lane run; nothing records what a daemon spawned, so the process arm reconstructs ownership from the OS four ways — and the three CLI daemon lanes are device-free, so it only ever compared an empty set to an empty set. Move that arm to test/integration/support/daemon-owned-process-probe.ts, the manual script that produced the #1109/#1324 evidence, and keep its rules pinned by a fixture test (no lane can run the probe, so its fixtures are the only CI guard on those shapes). The shipped oracle keeps every guarantee it had: surviving daemon at after-shutdown, the closed-session finalization rule, the state-dir allowlist, the CLI refusal, and the phase-discriminated types. Lane-path harness 852 → 548 LOC (-36%); repo total roughly flat, because the ownership reconstruction can be relocated but not deleted. That is the argument for the follow-up: once the daemon records owned child pids the way it records captures, the arm collapses to "read the record, assert they are dead" and moves back into the oracle. Re-proved after the move: #1324 against pre-fixa84caa818on an iPhone 16 simulator still goes red through the probe (simctl recordVideo, ppid 1, pgid = the dead daemon, 0-byte mp4), and clean once the orphan is reaped. Refs #1781 #1431 * test: guard the daemon-owned-process arm in CI on the live web lane Re-review of ce9f0eeb: moving that arm to a manual probe left the two failures B1 exists to prevent detectable only by hand, and a fixture test proves regexes recognize synthetic rows, not that the shipped route reaps what it spawned. Restore the arm to the shipped oracle (the three files return byte-identical to 3a5b9bef3) and give it a lane that can execute it: smoke-web-platform is the one CI route whose daemon owns real children — the managed agent-browser daemon and its Chrome fleet. After the normal smoke it reopens a session, stops the daemon with that session still open (the #1109 shape: an ordinary close reaps the fleet, so only an unclosed session can strand it), and requires that nothing owned outlives the browser idle window. Proven both ways locally: green in 59s, and red when the fleet is stranded (browser idle window raised past the settle budget) with the oracle naming 15 owned processes — the #1109 signature, in a lane that runs on every PR. Also: daemon-replace-exit-flush cleared `info` before the oracle ran, so a failed stop would skip the `finally` retry and remove the state dir while the daemon was still alive. Clear it only once the checkpoint passes. Refs #1781 #1431 #1882 * docs: align daemon leak coverage rationale * refactor(test): narrow daemon oracle to durable state leaks
124 lines
4.9 KiB
TypeScript
124 lines
4.9 KiB
TypeScript
import test from 'node:test';
|
|
import assert from 'node:assert/strict';
|
|
import fs from 'node:fs';
|
|
import os from 'node:os';
|
|
import path from 'node:path';
|
|
import { skipWhenLoopbackUnavailable } from '../../src/__tests__/test-utils/loopback.ts';
|
|
import { stopProcessForTakeover } from '../../src/daemon/daemon-process.ts';
|
|
import { isProcessAlive } from '../../src/utils/host-process.ts';
|
|
import { assertNoDaemonLeaks } from './support/daemon-leak-oracle.ts';
|
|
import { runCliJson } from './test-helpers.ts';
|
|
|
|
// #1596: a CLI command that finds its recorded daemon unreachable replaces it
|
|
// (`Replacing daemon (pid N, vX) in <state-dir>: unreachable`) and retries
|
|
// against a fresh one. Three field runs died with zero further agent actions
|
|
// immediately after that replace plus a SESSION_NOT_FOUND (the fresh daemon
|
|
// has no sessions yet, which is expected). This file locks down that a
|
|
// replace-mid-command always ends in a normal, fully-delivered structured
|
|
// error rather than a truncated or hung process.
|
|
|
|
type DaemonInfo = {
|
|
pid: number;
|
|
processStartTime?: string;
|
|
};
|
|
|
|
test('daemon replace mid-command returns a structured, parseable error and exits normally', async (t) => {
|
|
if (await skipWhenLoopbackUnavailable(t)) {
|
|
return;
|
|
}
|
|
|
|
const stateDir = fs.mkdtempSync(path.join(os.tmpdir(), 'agent-device-replace-exit-flush-'));
|
|
let info: DaemonInfo | null = null;
|
|
const daemonPids: number[] = [];
|
|
try {
|
|
// A real daemon, started by this codebase, so its recorded version/code
|
|
// signature legitimately match — the only way to reach the "unreachable"
|
|
// takeover reason (as opposed to a version/signature mismatch takeover).
|
|
const started = runCliJson(['session', 'list', '--json', '--state-dir', stateDir]);
|
|
assert.equal(started.status, 0, `${started.stderr}\n${started.stdout}`);
|
|
|
|
info = readDaemonInfo(stateDir);
|
|
daemonPids.push(info.pid);
|
|
assert.equal(isProcessAlive(info.pid), true, 'expected the started daemon to be alive');
|
|
|
|
// Kill it out from under its own metadata: daemon.json stays put and
|
|
// still points at a pid that is now unreachable, reproducing the crash
|
|
// the field transcripts observed.
|
|
process.kill(info.pid, 'SIGKILL');
|
|
await waitForProcessDeath(info.pid);
|
|
|
|
const result = runCliJson(['close', '--json', '--state-dir', stateDir]);
|
|
|
|
assert.equal(result.status, 1, formatUnexpected('exit code', result));
|
|
assert.ok(
|
|
result.stderr.includes('Replacing daemon') && result.stderr.includes('unreachable'),
|
|
formatUnexpected('takeover notice on stderr', result),
|
|
);
|
|
assert.ok(result.json, formatUnexpected('parseable JSON stdout', result));
|
|
assert.equal(result.json.success, false, formatUnexpected('success:false', result));
|
|
assert.equal(
|
|
result.json.error?.code,
|
|
'SESSION_NOT_FOUND',
|
|
formatUnexpected('SESSION_NOT_FOUND', result),
|
|
);
|
|
// #1596 requirement: a hint pointing at `open` is always present, not
|
|
// just "fresh daemon, good luck" — this is the daemon.json truthfully
|
|
// having no sessions, which is expected; only the error's shape/delivery
|
|
// was ever in question.
|
|
assert.match(
|
|
result.json.error?.hint ?? '',
|
|
/open/i,
|
|
formatUnexpected('an `open` hint', result),
|
|
);
|
|
|
|
info = readDaemonInfo(stateDir);
|
|
daemonPids.push(info.pid);
|
|
await stopProcessForTakeover(info.pid, {
|
|
termTimeoutMs: 1_500,
|
|
killTimeoutMs: 1_500,
|
|
expectedStartTime: info.processStartTime,
|
|
});
|
|
// #1781 B1: neither the SIGKILLed daemon nor its replacement may leave owned
|
|
// processes or unclassified state-dir residue once both are gone. `info`
|
|
// stays set until this passes: `stopProcessForTakeover` is best-effort, so a
|
|
// failed stop must still reach the `finally` retry below rather than have
|
|
// the state dir removed out from under a daemon that is still running.
|
|
await assertNoDaemonLeaks({ stateDir, daemonPids, phase: 'after-shutdown' });
|
|
info = null;
|
|
} finally {
|
|
if (info) {
|
|
await stopProcessForTakeover(info.pid, {
|
|
termTimeoutMs: 1_500,
|
|
killTimeoutMs: 1_500,
|
|
expectedStartTime: info.processStartTime,
|
|
});
|
|
}
|
|
fs.rmSync(stateDir, { recursive: true, force: true });
|
|
}
|
|
});
|
|
|
|
async function waitForProcessDeath(pid: number): Promise<void> {
|
|
const deadline = Date.now() + 5_000;
|
|
while (Date.now() < deadline) {
|
|
if (!isProcessAlive(pid)) return;
|
|
await new Promise((resolve) => setTimeout(resolve, 25));
|
|
}
|
|
assert.fail(`daemon pid ${pid} did not die after SIGKILL`);
|
|
}
|
|
|
|
function readDaemonInfo(stateDir: string): DaemonInfo {
|
|
return JSON.parse(fs.readFileSync(path.join(stateDir, 'daemon.json'), 'utf8')) as DaemonInfo;
|
|
}
|
|
|
|
function formatUnexpected(
|
|
expected: string,
|
|
result: { status: number; stdout: string; stderr: string },
|
|
): string {
|
|
return [
|
|
`expected ${expected}`,
|
|
`status: ${result.status}`,
|
|
`stdout: ${result.stdout || '(empty)'}`,
|
|
`stderr: ${result.stderr || '(empty)'}`,
|
|
].join('\n');
|
|
}
|