mirror of
https://github.com/jackwener/OpenCLI.git
synced 2026-09-14 18:25:42 +08:00
feat/evaluate-function-overload
218 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
14b1dc86fc | feat(browser): add function form page evaluate | ||
|
|
fa9b38cd92 |
feat(reddit): add whoami, home, subreddit-info read commands (#1491)
* feat(reddit): add whoami, home, subreddit-info read commands Closes gap against jackwener/rdt-cli — three commands the existing 17 reddit adapters were missing: - `reddit whoami` — show the currently logged-in identity (fields: Username, ID, Post / Comment / Total Karma, Account Created, Gold, Mod, Verified Email, Has Mail, Inbox Count). Probes `/api/me.json` with two-pronged auth detection (401/403 OR `data.name` missing on 200 — Reddit returns 200 with an empty body for stale anon sessions, see PR #1428). - `reddit home` — personalized Best feed (`/best.json`). Distinct from the public `frontpage`/`r/all` command: enforces login via the same two-pronged auth check rather than silently degrading to the unauthenticated default feed. `--limit` accepts [1, 100] — out-of-range raises `ArgumentError` before navigation, no silent clamp. - `reddit subreddit-info` — subreddit metadata (Name, Title, Subscribers, Active Now, NSFW, Type, Description, Created, URL) from `/r/<X>/about.json`. Banned / private / quarantined / 404 subreddits raise `EmptyResultError` so the output table never holds a silent sentinel row. All three use Strategy.COOKIE + siteSession:'persistent' matching the existing reddit adapters, validate args upfront before `page.goto`, and use the 5-kind discriminated-union pattern (kind: auth/http/missing/ exception/ok) from PR #1428 to map page.evaluate results to typed errors on the Node side. Intermediate object keys deliberately avoid the declared columns (`field`/`value`/`rank`/etc.) per the silent-column-drop audit sediment from PR #1329. Tests: 28 new (whoami 6, home 9, subreddit-info 13); full reddit suite 38/38. Audits: typed-error-lint 189/189 (0 new), silent-column-drop 103/103 (0 new). Manifest 812 → 815. Refs: https://github.com/jackwener/rdt-cli * fix(reddit): tighten new read command failure contracts * fix(reddit): treat inaccessible subreddit info as empty |
||
|
|
eb59b7444d |
feat(ctrip): add hotel-search + flight browser-mode commands (#1481) (#1489)
* feat(ctrip): add hotel-search + flight browser-mode commands Closes #1481. Two new browser-mode commands on top of the existing public `search` / `hotel-suggest` pair: - `ctrip hotel-search <city> --checkin --checkout [--limit]` reads `window.__NEXT_DATA__.props.pageProps.initListData.hotelList` on `hotels.ctrip.com/hotels/list`. SSR-rendered first page ships ~13 entries; the server ignores `&pageSize=N` so limit caps at 30 with default 10. AuthRequiredError surfaces when Ctrip redirects to the captcha gate. - `ctrip flight <from> <to> --date [--limit]` searches one-way flights on `flights.ctrip.com/online/list/oneway-…`. The post-load XHR is not currently captured by the daemon network buffer (per the known daemon_capture_pipeline_bug_2026_05_07 in agent memory), so rows are pulled from `.flight-list > span > div` cards via a position-anchored innerText parser. A generic `buildScrollUntilJs(selector, target)` helper mirrors the PR #1487 xiaohongshu scroll-until pattern with the selector parameterised. Round-trip + airline filters are out of scope for v1. All argument validation (IATA / ISO date / city ID / limit range) fires upfront before any `page.goto`, per the PR #1387 boundary standard. No silent clamps, no sentinel rows: rows missing required fields are dropped, and end-state checks raise `ArgumentError` / `AuthRequiredError` / `EmptyResultError` as appropriate. The new `mapHotelRow` / `pickHotelMapCoords` / `buildFlightExtractJs` / `buildScrollUntilJs` helpers live in `clis/ctrip/utils.js` alongside the existing suggest helpers. Docs at `docs/adapters/browser/ctrip.md` now distinguish the public suggest commands from the browser-mode commands and document each command's columns + caveats. Verified: - 61/61 vitest tests in `clis/ctrip/ctrip.test.js` (including JSDOM exercises of `buildFlightExtractJs` and full `mapHotelRow` shape parity) - `check:typed-error-lint` 189/189 (0 new) - `check:silent-column-drop` 103/103 (0 new) - `build-manifest` clean — 812 entries total (was 810) * fix(ctrip): harden browser search failure contracts * fix(ctrip): tighten browser empty-vs-parser failures |
||
|
|
7df9b80dea |
fix(xiaohongshu+rednote): scroll until enough rows for --limit > 13 (#1471) (#1487)
* fix(xiaohongshu+rednote): scroll until enough rows are rendered instead of fixed 2x autoScroll
Both search adapters previously called `page.autoScroll({ times: 2 })` which
hard-capped extraction at ~13 notes (xiaohongshu lazy-loads ~5-7 notes per
scroll round) regardless of `--limit`. Reported in #1471: `--limit 40` still
only returned 13 results.
Replace with a dynamic `buildScrollUntilJs(targetCount, maxScrolls=15)`
helper that:
- counts visible `section.note-item` rows (excluding `.query-note-item`
related-search rows)
- breaks early when count >= target
- breaks early after 2 consecutive scrolls add no new rows (DOM plateaued,
feed exhausted)
- hard caps at 15 iterations to bound runtime
Exported from xiaohongshu and reused by rednote (same DOM shape) instead of
duplicating the IIFE.
Fixes #1471
* fix(xiaohongshu): tighten search scroll boundary
|
||
|
|
b476d2364f |
fix(doubao/ask): restore Assistant detection after 2026-05 DOM refactor (#1484)
* fix(doubao/ask): restore Assistant turn detection after 2026-05 DOM refactor Doubao reworked message-item wrappers and dropped all `receive-message` / `bg-g-receive-msg-bubble` markers from assistant turns. The legacy 6 `itemSelectors` (`item-kDun2N`, `union_message`, `message-block-container`, `data-message-id`, `bg-g-send-msg-bubble`, `bg-g-receive-msg-bubble`) match 0 elements on the new DOM, so `getTurnsScript` returned [] and `getDoubaoTurns` fell through to the whole-page transcript scraper. Assistant text came back as sidebar labels + history titles + adjacent conversation snippets concatenated with the real reply — silent SELECTOR failure (no thrown error). Two minimal changes in `clis/doubao/utils.js` `getTurnsScript`: 1. `itemSelectors`: prepend `[class*="inner-item-"]` and `[class*="top-item-"]` — the new 2026-05 wrappers. Outer wins via existing ancestor-keep dedup below, so we get one root per turn (not one per nested chunk). 2. `getRole`: add a third fallback branch — if the root matches `inner-item-*` / `top-item-*`, contains `.flow-markdown-body`, and has NO `bg-g-send-msg-bubble` marker (User detection still works), treat it as Assistant. `.flow-markdown-body` is already in `messageTextSelectors`, so text extraction kicks in unchanged. Test added asserting both new wrappers and the `.flow-markdown-body` assistant fallback are present in the generated script. Fixes #1478 * test(doubao): cover refactored assistant turns |
||
|
|
6d84009ee8 |
fix(youtube): request srv3 format for caption URLs (#1420) (#1422)
* fix(youtube): request srv3 format for caption URLs (#1420) YouTube may return empty responses when caption URLs lack an explicit format parameter. This adds fmt=srv3 (standard YouTube XML caption format) to the caption URL when no fmt parameter is already present, with a fallback to the original URL if srv3 also returns empty. Also adds HTTP status checking before reading the response body, preventing silent failures on non-200 responses. Fixes #1420 * fix(youtube): preserve caption fetch failures --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
150551be8c |
feat(reddit): add reply command for replying to comments (#1428)
* feat(reddit): add reply command for replying to comments Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(reddit/reply): replace silent-sentinel rows with typed errors reply.js originally mirror-copied comment.js's failure pattern: returning [{ status: 'failed', message: 'HTTP 403' }] on auth/HTTP/Reddit errors and relying on the caller to inspect the row instead of throwing. That's the 'silent-sentinel' anti-pattern from typed-errors.md — failures should surface as typed errors so an agent can actually branch on them. Round 21 lesson (f) — "grandfathered-not-exempt + helper-refactor boundary is new" — applies: comment.js / upvote.js / save.js can stay grandfathered, but a brand-new file does not inherit that exemption. Changes: - Throw AuthRequiredError when /api/me.json or /api/comment returns 401/403, or when /api/me.json returns 200 but data.name is missing (stale anon session — empty modhash alone isn't a strong enough signal). - Throw CommandExecutionError for non-2xx HTTP and for non-empty data.json.errors (e.g. RATELIMIT, NO_TEXT, TOO_OLD). - Drop the over-defensive `if (!page) throw ...` — registry guarantees a page object when browser:true. - Intermediate result object uses `kind` discriminator + `detail` / `httpStatus` / `where` keys that don't overlap with columns ['status','message'], so the silent-column-drop audit stays quiet (per PR #1329 sediment). Verified: - npx tsc --noEmit clean - node scripts/check-typed-error-lint.mjs → 189/189, 0 new - node scripts/check-silent-column-drop.mjs → 103/103, 0 new - npx vitest run clis/reddit src/convention-audit → 11/11 pass - node ./dist/src/main.js validate → 0 errors Success path is unchanged: still returns [{ status: 'success', message: 'Reply posted on t1_<id>' }]. * fix(reddit): harden reply command contract * fix(reddit): reject suffixed reply urls --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
64ac362a40 |
feat(rednote): add rednote.com adapter mirroring xiaohongshu read commands (#1136) (#1475)
* feat(rednote): add rednote.com adapter mirroring xiaohongshu read commands (#1136) Implements rednote.com support as discussed in issue #1136. The mainland xiaohongshu adapter stays in place; international users redirected to www.rednote.com now have a CLI without a copy-pasted adapter. Issue #1136 documents that xiaohongshu and rednote share DOM selectors, URL paths, API paths, response schema, cookies, and the xsec_token auth mechanism. The only material differences: Layer xiaohongshu rednote Web host www.xiaohongshu.com www.rednote.com API host edith.xiaohongshu.com webapi.rednote.com Security host fe-static.xhscdn.com as.rednote.com Cookie root .xiaohongshu.com .rednote.com Search gate Inline text Full-screen modal + text ## Architecture (minimal) `clis/xiaohongshu/*` keep all selector / regex / extraction logic. Each command file is touched minimally to export the IIFE or pipeline so the sibling adapter can reuse it: search.js + export const buildSearchExtractJs(webHost) + export const command = cli({...}) note.js + export const NOTE_EXTRACT_JS + export const command = cli({...}) comments.js + export function buildCommentsExtractJs(withReplies) + export parseCommentLimit + export const command = cli({...}) download.js + export function buildDownloadExtractJs(noteId) (CDN allowlist now includes rednote alongside xhscdn) + export const command = cli({...}) user.js + export const USER_SNAPSHOT_JS + export const command = cli({...}) feed.js + export function buildFeedPipeline(webHost) + export const command = cli({...}) notifications.js + export function buildNotificationsPipeline(webHost) + export const command = cli({...}) note-helpers.js buildNoteUrl now accepts `cookieRoot` + `signedUrlHint` options (defaults preserved so xhs callers and tests are unchanged) user-helpers.js buildXhsNoteUrl / extractXhsUserNotes accept an optional `webHost` argument (default xhs) The `export const command = cli({...})` pattern matches twitter/lists.js and clis/discord-app/*; without it the build-manifest scanner attributes xhs's command to whichever rednote sibling triggered the transitive import first. ## clis/rednote/ — thin shims Each rednote command file imports the relevant builder / constant from its xiaohongshu sibling and calls `cli()` with the rednote host triple. No selectors, regexes, or extraction logic are duplicated. search.js imports buildSearchExtractJs + noteIdToDate declares its own WAIT_FOR_CONTENT_JS (modal + text login-gate variants — the one xhs behaviour that genuinely differs) note.js imports NOTE_EXTRACT_JS + buildNoteUrl + parseNoteId comments.js imports buildCommentsExtractJs + parseCommentLimit + buildNoteUrl + parseNoteId download.js imports buildDownloadExtractJs + buildNoteUrl + parseNoteId user.js imports USER_SNAPSHOT_JS + extractXhsUserNotes + normalizeXhsUserId ## Scope (initial) Ships the five commands verified live against the user's logged-in rednote.com session: search / note / comments / user / download. `feed` and `notifications` are intentionally left out. Both rely on intercepting the xiaohongshu Pinia store at the `homefeed` / `you` capture pattern; live verification on rednote returns `tap → dict (error)` for the feed step, so shipping them would surface a broken contract. The mainland xiaohongshu commands continue to work. Adding the rednote-side feed / notifications is straightforward follow-up work once someone with rednote access maps the network surface. Creator-center commands (publish, creator-*) have no rednote counterpart and stay xiaohongshu-only, per the reporter's note in #1136. ## Verification - clis/xiaohongshu/ + clis/rednote/: 103/103 tests green - npx tsc --noEmit: clean - npm run build: 807 manifest entries (xhs 13 + rednote 5 + everything else preserved) - silent-column-drop / typed-error-lint: 103 / 189 baseline entries, no new violations - Live verify against the user's rednote.com session: rednote search "travel" --limit 1 → real note row rednote note <signed-url> → 7 field/value rows rednote comments <signed-url> --limit 3 → 3 top-level rows rednote user 5b21f6564eacab3b38f05c39 --limit 2 → 2 profile notes Spaced 15–30s between runs per the xhs/rednote rate-limit guidance; no write commands invoked. Regression check: xiaohongshu/feed on the existing mainland session still returns the standard 6-field rows after the refactor. Closes #1136 * fix(rednote): tighten adapter failure boundaries --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
b262d8ffd5 |
feat(chatgpt): support local image uploads (#1476)
* feat(chatgpt): support local image uploads * chore: refresh cli manifest * fix(chatgpt): harden image upload flow * fix(chatgpt): validate image uploads before navigation * fix(chatgpt): keep send fallback click in sync --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
467fdd0b62 | refactor(adapter): rename site browser reuse to persistent sessions (#1462) | ||
|
|
9c06e84c89 | refactor(browser): replace workspaces with sessions (#1461) | ||
|
|
1d2e606498 |
perf(chatgpt): replace fixed-sleep waits with selector-based readiness (D3) (#1456)
Continues the wait→event sweep started in #1449 (deepseek) / #1452 (claude). Same 3-bucket classification across the chatgpt adapter: CONVERT (5) - utils.js ensureOnChatGPT/startNewChat: 2s settle → wait({selector: composer, 8s}) - utils.js getConversationList: openSidebar 1.5s + fallback goto 2.5s → selector - detail.js: post-/c/<id> goto 2s → wait({selector: message bubble, 10s}) DELETE (6) - ask/send/read.js: standalone 2s settle after ensureOnChatGPT/startNewChat (those helpers now wait for composer internally — settle is redundant) - utils.js sendChatGPTMessage: 0.5s post-closeBtn + 1.5s pre-composer-focus - utils.js getConversationList: 2s settle after ensureOnChatGPT (helper waits for composer; we re-check sidebar selector independently) KEEP (6) - utils.js sendChatGPTMessage: ProseMirror React debounce ticks - utils.js waitForChatGPTResponse: streaming response polling cadence Verification: - npx vitest run clis/chatgpt → 20/20 (4 files) - Full vitest → 3379 pass / 1 skip (2 errors are the unrelated daemon EADDRINUSE flake also seen on #1449/#1452/#1454) - tsc --noEmit clean / typed-error-lint 189 / silent-column-drop 103 unchanged - npm run build → 802 manifest entries Diff: +50 / -13 across 5 files (ask.js, detail.js, read.js, send.js, utils.js). No typed-error harmonization needed — chatgpt's existing helpers already gate through ensureChatGPTLogin (AuthRequiredError) and ensureChatGPTComposer (CommandExecutionError) correctly. |
||
|
|
6d87142821 |
perf(reddit): opt 13 browser adapters into shared site-tab lease (#1455)
* perf(reddit): opt 13 browser adapters into shared site-tab lease
Adds `browserSession: { reuse: 'site' }` to every reddit adapter that
already runs `browser: true` on `domain: 'reddit.com'`. Same metadata-only
follow-up to the twitter sweep merged in #1454 — the framework's
`shouldRunPreNav` short-circuit (src/execution.ts:190) skips the redundant
domain-root pre-nav when a sibling adapter already has the tab on
reddit.com, and idle-bound tabs are reused under the `site:reddit` bucket
until expiry.
Scope (13 files, all on `domain: 'reddit.com'` + `Strategy.COOKIE`):
- read (9): frontpage / popular / saved / search / subreddit / upvoted /
user / user-comments / user-posts
- write (4): comment / save / subscribe / upvote
Excluded:
- `hot.js` (no browser:true — public Reddit JSON API, no tab)
- `read.js` (Strategy.COOKIE but no browser:true — non-browser pipeline)
No logic changes; only metadata + manifest regeneration.
Verification:
- npm run check:typed-error-lint → 189/189 unchanged
- npm run check:silent-column-drop → 103/103 unchanged
- npm run test:adapter → 264/264 passed (2146 tests)
- npx vitest run --project unit → 72/72 passed (unrelated EADDRINUSE
flake on daemon.test.ts port 19825, also seen on #1454/#1452)
- tsc --noEmit clean
* fix(reddit): include read in site browser session reuse
|
||
|
|
674f0e1105 |
feat(openreview): add author command for ID-explicit publication lookup (#1365)
* feat(openreview): add author command for ID-explicit publication lookup
Closes the missing leaf in the openreview adapter. Among the public-strategy
academic adapters, dblp and arxiv both already ship an `author` command for
ID-explicit publication lookup; openreview only had `search` (full-text),
`paper` (detail by note id), `reviews` (thread by forum id) and `venue`
(listing by invitation / venue text). There was no way to ask "give me every
submission this author put on OpenReview, newest first."
`openreview author <profile>`:
- takes a canonical profile id (`~First_LastN`); validated by
`requireProfileId` so a dblp PID or a bare name fails before any
network call,
- hits `/notes?content.authorids=~<id>&limit=<n>&sort=cdate:desc`,
- returns rank-ordered rows with the same shape as `openreview search`
(id / title / authors / venue / pdate / url),
- throws `EmptyResultError` when the profile has no public submissions
instead of returning an empty list,
- inherits the typed-error envelope from `openreviewFetch` so network
failure, non-200, malformed JSON, and in-band error envelopes all
surface as `CommandExecutionError`.
Tests: 6 new `it` blocks plus 1 updated registration test in
`clis/openreview/openreview.test.js`.
- `requireProfileId` (1 block, 9 assertions): accepts canonical
`~First_LastN`, `~Bo_Liu17`, and a multi-segment middle-name id;
rejects empty, whitespace, missing tilde, missing trailing number,
embedded space, and a dblp-style PID.
- 5 author runtime cases covering pre-network ArgumentError, empty
result, non-200, fetch network error, and the happy path with a
request-shape assertion (`content.authorids` filter + `cdate:desc`
sort).
- Registration test extended to expect five commands and lock the new
`columns` contract.
Manifest auto-regenerated to register the new command.
Live-verified end to end against `~Yoshua_Bengio1`: the most recent ICLR
2026 workshop submissions return with the expected fields. A malformed
profile is rejected before any HTTP call. A nonexistent profile yields
`EMPTY_RESULT`.
* fix(openreview): accept real profile id slugs
---------
Co-authored-by: jackwener <jakevingoo@gmail.com>
|
||
|
|
2034e90337 |
fix(chatgpt): use locale-stable send button selector (#1354)
Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
a77d9c930a |
perf(claude): replace fixed-sleep waits with selector-based readiness (#1452)
Convert 9 of 18 page.wait(N) calls in clis/claude/ from fixed-duration
sleeps to event-based readiness checks (page.wait({ selector, timeout }),
backed by MutationObserver). Mirrors the deepseek D1 template (PR #1449).
* Page-ready waits (5 converted): utils.js:29 (ensureOnClaude composer),
utils.js:114 (getConversationList recents links), new.js:19 (composer),
detail.js:24 (.font-claude-response message bubble), send.js:26 (composer).
Each resolves as soon as the selector matches, swallowing the timeout so
downstream typed-error helpers (ensureClaudeLogin / ensureClaudeComposer
/ EmptyResultError) still surface the right error when the selector
never mounts (login redirect, empty conversation, etc).
* Dropdown waits (2 converted): utils.js:150 (selectModel post-trigger),
utils.js:178 (setAdaptiveThinking post-trigger). Wait for menuitemradio
/ menuitem to mount instead of a fixed 0.6 s sleep.
* Resume conversation wait (1 converted): ask.js:51 — wait for the resumed
message bubble (MESSAGE_SELECTOR) instead of a fixed 2 s sleep.
* Settle/redundant waits removed (6): ask.js:55 standalone settle (next
ensureClaudeComposer queries composer presence directly via getPageState);
ask.js:83 / ask.js:90 post-toggle settles (next CDP eval flushes React
state between roundtrips); ask.js:103 pre-waitForResponse settle (the
polling loop's first 3 s tick already covers this); read.js:20 post-
ensureOnClaude sleep (ensureOnClaude now waits for the composer selector
itself); send.js:29 post-ensureOnClaude sleep (same).
Three remaining page.wait(N) calls are kept: utils.js:231 post-input
1.2 s React debounce inside sendMessage (the ProseMirror editor needs a
debounce window before the send button enables; reducing this risks
silent send-button-disabled drops), and the 3 s / 1 s polling ticks in
waitForResponse / waitForFilePreview (already polling patterns, out of
scope for D-track wait→event sweep).
Targeted tests: clis/claude + src/browser 389/389 pass; tsc clean;
build clean; typed-error 189/189 baseline (no new); silent-column-drop
103/103 baseline (no new).
D2 in the LLM-adapter wait→event sweep started by deepseek (D1, #1449).
|
||
|
|
cb64192f06 |
perf(deepseek): replace fixed-sleep waits with selector-based readiness (#1449)
Convert 10 of 18 `page.wait(N)` calls in clis/deepseek/ from fixed-duration
sleeps to event-based readiness checks (`page.wait({ selector, timeout })`,
backed by MutationObserver):
* Page-ready waits (5): utils.js:46, ask.js:38/52, detail.js:28, new.js:19
now wait for the composer textarea (TEXTAREA_SELECTOR) or message bubble
(MESSAGE_SELECTOR) to mount before continuing. Resolves as soon as the
selector matches instead of always sleeping the full duration.
* Settle/redundant waits removed (5): ask.js:56 standalone settle (already
covered by upstream selector waits); ask.js:79/105 post-toggle settles
(next CDP eval gives React time to flush aria-checked updates); ask.js:118
pre-waitForResponse settle (the polling loop's first 3 s tick already
covers this); read.js:19 post-ensureOnDeepSeek sleep (ensureOnDeepSeek
now waits for the textarea selector itself).
* `new.js` now throws CommandExecutionError when the composer fails to
mount within 8 s instead of silently returning "New chat started" on a
half-loaded or logged-out page.
Eight remaining `page.wait(N)` calls are kept: in-loop polling ticks in
waitForResponse / pickResumeUrl / getConversationList / waitForFilePreview
/ send-button-enable polling (these are already polling patterns and
out of scope for D1), and the native-input flush + textarea-mount poll in
send.js.
Targeted tests: clis/deepseek 49/49, src/browser 355/355 pass; build,
typecheck, typed-error and silent-column-drop audits clean.
Proof template for the LLM-adapter wait-cleanup follow-ups.
|
||
|
|
833c1c872f |
perf(twitter): enable browserSession reuse:site on 17 read-only adapters (PR B) (#1454)
Read-only Twitter/X adapters now declare `browserSession: { reuse: 'site' }`,
matching the LLM-site adapters (claude/gemini/yuanbao/etc.) and unblocking
the perf wins WAWQAQ called out for the 35s→9s/3.4s thread.js progression
(#OpenCLI:3889b5cf):
- Tab lease shared across calls under `site:twitter` until idle expiry, so
the second-and-later command pays no cold-start tab cost.
- Framework's domain-root pre-nav (`https://x.com`) is skipped on subsequent
calls when the reused tab is already on x.com (`shouldRunPreNav` →
`isDomainRootPreNav` + `urlMatchesDomain` short-circuit at
`src/execution.ts:190`).
Files (17 read-only adapters):
- Strategy.COOKIE × 13: article, bookmark-folder, bookmark-folders,
bookmarks, download, following, likes, list-tweets, lists, profile,
thread, timeline, trending, tweets
- Strategy.UI × 1: followers
- Strategy.INTERCEPT × 2: notifications, search
Insertion point in each file: after `browser: true,` (or after `strategy:`
in download.js which omits the explicit `browser:` field), matching the
convention used by yuanbao/read.js, claude/read.js, etc.
Manifest regenerated (cli-manifest.json: +85/-17 — 17 entries gain the
`browserSession: { reuse: "site" }` block).
Verification:
- npx tsc --noEmit clean
- npx vitest run clis/twitter → 218/218 pass (25 files)
- npx vitest run src/convention-audit.test.ts → 8/8 pass
- typed-error-lint baseline 189/189 (no new violations)
- silent-column-drop baseline 103/103 (no new violations)
Scope notes (intentionally NOT in this PR):
- Write adapters (post/reply/quote/like/retweet/bookmark/follow/list-add/
list-remove/delete/hide-reply/block/accept/follow) are kept as one-shot
by default — `reuse: 'site'` for write paths is a separate decision
about action idempotency under tab reuse.
- The thread.js / timeline.js comments still say "Cookie context
auto-established by framework pre-nav"; the deeper truth (CDP
`getCookies({url})` is origin-independent) was a framing nit on PR C
(#1451) — left as a doc-only follow-up to keep this PR's diff focused
on the perf gain.
Refs: #OpenCLI:3889b5cf (WAWQAQ msg=fa209a2c, msg=35c90460, msg=838128ef
"你们继续做啊… 后面还有那么多其他的东西呢")
|
||
|
|
ff7d741a4c |
perf(twitter): drop redundant goto+wait — framework auto pre-navs (PR C) (#1451)
* perf(twitter): drop redundant goto+wait — framework auto pre-navs (PR C)
Twelve twitter read adapters did `await page.goto('https://x.com'); await
page.wait(2~3)` purely to establish cookie context for the subsequent
`document.cookie` read. After PR #1450 hoisted those reads to
`page.getCookies({url})` (which queries the CDP cookie store directly,
no navigation needed), the explicit goto+wait became dead.
The framework already pre-navigates to `https://${domain}` for any
adapter declaring `Strategy.COOKIE + domain` (`src/registry.ts:191`),
so the cookie store is populated before `func` runs. The 2-3s
`page.wait` was the slowest part of the redundant call.
Files (all read-only, all ct0/cookie-only):
- bookmark-folder / bookmark-folders / bookmarks
- following / likes / list-add / list-remove / list-tweets / lists
- thread / timeline / tweets
Out of scope (kept as-is): goto calls that navigate to a *specific*
URL needed for content/SPA shell — `trending` (`/explore/tabs/trending`),
`notifications` (`/home`), `article` (`/i/article/{id}`), `profile`
(`/${username}`), and `list-add` line 133 (`/${username}` for UI ops).
Verification:
- npx tsc --noEmit ✓
- npx vitest run clis/twitter → 216/216 ✓
- typed-error-lint 189/189, 0 new ✓
- silent-column-drop 103/103, 0 new ✓
* fix(twitter): keep list UI root navigation
|
||
|
|
60dbbd4baa |
perf(adapters): hoist cookie reads to page.getCookies (Tier 1, 25 files) (#1450)
* perf: replace document.cookie reads with page.getCookies({domain}) (Tier 1 cookie API sweep)
Prior pattern in 25 adapter files round-tripped through `page.evaluate(\`document.cookie.split…\`)` to extract a single cookie value (CSRF token, session ID, etc.). CDP's `page.getCookies({domain})` reads the cookie store directly with zero JS-execution overhead.
Files touched (sites: twitter / linkedin / maimai / youtube):
- twitter (15): thread, timeline, list-add, bookmark-folders, following, list-tweets, bookmarks, list-remove, tweets, bookmark-folder, likes, lists, trending — direct 4-line replacement (cookie was outside `page.evaluate`); article, profile — hoisted ct0 read OUT of `page.evaluate` and threw `AuthRequiredError` upfront so unreachable in-evaluate auth branches got cleaned up too.
- linkedin/search.js — JSESSIONID was read inside the per-batch fetch loop's `page.evaluate`; hoisted once before the loop and pass `csrf` value into the template via `JSON.stringify`.
- maimai/search-talents.js — csrftoken cookie hoisted via getCookies; meta-tag fallback preserved inside `page.evaluate` (reached only when no cookie). Also converted the `page.evaluate(async (body) => …, body)` Playwright-style call to OpenCLI's template-string form so the helper actually runs.
- youtube — `SAPISID_HASH_FN` (used by like / unlike / subscribe / unsubscribe) reworked: sapisid is now passed in as a parameter; new `readYoutubeSapisid(page)` helper reads it via CDP. The HMAC-SHA1 compute still happens browser-side (Web Crypto), only the cookie read is hoisted.
Tests updated where mocks specifically referenced `document.cookie` (twitter following / bookmark-folder / bookmark-folders) to mock `getCookies` instead.
Verification:
- `npx tsc --noEmit` clean
- `npx vitest run clis/twitter clis/linkedin clis/youtube` → 264/264 pass
- typed-error-lint 189/189 (no new violations)
- silent-column-drop 103/103 (no new violations)
Scope notes (not in this PR):
- `goto + wait` redundancy and `browserSession: { reuse: 'site' }` rollout are scoped to follow-up PRs B and C per the #OpenCLI:3889b5cf thread plan.
- `document.cookie.match(...)` patterns (instagram 8 / xiaoe / qwen / hupu / tiktok / 1point3acres — ~13 files) are outside the original \`document.cookie.split\` audit scope and will follow as a Tier 1 expansion sweep.
* fix(adapters): read auth cookies by url scope
|
||
|
|
85ea18c93b |
feat(dianping): resolve unknown cities live from www.dianping.com (#1429)
* feat(dianping): resolve unknown cities live from www.dianping.com
The static CITY_ID map in clis/dianping/utils.js only covers ~20 cities,
so passing --city 汕头 (or any other Chinese name / pinyin slug not on
that list) fails with ArgumentError. Adding the missing cityIds by hand
doesn't scale to dianping's full city list and silently goes stale when
the site renumbers cities.
This change adds an async resolver that falls back to dianping.com when
the static map misses:
- Numeric input → pass through unchanged.
- Static map hit → fast path, no network (utils.CITY_ID untouched).
- Pinyin slug (e.g. "shantou") → goto /<slug>, parse cityId out of
any /search/keyword/{id}/ link rendered on the per-city landing page.
- Chinese name (e.g. "汕头") → goto /citylist, walk anchors to build a
Chinese-name → pinyin map, then resolve the slug as above.
Resolved (input → cityId) pairs are memoized per-process so repeat
searches skip both navigations.
Implemented as a new module (clis/dianping/cityResolver.js) so utils.js
stays minimal and the existing synchronous resolveCityId / CITY_ID API
keeps working for direct callers and tests.
Tested:
- Unit tests cover null/numeric/static fast paths, pinyin fallback +
cache, Chinese-name fallback via /citylist + cache for both forms,
rejection of garbage input, rejection of Chinese names not on
/citylist, and CommandExecutionError when the per-city page lacks
a /search/keyword/{id}/ link.
- JSDOM tests cover the pure DOM extractors (buildCitylistMap and
extractCityIdFromPage) against curated HTML fixtures.
- npm test: 3196 passed, 1 skipped (no new failures).
- npx tsc --noEmit: clean.
- opencli validate: 0 errors.
* fix(dianping): require city resolver links to be authoritative
---------
Co-authored-by: jackwener <jakevingoo@gmail.com>
|
||
|
|
962842cd59 |
fix(douyin): handle empty response body in browserFetch (#1408)
* fix(douyin): handle empty response body in browserFetch (#1405) browserFetch calls res.json() directly, which throws SyntaxError when the API returns an empty body (content-length: 0). This happens when the Douyin hashtag search endpoint returns HTTP 200 with no content. Fix: read response as text first, return null for empty bodies, then throw a descriptive CommandExecutionError at the caller level. Fixes #1405 * fix(douyin): wrap browser fetch parse failures --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
4475d4efe3 |
feat(twitter): P1+P2+P3+P4+P5 — search filters, bookmark folders, engagement scoring, sibling dedupe + help docs (#1406)
Round 21 follow-up to #1400 (P0 write-action symmetry, merged `644d4517`). 5 features + help docs unified into one PR per WAWQAQ "全部合成一个 PR" directive. ## Scope - **P1** (`cf10c098`): `twitter search` `--from / --has / --exclude / --product` filters, mapping to X `from:` / `filter:` / `-filter:` / `f=` operators; legacy `--filter top|live` preserved (--product win on conflict) - **P2** (`a484a69a`): new `twitter bookmark-folders` + `bookmark-folder <id>`; X Premium GraphQL `bookmarkFoldersSlice` + `BookmarkFolderTimeline`; queryId 三层 fallback (placeholder.json → client-web bundle → pinned constants) - **P3** (`f209f914`): `--top-by-engagement N` to 7 tweet-shaped read commands (search/timeline/likes/bookmarks/list-tweets/tweets/thread); single helper in `utils.js`; formula `likes×1 + retweets×3 + replies×2 + bookmarks×5 + log10(views+1)×0.5`; **N=0 reference equality no-op** → existing 157 twitter tests 0 churn - **P4** (`89283fa0`): `TWITTER_BEARER_TOKEN` + composer image helpers extracted to `utils.js` (12 GraphQL adapter dedup); reply hardening; quote adds `--image` - **P5** (`a3d10a48`): sibling article-scope helper extracted to `shared.js` (9 write commands reuse, dedup with #1400 P0 invariant) - **docs** (`2a358d80`): help-doc precision (positional-omitted defaults + download/bookmarks/notifications/timeline/lists description thicken; concurrent #1401/#1403 wording preserved) 47 files / +2594/-470. Tests **96 → 216 (+120)**, manifest 798 → 801 (+3), typed-error-lint 190 → 189 (resolved 1 grandfathered sentinel). ## Iteration history (3 review fix commits on top of 6 author commits) - `7f93779b` — codex-mini1 lead fix1: 3 blocker bundle (P5 host invariant + P2 safe-id + sentinel removal + P1 fallback fail-fast) - `2a29ecc6` — codex-mini1 lead fix2: P3 help formula consistency (doc/help text matches actual `log10(views+1)×0.5`) - `df4dcd76` — codex-mini1 lead fix3 (F-P-1 aux catch): P2 `bookmark-folder --limit` upfront validation (`Number(kwargs.limit ?? 20)` + reject non-positive/non-integer + regression `0/negative/fractional/NaN` + `page.goto` zero-call assert) ## 4 progressive blockers caught (codex-mini1 lead 3 rounds + F-P-1 aux 1 round) 1. **P5 host invariant gap** (lead): article-scope helper preserved exact `/status/<id>` path but ignored link host → off-domain `https://evil.com/alice/status/<target>` would satisfy `__twHasLinkToTarget`. Fixed: `https` + X/Twitter host or subdomain + exact `/status/<id>` or `/i/status/<id>` path; query/hash allowed; off-domain/host-suffix/non-https/path-suffix/substring-id rejected; JSDOM positive + 5 negative anchors. 2. **P2 listing→detail round-trip + sentinel** (lead): `bookmark-folders` accepted opaque IDs but `bookmark-folder <id>` only accepted numeric → round-trip broken; new `author: 'unknown'` sentinel created fabricated author URL. Fixed: `[A-Za-z0-9_-]+` opaque safe-id (rejects `/`, `?`, `%`, spaces) + `resolveTwitterQueryId()` sanitization for queryId resolution; sentinel removed → empty author + canonical `/i/status/<id>` URL. 3. **P1 fallback silent tab miss** (lead): pushState fail → fallback typing into search box, `clickProductTabIfNeeded()` silent return on tab not found → user `--product photos` silently degraded to Top results. Fixed: throw `CommandExecutionError` when requested `--product` tab cannot be selected + invalid `--from` / `--limit` upfront pre-nav reject + double-direction tests. 4. **P2 limit silent normalize** (aux): `const limit = kwargs.limit || 20` → `--limit 0` silent → 20; negative/non-integer pre-IO unchecked. Fixed: `Number(kwargs.limit ?? 20)` + require positive integer before `page.goto` + regression covers `0/negative/fractional/NaN` + `page.goto` zero-call. ## Cultural sediment (Round 21 audit checklist 7 rules / 6 dimensions) This PR **immediately validated 4 of 7 rules** in review pipeline: - (b) silent-clamp class — P1 fallback silent tab miss (silent semantic-downgrade) + P2 `|| 20` silent normalize - (e) ID exact-not-substring — P5 host invariant (was only path-exact, not host-exact) - (f) grandfathered-not-exempt — P5 helper-refactor boundary lost host invariant + P2 new adapter inherited grandfathered `'unknown'` sentinel - (g) fallback-must-have-success-criterion — P1 fallback path missing post-condition assertion 7 rules / 6 dimensions: - (a) cross-grep sibling URL pattern — structural - (b) silent-clamp class — failure mode (input) - (c) broad querySelector → article-scoping — scope - (d) missing-validation early reject — boundary - (e) ID exact-not-substring — identity - (f) grandfathered-not-exempt (corollary: applies to new file + new helper-refactor boundary; not original-file line-edit) — time-axis - (g) fallback-must-have-success-criterion (sub-rule g': fallback unit test must include post-condition assertion, not just "doesn't throw") — failure mode (output) **Cross-PR validation 4-chain on meta-anchor "Structural exactness for identity matching"**: - #1391 URL layer (`isFacebookAuthRedirectPath`: top-level anchor + `\.php` + `(/|$)` segment edge) - #1392 URL parser layer (`parseGrokSessionId`: bare UUID exact / URL host-exact-or-subdomain + path-exact) - #1400 DOM layer (article-scoping: status-id `/\/status\/${id}(?:\/|$)/` regex / segment-array exact) - #1406 P5 helper-refactor boundary (full URL invariant in shared helper: host+path re-anchored after extraction) - Common invariant: boundary-lock structural shape; **fuzzy match is silent-failure 温床**; lesson lifecycle = surface-shift not add-and-forget. **Audit framework self-discipline**: each rule must have grep-able detection signal, otherwise rule degenerates to mantra. Framework is "7 rules + sub-instance pattern in new surface", not frozen 7 rules. **Round 17 race-mitigation 第 9 连续 race-free execution**: standard alternation cadence (#1400 A 组 → #1406 B 组), lead final + aux final + `@pr-monitor squash?` trigger, pr-monitor proactive ack + serial squash, lead silent on closeout. ## Validation gates (final head `df4dcd76`) Local: Twitter adapter tests `25 files / 216 tests`, focused P1/P2/P3/P5 tests `99/99`, `node --check` touched runtime, `npx tsc --noEmit`, `npm run build`, manifest 801 entries, typed-error-lint `189/189`, silent-column-drop `103/103`, doc-coverage `140/140`, docs:build clean, listing-id advisory `13` unchanged (wikipedia/trending residual non-Twitter), `git diff --check` clean. GitHub: build×3 (ubuntu/macos/windows) SUCCESS, unit-test shards SUCCESS, bun-test SUCCESS, adapter-test SUCCESS, audit SUCCESS, doc-coverage SUCCESS, docs-build SUCCESS, smoke-test skipped, PR `CLEAN/MERGEABLE`. Reviewers: - Lead: @codex-mini1 (3 fix rounds, all caught proactively + amend P3 help consistency) - Aux: @First-principles-1 (better-solution triangulation on P2 queryId 三层 fallback + P5 invariant + P3 N=0 reference no-op + caught P2 limit silent normalize) - Author: @opencli-user (5-feature scope + 7-rule sediment co-author + corollary contributor) |
||
|
|
407b559a83 |
feat(help): hard-gate empty positional help text + fix 18 offenders (#1403)
Why
- `opencli twitter followers --help` rendered:
Arguments:
user
with a blank trailing column. Both humans and agents could not
recover the parameter's purpose without reading source. WAWQAQ
surfaced this directly: "没有说明当后面的 followers [user] [options]
如果都没填的时候,获取的是什么?"
- This is metadata completeness, not stylistic taste. Failing closed
is the only way to keep the help surface trustworthy as adapters
land.
What
- src/build-manifest.ts: add `findManifestMetadataIssues()` that flags
any positional with empty / whitespace-only / missing `help`. Wired
into `main()` after the import-failures gate; build aborts non-zero
with a per-arg report (`site/cmd positional "name" (sourceFile)`).
- src/build-manifest.test.ts: cover the gate (positives + negatives,
scoped strictly to positionals — named flags are intentionally
out-of-scope).
- 18 adapter offenders (16 required + 2 optional) get explicit help
text:
twitter: followers/following/list-add/list-remove/list-tweets/
search/thread
reddit: search/subreddit/user/user-comments/user-posts
douyin: stats/update
bilibili: subtitle
jike: search
Optional positionals (`twitter followers/following [user]`) now
document the omit semantics — fetches the currently logged-in
account.
- CHANGELOG: document the build gate and the offender list.
Out of scope (planned follow-ups)
- Semantic-quality advisory: optional positional help should also
contain `default / omit / current / logged-in / required unless …`
keywords. That belongs to the planned Arg metadata v2 work
(`when_omitted / when_present / value_format` 3-field schema).
- Named-flag `help` quality. Named flags carry the flag name itself
in help, so a missing `help` is not as opaque; if we want to gate
those too, do it as a separate, intentional decision.
Validation
- `npm run build` → 799 entries, clean.
- `npm run typecheck` → clean.
- `npx vitest run --project unit --project adapter` → 257 + 4 files,
all green (build-manifest 13 tests, manifest gate added).
- Smoke: temporarily reverted `followers.js` help to empty → build
aborts with the exact `twitter/followers positional "user" (...)`
line; restored, build is clean again.
- `npm run check:silent-column-drop` and `check:typed-error-lint`
baselines unchanged.
|
||
|
|
644d45177b |
feat(twitter): add unlike + retweet + unretweet + quote (write-action symmetry P0) (#1400)
Round 21 P0 — Twitter write-action symmetry (4 of 4: unlike, retweet, unretweet, quote). ## Scope Closes write-action gap with existing siblings (`like`, `bookmark`, `unbookmark`, `delete`): - `unlike` (UI strategy, navigateBefore:true) - `retweet` (UI strategy) - `unretweet` (UI strategy) - `quote` (UI strategy, `/compose/post?url=` route — same family as `reply.js` `/compose/post?in_reply_to=`) +745/-0 in initial commit, plus 3 progressive review fixes. Final: 4 adapters + 4 tests; modified `shared.js`, `shared.test.js`, manifest, docs. ## Iteration history (4 heads, 102/102 tests on final) - `07836783` — initial 4 adapters + 4 tests, 96/96 - `55a89776` — fix #1: shared `parseTweetUrl()` URL invariant + quote post-submit verify (102/102) - `dc9eab66` — fix #2: article-scoping for unlike/retweet/unretweet (delete.js sibling pattern) - `8809d2c1` — fix #3: exact status-id matching (`match?.[1] === tweetId`) + quote-card exact id guard ## 4 progressive blockers caught (codex-mini0 lead + F-P-0 aux) 1. **URL validation (silent-clamp class)**: original passed any host containing `/status/<id>`. Fixed: `parseTweetUrl()` requires `https` + Twitter/X exact host + exact `/<user|i>/status/<id>` path; host-suffix, embedded URL, path-suffix all `ArgumentError` pre-nav. 2. **Quote silent-success illusion**: original click-implies-success without composer/toast verify. Fixed: pre-submit quoted-card exact id render assertion + post-submit success toast OR composer-clear assertion, otherwise return failed row. 3. **Broad querySelector scoping (delete.js sibling pattern)**: original state probe + click + post-click verify on conversation pages picked first matching button. Fixed: scope to `article` containing requested exact status id (sibling `clis/twitter/delete.js:22-23` pattern). 4. **Substring vs exact status-id matching**: `/status/123` substring-matched `/status/1234`. Fixed: regex `/\/status\/${id}(?:\/|$)/` segment-edge anchor + `match?.[1] === tweetId` exact compare. ## Cultural sediment (Round 21) **Audit checklist 5 rules (pre-write upstream selection net)**: 1. cross-grep sibling URL-construction patterns before adopting 2. silent-clamp class detection (any normalize-then-trust path) 3. broad querySelector → article-scoping requirement 4. missing-validation early reject before navigation/IO 5. ID-based DOM/URL matching exact-not-substring **Augment framing**: Round 21 audit-first 是 Round 18 字面量 self-check 的 **upstream pre-write 阶段**, 两者作用阶段不同, 共存比替换稳。 **Meta-anchor "Structural exactness for identity matching"** unifying: - URL layer (#1391 isFacebookAuthRedirectPath: `\.php` + `(/|$)` segment edge) - URL parser layer (#1392 parseGrokSessionId: bare UUID exact / URL host-exact-or-subdomain + path-exact) - DOM layer (#1400 article-scoping: status-id `/\/status\/${id}(?:\/|$)/` regex or pathname segment-array exact compare) Common invariant: boundary-lock structural shape, 不 trust substring 模糊 — fuzzy match 是 silent failure 温床。 ## Validation gates (final head `8809d2c1`) Local: Twitter tests 102/102, `node --check` touched files, `npx tsc --noEmit`, `npm run build`, typed-error-lint 189/189, silent-column-drop 103/103, doc-coverage 140/140, docs:build clean, listing-id advisory unchanged 13, `git diff --check` clean, merge-tree clean. GitHub: build×3 (ubuntu/macos/windows) SUCCESS, unit-test×2 shards SUCCESS, bun-test SUCCESS, adapter-test SUCCESS, audit SUCCESS, doc-coverage SUCCESS, docs-build SUCCESS, smoke-test skipped, PR `CLEAN/MERGEABLE`. ## Strategy/UI boundary (better-solution verdict) UI write path acceptable for P0 symmetry (matches existing Twitter write siblings). GraphQL write migration + structured `idempotent:true` flag are cross-sibling upgrades, P5 candidate, not P0 blockers. Round 17 race-mitigation 第 8 连续 race-free execution (this round absorbed author scope-uncertainty hold-then-retract event without producing actual race). Reviewers: - Lead: @codex-mini0 (4-round iteration, all blockers caught) - Aux: @First-principles-0 (better-solution triangulation, scope-discipline verdict, regression invariants) - Author: @opencli-user |
||
|
|
bf914f20f1 |
fix(grok): replace sentinel rows + silent-clamp with typed errors, deliver image cmd (#1397)
fix(grok): replace sentinel rows and deliver image command |
||
|
|
6f45db1be9 |
fix(manifest): rescue 11 desktop adapter commands from factory pattern (#1396)
fix(manifest): rescue desktop factory commands |
||
|
|
abfd0e2180 |
fix(youtube): use watch page HTML for transcript captions (#1378)
Fixes #1376 — YouTube transcript command failed with `No captions available for this video` for all videos. ## Root cause Transcript adapter used InnerTube `/youtubei/v1/player` API with Android client context (`clientName: 'ANDROID'`, version `20.10.38`) to retrieve caption track URLs. YouTube has restricted/deprecated this approach; the Android client no longer reliably returns captions data. ## Fix Replace Step 1 (caption track retrieval) with watch page HTML bootstrap parsing — fetch `/watch?v=...` with cookies and extract `ytInitialPlayerResponse.captions.playerCaptionsTracklistRenderer`. This is the same approach used by sibling `clis/youtube/video.js`, so it's an alignment to existing site-local stable pattern, not a new invention. ## 2 head iteration - `cf77f5e8` initial fix (Step 1 caption retrieval switch + 18/18 unit tests) - `bb30788c` lead test hardening — source-contract regression test in `transcript.test.js`: - **positive lock**: must fetch `/watch?v=...`, parse `ytInitialPlayerResponse`, read `playerCaptionsTracklistRenderer` - **negative lock**: must NOT use `/youtubei/v1/player` or `clientName: 'ANDROID'` (prevents regression) - stale Android-InnerTube file header comment also updated ## Better-solution evaluation - Official YouTube Data API captions surface (`developers.google.com/youtube/v3/docs/captions/download`) is owner-authorized API, NOT a public transcript replacement - yt-dlp also relies on watch-page bootstrap path - Existing `youtube/video.js` already uses the same `ytInitialPlayerResponse` extraction → this PR aligns transcript with stable site-local pattern instead of inventing a new path ## Typed failure / no-silent-empty boundaries - watch HTML HTTP failure / missing `ytInitialPlayerResponse` / no `captionTracks` → `CommandExecutionError` (typed fail) - Empty parsed XML → `EmptyResultError` (existing path, preserved) - `Strategy.COOKIE` matches YouTube adapter family + `video.js`; cookies/session/consent unavailable → typed fail not silent empty success illusion ## Diff containment Runtime change limited to Step 1 caption track discovery. XML fetch, segment parsing, chapters, raw/grouped formatting all unchanged. ## Verification Local: YouTube adapter tests `19/19` (+1 from new test), `npm run build`, typed-error-lint `192/192`, silent-column-drop `103/103`, doc coverage `140/140`, `docs:build`, listing-id advisory unchanged `13`, `git diff --check`, merge-tree clean. GitHub: build × 3 OS, unit × 2 shards, bun-test, adapter-test, audit, doc-coverage, docs-build all SUCCESS. PR CLEAN/MERGEABLE. Author: kagura-agent (fork). Lead: codex-mini0. Aux: First-principles-0. Coordination: pr-monitor. |
||
|
|
195333ff8a |
fix(xiaohongshu): improve image publishing — creator-center URL + tab priority + DataTransfer fallback (#1380)
Xiaohongshu image-note publishing reliability fixes for creator center UI (legacy raw-Error write command, not a typed-error migration). ## 3 changes (one publish-path repair) 1. **Open creator publish in image mode**: append `target=image` to the publish URL so it loads directly in image mode instead of default 2. **Exact `图文` tab priority**: prefer exact tab text matching before broad `startsWith/includes`, reducing parent-container misclicks while keeping fallback for UI wording variants 3. **DataTransfer fallback for `Chrome Not allowed`**: when CDP `setFileInput` returns the permission/bridge denial error, fall through to the existing DataTransfer upload path (CDP-first remains primary to avoid base64 bridge/payload limits) ## Lead hardening (`edf8107d`) Added `clis/xiaohongshu/publish.test.js` regression coverage for all three claimed behaviors: - `target=image` creator URL locked - exact tab text matched before broad fallback - `Chrome Not allowed` falling into DataTransfer path ## Better-solution evaluation (lead + aux 一致) - **CDP-first kept**: CDP avoids base64 payload/bridge limits; `Not allowed` is a known permission failure class where fallback is appropriate. DataTransfer-first would weaken the common path and reintroduce large-payload fragility. - **Exact tab text first**: XHS creator markup is private and volatile, selector-only alternative not clearly more stable. Exact text reduces misclicks while broader fallback + post-click `video_surface` check preserve resilience for wording shifts. If exact text disappears, command fails fast with screenshot instead of silent video-mode publish. - **Scope boundary self-imposed**: not expanding to typed-error migration (publish.js is legacy raw-Error and typed-error-lint already accounts for it). ## Verification Local: xiaohongshu publish tests `12/12`, typecheck, build/manifest, docs:build, typed-error-lint `189/189`, silent-column-drop `103/103`, doc coverage `140/140`, node --check, git diff --check. GitHub: build × 3 OS, unit shards, bun-test, adapter-test, audit, docs-build, doc-coverage all SUCCESS. PR CLEAN/MERGEABLE. Author: E2ern1ty (fork). Lead: codex-mini1. Aux: First-principles-1. Coordination: pr-monitor. |
||
|
|
3b585fb4d1 |
feat(grok): add browser chat baseline commands (read/history/detail/new/send/status) (#1392)
Phase 3 — Grok adapter baseline (LLM browser-chat command family, parallel to ChatGPT/Qwen/Yuanbao). ## Surface 6 commands: `status` / `history` / `read` / `detail` / `new` / `send`. Site-local `clis/grok/utils.js` justified by 6 commands sharing helpers, not over-abstraction. ## 4-head review iteration 1. **`b4e81bad`** — initial baseline (12 Grok/shared files) 2. **`0a8112fc`** — mechanical rebase (CHANGELOG conflict only, all 12 Grok files preserved business-equivalent through rebase) 3. **`481e87e2`** — security fix: `parseGrokSessionId()` SSRF-shape vulnerability close — switched from regex string match to `new URL()` parser with branch separation: - Bare UUID mode: only exact UUID shape (no URL/query suffix accepted) - URL mode: requires `https` scheme + exact `grok.com` or subdomain host + exact `/c/<uuid>` path 4. **`a082023c`** — test-only hardening: 2 additional negative anchors covering existing implementation rejections (bare UUID `?next=abc` query tail / `grok.com.evil.com` host-suffix trick) ## Negative anchor coverage (8 cases) http / off-domain / fakegrok / host-suffix subdomain / embedded URL / path suffix / UUID-tail / bare query tail ## Better-solution evidence form LLM browser-chat family pattern (matching ChatGPT/Qwen/Yuanbao baseline) + 5 live probes — not first-site hostile scrape. TipTap editor API send seam (`editor.commands.focus/clearContent/insertContent`) is correct boundary because Grok ignores DOM input events; isolated in `sendMessage()`. Lack of full TipTap mock = residual risk, not blocker. ## Invariants locked - `parseGrokSessionId()` URL parser branch separation (bare UUID exact / URL exact path) - `history --limit` rejects invalid/out-of-range - `status` uses `null` for unknowns (no fabrication) - Bubble extraction preserves image-only assistant turns (no silent HTML-only drop) - Auth/empty semantics aligned with LLM browser-chat baseline family ## Verification Local: Grok adapter tests `28/28`, typecheck, build/manifest, docs:build, typed-error-lint `189/189`, silent-column-drop `103/103`, doc coverage `140/140`, listing-id advisory `13` unchanged, diff-check clean. GitHub: build ubuntu/macos/windows × unit-test 1/2 + 2/2, bun-test, adapter-test, audit, doc-coverage, docs-build all SUCCESS. PR CLEAN/MERGEABLE. Lead: codex-mini1. Aux: First-principles-1. Coordination: pr-monitor. |
||
|
|
b2ebe211d1 |
feat(yuanbao): add browser-web baseline commands (status/read/detail/history/send) (#1394)
Wire up the standard browser-LLM command surface for Yuanbao, matching the recently shipped chatgpt + claude + qwen baselines: - status — login + current model + (agentId, convId) + URL - read — render the visible conversation as User/Assistant rows - detail — open `<agentId>/<convId>` and read its messages - history — list sidebar conversations with stable IDs - send — fire-and-forget, returns once the send button has been clicked Refactor `ask.js` to share helpers (`sendYuanbaoMessage`, `normalizeBooleanFlag`) with the new commands via `shared.js`, keeping the public ask behavior intact. Notable bits: - `parseYuanbaoSessionId` accepts only full chat URLs or `<agentId>/<convId>` pairs — Yuanbao chat URLs encode both, and silently opening the wrong agent on a bare UUID is a worse failure mode than throwing. URL regex anchored with `(?:[/?#]|$)` so 37+ char tails reject rather than truncate. - `sendYuanbaoMessage` polls the send button (up to 3s) for the React re-render that drops `style__send-btn--disabled___*` after composer input — a fixed wait raced the debounce and produced silent no-op clicks. - `getYuanbaoMessageBubbles` uses `data-conv-id`/`data-conv-idx`/ `data-conv-speaker` attributes for stable per-turn identity (was relying on innerHTML alone). - Status surfaces both human label (`Yuanbao`) and `dt-model-id` (`hunyuan_gpt_175B_0404`) — sentinel strings would silently look like a real model name; null is the typed-unknown signal. Verified: 25 unit tests pass; targeted live smoke for status/read/detail/ history/new/send + ask round-trip on yuanbao.tencent.com. |
||
|
|
b9b87a5c64 |
refactor(facebook/notifications): pipeline→func + typed errors + 7-col contract + runtime upfront limit (Phase 3 P5, #1391)
First Facebook adapter — Pattern C HTML scrape (lead 5 + author 4 = 7 endpoint family probe matrix dual-source negative evidence: graphql×3 / m.facebook redirect / login.php / checkpoint.php / fetch-patch / Messenger relay / ajax legacy 全 unauth 不可达, DOM walk over rendered notification rows + path-anchored auth detection 是当前 reviewable boundary). Caller-visible delta: 3 cols (index/text/time) → 7 cols (+unread/+url/+notif_id/+notif_type). [Bug fix] — 5 silent failures resolved - silent-bad-shape: text.substring(0,150) → full body via per-row 'Mark as read' aria-label - silent-bad-shape: time || '-' sentinel → string|null typed unknown - silent-column-drop: unread badge / anchor href / notif_id / notif_t 暴露 - silent-empty-row: /login(.php)? + /checkpoint(.php)? redirect 返 [] → AuthRequiredError; empty/no-recoverable-text → EmptyResultError - silent-clamp: limit 越界 silent clamp → ArgumentError (1-100), upfront before any navigation (navigateBefore: false) [Structural refactor] - pipeline → cli() func form + Strategy.COOKIE + navigateBefore: false (runtime upfront invariant 与 #1387 standard 拉齐) - module-level pure exports: normalizeNotificationsLimit, stripMarkAsReadPrefix, stripAnchorChrome, parseNotifQuery, extractNotificationRowsFromDoc, isFacebookAuthRedirectPath, buildNotificationsScript - Live IIFE 通过 \${fn.toString()} 嵌入 (dianping #1313 / hupu #1387 / xiaoe #1388 lineage) - Locale 表 6 prefix / 4 badge label 显式列出 - AUTH_REQUIRED: sentinel → Node-side AuthRequiredError mapper [Typed-error hardening] - Path-anchored auth helper: isFacebookAuthRedirectPath(/^\/(?:login|checkpoint)(?:\.php)?(?:\/|\$)/i) — domain-invariant-first encoding (FB top-level auth-only invariant), 排除 /loginhelp /help/login /account/login/identify - Three-layer navigateBefore=false invariant lock: registration assertion + manifest absence + executeCommand runtime page.goto-zero-call (test layer 与 invariant layer 完整对齐) - Row-level silent-empty-row defense: anchor rows with no recoverable body text 直接 skip, 不 emit text:null success row [Doc fix] - docs/adapters/browser/facebook.md notifications enrichment + Output table (列类型 / null vs sentinel 语义) + auth/empty error contract - Boy Scout audit: cross-checked profile / feed / search / marketplace-listings / marketplace-inbox 例 commands 与 args 定义一致 Tests - notifications.test.js 39/39 + src/execution.test.ts 21/21 - Anti-pattern regression guards: not.toMatch(/text\.substring\(0,\s*150\)/) + not.toMatch(/time\s*\|\|/) - JSDOM frozen-fixture (slim 13 lines, 0 blank): header listitem skip / full text / unread badge / query parsing / null time / blank-row skip / relative href absolute / 19-case auth path matrix - typed-error-lint baseline 192 → 191 (silent-sentinel resolved 1) Review iterations (4 head, A 组 codex-mini0 lead + First-principles-0 aux): 1. |
||
|
|
381f095706 |
feat(qwen): add detail command + fix stale message bubble selector (#1390)
* feat(qwen): add detail command + fix stale message bubble selector `getMessageBubbles` was matching `[data-msgid="<id>-question|answer"]` from an older Qianwen frontend. The reshipped DOM no longer carries that attribute on chat turns; `[data-message-id]` now lives on citation cards inside assistant responses, so the old selector silently returned an empty list and `qwen read` had been silently broken. Rewire to walk `[data-chat-question-wrap]` and `[data-chat-answers-wrap]` in DOM order (correct Q/A interleaving) and synthesize stable IDs from the nearest sibling `data-req-id` so `waitForAnswer.seenAssistantId` and read/ask/detail dedupe paths keep working. Verified live against an existing conversation: 3 user turns + 3 assistant turns extracted; old selector returned 0. `qwen detail <id|url>`: open a specific conversation by ID or full chat URL, poll up to 20s for the transcript to render, return Role/Text rows. Adds `parseQianwenSessionId` (5 unit tests covering ID/URL parsing + ArgumentError on malformed input). Reuses the same site-level browser session as `read`/ `ask` so consecutive calls continue in the same Qwen tab. - clis/qwen/detail.js (new) - clis/qwen/utils.js (parseQianwenSessionId + getMessageBubbles rewire) - clis/qwen/utils.test.js (new) - docs/adapters/browser/qwen.md (detail entry + options/columns) - cli-manifest.json (regenerated) * fix(qwen): anchor URL regex to reject 33+ hex tail truncation codex-coder review on PR #1390 caught that `https://www.qianwen.com/chat/<33+ hex>` would silently truncate to the first 32 chars and open the wrong conversation. Adds end-of-input / slash / query / fragment boundary to the URL match group and two new unit-test cases (digit tail + letters tail) covering the truncation gap. |
||
|
|
99986c3101 |
feat(chatgpt): add browser chat baseline commands
Add ChatGPT web ask/send/read/history/detail/new/status alongside existing image support. Tighten ChatGPT web helper selectors and typed error contracts, update docs/changelog, regenerate manifest, and seed local ChatGPT verify fixtures for ask/read. |
||
|
|
6f7eb6a76a |
refactor(xiaoe x3): pipeline→func + typed errors + content silent-drop fix (Phase 3 P1)
Phase 3 P1 (xiaoe catalog/courses/content) — pipeline→func refactor + typed-error hardening + content silent-drop bug fix + URL upfront validation + inherited legacy doc fix。
## Tags (PR body honesty 演进 dual-nature framing 试用)
- **[Bug fix]** `xiaoe/content` silent-column-drop (caller-visible delta)
- **[Structural refactor]** `xiaoe/catalog` + `xiaoe/courses` pipeline→func 包壳 (parity by construction, IIFE 字节级保留)
- **[Typed-error hardening]** 三 func `page.goto` + `page.evaluate` failure 包成 `CommandExecutionError`; `content/catalog` URL upfront `ArgumentError` (missing/malformed/non-https/off-domain) before navigation
- **[Doc fix]** `docs/adapters/browser/xiaoe.md` `courses --limit 10` (legacy doc 错误 inherit) + `--url` wording → 实际 positional `url` (manifest aligned)
## Per-tag detail
### [Bug fix] content silent-column-drop (real caller-visible bug)
adapter 名"提取小鹅通图文页面内容为文本", IIFE 返 `{title, content, content_length, image_count, images}`, 但 columns 只声明 `[title, content_length, image_count]` → `content` (那段文本本身) 被 silent drop。**用户拿到 "1234 chars" 但拿不到那 1234 chars** — adapter 名字撒谎了。
- Fix: 公开列 `[title, content, content_length, image_count]`, `content` 真 caller-visible delta
- Choice A (vs B reshape): legacy `images` 是 `JSON.stringify(slice(0, 20))` 截断/stringified 坏合同, **不暴露成新列** (避免把 silent-bad-shape 升级成公开坏合同), 留 follow-up 另开 explicit media/images contract
- `image_count` 用 `countXiaoeImages(doc)` 全页计数, 不 slice (既有 metadata 质量修正)
### [Structural refactor] catalog + courses pipeline→func wrapper (parity by construction)
- `pipeline:[]` form → `func` form
- IIFE body 字节级保留 (Xiaoe 没 public REST, Vue 私有 runtime 是唯一稳定 hook, JSDOM 复刻不了 Vue tree)
- Pure helpers extracted: `pickContentText`, `countXiaoeImages` (content) / `typeLabel`, `buildItemUrl`, `chapterUrlPath` (catalog) / `buildCourseUrl` (courses)
- IIFE 通过 `\${fn.toString()}` 嵌同一份代码 (dianping #1313 / hupu #1387 同模式)
- No live verify acceptable: IIFE 字节级保留 + helper 全 unit-test + manifest column shape 不变 = 行为 parity by construction
- `buildScript` 反向断言 `images.slice(0, 20)` legacy anti-pattern 不出现 (anti-pattern regression guard, 同 #1387 `documentElement.outerHTML` 反向 guard)
### [Typed-error hardening] 三 func navigation + evaluate boundary
- `requireXiaoePageUrl()` for `content/catalog`: missing/malformed/non-https/off-domain URL → upfront `ArgumentError` before `page.goto` (test asserts `expect(page.goto).not.toHaveBeenCalled()`)
- `content/catalog/courses`: `page.goto` moved inside try, navigation/evaluate failures both wrap as `CommandExecutionError`, no raw CDP/browser error path leaks
- Empty shell stays `EmptyResultError` (no reliable login-wall signal to justify `AuthRequiredError`, 避免 false positive — 应用 #1384 secUid 教训)
### [Doc fix] inherited legacy doc errors
- `xiaoe courses --limit 10` example removed (no `--limit` arg in manifest, legacy doc 错误 inherit)
- positional `url` wording aligned with manifest (was incorrectly `--url`)
- 同 #1386 positional docs 教训, 但延伸到 "继承 legacy doc 错误也是新 PR 责任" (Boy Scout typed-error hardening 在 doc 层延伸)
## Tests: 46/46 green
- 3 cmd registration contract
- pure helper unit tests (selector chain / image filter / URL priority / type label fallback / no synthetic URL)
- `buildScript` invariants (`images.slice(0, 20)` 反向断言)
- wire tests: ArgumentError upfront (BEFORE page.goto), EmptyResultError empty rows + empty content, CommandExecutionError navigation/evaluate failure, rows verbatim happy path
## Lint gates
- typed-error-lint 190/190 (no new) ✓
- silent-column-drop 103/103 (no new) ✓ (注: `pipeline:[]` IIFE string template AST walker 看不进, lint follow-up scope)
- doc-coverage 140/140 ✓
- listing-id-pairing advisory unchanged 13 ✓
## GitHub checks (head
|
||
|
|
e610260705 |
refactor(hupu/hot): pipeline→func + querySelectorAll + 4 enrichment columns (Phase 3 P3)
Phase 3 P3 (hupu/hot) — pipeline→func refactor + 2 真 bug 修 + 4 列 enrichment + JSDOM-frozen-fixture test pattern (#1313 复用) + anti-pattern regression guard。
## Summary
- Pipeline form (`pipeline:[]` + `documentElement.outerHTML` regex) → `func` form (`querySelectorAll('.t-info')` DOM walk)
- **Bug 1 修**: outerHTML regex 静默漏行 (markup 抖动就漏, mocked test 抓不到)
- **Bug 2 修**: regex 抓所有 9-digit 锚点 → ~70 个 anchor 但页面只 render 60 个 `.t-info` row → legacy adapter 每次返 ~10 个 phantom 行 (导航链接 conflated 成 thread 行)
- **4 enrichment columns** (4→8): `lights` (亮 count int|null, 万 expanded `1.2万→12000`) / `replies` (回复 count int|null) / `forum` (per-row sub-section) / `is_hot` (bool 暴露 hupu \" hot\" marker, 不 filter 行序保持页面顺序)
- columns/manifest/docs sync: `[rank, tid, title, lights, replies, forum, is_hot, url]`,`null` vs `0` 语义清楚
## Typed errors
- `--limit` 上游 `ArgumentError` for 0/-1/>100/1.5/non-numeric (BEFORE `page.goto`,**不 silent clamp**)
- 空页 `EmptyResultError`
- `page.evaluate` failure 包成 `CommandExecutionError` (test regression locked)
## JSDOM frozen-fixture test pattern (#1313 复用)
- 抽 `extractHupuHotRowsFromDoc(doc, limit, parseCount)` 为 module-level pure export
- in-page IIFE 通过 `\${fn.toString()}` 嵌同一份代码
- JSDOM test 直接调 export against `__fixtures__/hot-home.html` (slim 6-row hand-crafted fixture)
- 17/17 tests green (contract / normalize / parseCount / extract / buildHotScript invariants / wiring / phantom-anchor exclusion / evaluate-error envelope)
## Anti-pattern regression guard (#1313 fixture pattern 延伸)
- `buildHotScript` 反向断言 `not.toContain('documentElement.outerHTML')` 锁不回退到旧 broad regex
- `buildHotScript` 反向断言 `not.toContain('regex.exec')` 同向锁
- fixture 顶部 `.t-info` 外的 9-digit phantom anchor `639999999` 反向锁: 旧 broad regex 会抓到, 新 `.t-info` extractor 不抓 — 把 fixture 反向验证从断言层升到证据层
## Better-solution check (live probe evidence-based)
DOM `.t-info` = 60 visible rows, `window.\$\$data.pageData.threads` = 70 (10 hidden/non-rendered)。对"首页可见 hot rows" 任务, DOM walk 比 bootstrap JSON 更贴 source of truth (后者会引入 hidden/不渲染条目)。这条 60 vs 70 数字是设计决策的硬 justify, 不是设计意见。
## Lint gates
- typed-error-lint 190/190 (no new) ✓
- silent-column-drop 103/103 (no new) ✓
- doc-coverage 140/140 ✓
- listing-id-pairing advisory unchanged 13 ✓
## GitHub checks (head
|
||
|
|
464de7059e |
refactor(tiktok): write commands -> button-walker Route 1 with typed errors (Phase 3 P0.5)
Phase 3 P0.5: refactor 3 TikTok write commands (comment, follow, unfollow) from time-window-wait UI flow to a button-walker + state-verification path with a typed-error boundary, sharing a parallel helper structure to the #1384 read PR. Two-layer helper boundary (clis/tiktok/utils.js extension): - BUTTON_WALKER_HELPERS (browser side): button-walker (locate / pre-click state read / click / state-verify post-click) + cleanText reuse + cookie/auth-secUid plumbing for write-auth + plain Error throws on contract violations - throwButtonWalkerError() (Node side): map browser-thrown errors -> typed CommandExecutionError (button missing / state-verify fail / captcha / rate-limit / navigation/eval/empty-row defensive failures) / AuthRequiredError (cookie + viewer secUid) / ArgumentError (upfront input validation). Explicitly NO EmptyResultError mapping (button contract violation is not an empty result, per #1384 R4 lesson on auth-vs-empty classification). Per command: - comment <video-url> <text>: button-walker click + state-verify by checking comment-list state (not wait-2s) - follow <username>: pre-click state read distinguishes idempotent fast path (`already-following` / `already-friends`) from post-click success (`followed`). Post-click result causality preserved (post-click never returns `already-*`). - unfollow <username>: pre-click `already-not-following` fast path; post-click `unfollowed`. result enums (per row): - comment: `posted` (no idempotent path - comments cannot dedupe) - follow: `followed` | `already-following` | `already-friends` (last two pre-click only) - unfollow: `unfollowed` | `already-not-following` (last one pre-click only) retryable contract (in hint string `retryable=<bool> reason=<...>`): - comment failures: retryable=false reason=server-fan-out - follow/unfollow failures: retryable=true reason=idempotent (server-side dedupe is safe) Lead push iterations during review (codex-mini1 maintainer-fixes-directly): - |
||
|
|
9a7dd44b3e |
refactor(tiktok): 6 read commands -> page-context API (Phase 3 P0, absorbs #1382)
Phase 3 P0: refactor 6 TikTok read commands (explore, following, friends, live, notifications, user) from DOM/network-intercept to TikTok web's own page-context API endpoints, sharing one helper boundary. Helper boundary (clis/tiktok/utils.js): - BROWSER_HELPERS: in-browser fetchJson + cleanText + asNumber (null/'' -> null preserve missing-vs-zero distinction) + cookie/msToken plumbing - VIDEO_ITEM_NORMALIZER: normalize page-context item -> row shape - assertTikTokApiSuccess(data, label): unify TikTok in-band envelope (status_code/statusCode != 0; code 8 or auth-looking message -> AUTH_REQUIRED; other -> upstream label API failed) - throwTikTokPageContextError() (Node side): map browser-thrown errors -> AuthRequiredError / EmptyResultError / CommandExecutionError Per command: - explore: /api/recommend/item_list/ pagination, --limit upfront ArgumentError - following: /api/user/list/ relationships - friends: /api/user/list/ + cross-filter - live: /api/live/discover/ feed - notifications: /api/notice/multi/ (status 8 -> AUTH_REQUIRED) - user (absorbed from #1382): secUid resolve via __UNIVERSAL_DATA_FOR_REHYDRATION__ -> /api/user/detail/, /api/post/item_list/ pagination, /api/search/general/full/ exact-author fallback. !secUid -> EmptyResultError (NOT AuthRequiredError; auth still covered by HTTP 401/403 + envelope status_code 8/auth-looking msg). source field = bootstrap | profile-api | search-fallback in row/columns/manifest/docs/tests. Closes #1382 (absorbed; #1382 closed without separate merge per WAWQAQ direction). Validation: - clis/tiktok/ tests: 38/38 - typed-error-lint: 190/190 - silent-column-drop: 103/103 - doc-coverage: 140/140 - docs:build pass, manifest no drift - GitHub gates: build x3 / unit x2 / bun / adapter-test / audit / doc-coverage / docs-build all SUCCESS, smoke skipped, MERGEABLE Reviewers: codex-mini0 (lead, push 4 boundary fixes |
||
|
|
b327da5b3c | feat(llm): reuse browser sessions by site (#1385) | ||
|
|
fa7851bb9a | feat(browser): add adapter session reuse (#1383) | ||
|
|
d527571b7d |
test(gov-policy): JSDOM-against-frozen-fixture tests for in-browser extractors (#1340)
* test(gov-policy): JSDOM-against-frozen-fixture tests for in-browser extractors Applies the pattern documented in skills/opencli-adapter-author/references/jsdom-fixture-pattern.md (introduced in #1319 alongside the dianping reference test in #1313) to the gov-policy adapter. Refactor: the inline IIFE inside `page.evaluate` template literal is hoisted to a top-level `extractSearchRows` / `extractRecentRows` function using bare `document` / `location`. Same code now runs identically in: - the live browser (injected via `${extractor.toString()}`) - JSDOM unit tests (with `globalThis.document` / `globalThis.location` swapped) Tests: - 6 new cases in clis/gov-policy/gov-policy.test.js (was commands.test.js). - 3 representative search result cards (1 with real article snippet, 2 with only publish-time in `.description`) and 5 recent listing rows in the fixtures. - ok:false fallback path covered for both extractors. - Lock-in: `要闻` type-tag prefix fusion in title and empty-source contract on recent listings (no `.source` / `.from` elements on that page) are asserted explicitly so a future selector tweak can't silently change them. Reverse-validated against two buggy variants per the reference doc: breaking the title selector and stripping the `要闻` prefix both fail the JSDOM assertions with helpful diffs. Fixture sanitization follows the reference doc step-by-step: scripts / styles / iframes / comments / preload links stripped, image srcs replaced with `placeholder.png`, trimmed to the minimum subtree that exercises the extractor (3 search items, 5 recent rows), all whitespace-only lines removed. * fix(gov-policy): use typed errors for touched commands --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
c6d5da54ee | feat(web): add exhaustive same-origin frame mode (#1373) | ||
|
|
124adf73d1 | fix(web): avoid duplicate iframe diagnostics (#1372) | ||
|
|
829edfea3a | fix(web): include relevant iframes outside main content (#1371) | ||
|
|
67cde0e263 |
enrich(coupang): product detail cmd + replace silent clamp/sentinel/Error with typed errors (#1370)
* enrich(coupang): add product detail cmd + replace silent clamp/sentinel/Error with typed errors Two enrichment changes plus three silent-failure fixes on top of existing search / add-to-cart. New cmd: coupang product ───────────────────────── Pairs with search as the listing↔detail round-trip target. Reads a logged-in product page and extracts a single canonical row with price, original_price, discount_rate, rating, review_count, seller, brand, rocket, delivery_promise, image_url, url. Three-source extractor (JSON-LD Product schema → bootstrap globals → DOM) merged in priority order, mirroring the search.js pattern. The columns use string|null typing — null means "upstream did not provide this field on this product" (e.g. some items have no original_price). Failures (login wall / page mismatch / page failed to render) raise typed errors instead of silently returning empty rows, so callers can treat any returned row as real data. Search column shape: added product_id ───────────────────────────────────── Listing must pair with detail by id. The data was already extracted by normalizeSearchItem; only the columns array needed updating so the field projects through to the rendered row. Per the listing-id-pairing convention (PR #1297) the new column lets agents round-trip rows directly into `coupang product` without re-scraping URLs. Silent-failure fixes ──────────────────── 1. search --limit silent clamp. Old: `Math.min(Math.max(Number(kwargs.limit||20),1),50)` silently rewrote `--limit 999` to 50 and `--limit 0` to 1. New: `parseLimitArg(raw, 20, 50)` throws ArgumentError on out-of-range / non-integer / negative input. Same convention as the typed-fail-fast memory & PR #1289. 2. search --page silent clamp. Old: `Math.max(Number(kwargs.page||1),1)` silently lifted negative pages. New: parsePageArg throws ArgumentError on non-positive input. 3. Generic `throw new Error(...)` → typed errors. - Empty query, unsupported --filter, missing --product-id/--url → ArgumentError - Login wall detection → AuthRequiredError('coupang.com', ...) - Empty result / filter-not-rendered → EmptyResultError - PRODUCT_MISMATCH / OPTION_REQUIRED / button-not-found / unknown ack failure (add-to-cart) → CommandExecutionError - The PRODUCT_MISMATCH and `actualProductId || 'unknown'` sentinel were also fixed (silent-sentinel was the audit hit there). Coverage ──────── - 21 contract assertions in clis/coupang/coupang.test.js covering parseLimitArg / parsePageArg (no silent clamp), registry shape (search has product_id, product is read-class with expected columns, add-to-cart is write-class), and typed-error pre-flight rejections (empty query / bad filter / out-of-range limit & page / missing detail args). - Manifest 763 → 764 (+1 entry: coupang/product). - Audits: typed-error-lint 196 → 194 (resolved 2 silent-clamp/sentinel baseline entries; baseline updated). silent-column-drop 103/103 unchanged. * fix(coupang): tighten product id and browser errors * fix(coupang): require real product urls |
||
|
|
a5a3248a77 |
refactor(linux-do): remove deprecated hot/category/latest compat shims (#1368)
* refactor(linux-do): remove deprecated hot/category/latest compat shims
The three shims have been pure backward-compat wrappers since linux-do/feed
became the unified entrypoint. With no stable release commitment to preserve,
they are pure surface cost: 3 manifest entries, 3 deprecated branches in help
output, and a `buildLinuxDoCompatFooter` helper that exists only to feed them.
- delete clis/linux-do/{hot,category,latest}.js
- drop now-orphaned `buildLinuxDoCompatFooter` from feed.js and unexport
`executeLinuxDoFeed` (no external consumers remain)
- remove the Compatibility section in docs/adapters/browser/linux-do.md
- regenerate cli-manifest.json (-125 lines)
BREAKING CHANGE: `opencli linux-do hot|category|latest` are removed. Use
`opencli linux-do feed --view top --period <period>`,
`opencli linux-do feed --category <id-or-name>`, and
`opencli linux-do feed --view latest` instead.
* fix(linux-do): finish compat shim removal
|
||
|
|
dcaae37068 |
refactor(registry): remove dead adapter metadata (#1369)
* refactor(registry): remove dead adapter metadata * docs(changelog): note header strategy removal |
||
|
|
12d88e4b23 |
refactor(runtime): unify command timeout into a single --timeout arg (#1364)
* refactor(runtime): unify command timeout into a single --timeout arg Drop the cli-level `timeoutSeconds` build-time ceiling field. A command now opts into runtime-enforced timeouts purely by declaring an arg named `timeout`; the user-facing `--timeout` value (its default or override) is the single authoritative knob, used both by the adapter polling loop and by the runtime ceiling (with a 30s padding for return + closeWindow + trace export). Behavior: - Browser commands without a `--timeout` arg fall back to OPENCLI_BROWSER_COMMAND_TIMEOUT (default 60s, unchanged). - Non-browser commands without a `--timeout` arg now run unbounded rather than against the previously implicit `timeoutSeconds` cap. Affected commands keep their old caps via newly added `--timeout` args. - LLM adapters (gemini/claude/deepseek/doubao/qwen/yuanbao ask) keep their current `--timeout` defaults; the runtime ceiling is now strictly more generous (userTimeout + 30s vs. the previous 180s cap), so `--timeout 600` actually buys 600s of polling rather than dying at 180s. Closes the design discussion that started from PR #1227, which proposed a per-site `OPENCLI_GEMINI_ASK_TIMEOUT` env var to work around the same underlying mismatch. * fix(timeout): wire --timeout arg into chatgpt/gemini image adapter polling codex-coder review on PR #1364 caught that the new --timeout arg I added to chatgpt/image and gemini/image only drove the runtime ceiling — the adapter still hardcoded `const timeout = 120`, so users passing --timeout 240/600 saw runtime allow 270s/630s but the adapter stop polling at 120s. That recreated the same single-knob mismatch this PR was meant to delete. Also add the browser-path runWithTimeout assertion codex-coder flagged as missing: a browser command with --timeout default=5 must call runWithTimeout with timeout: 35; a browser command without --timeout arg must fall back to DEFAULT_BROWSER_COMMAND_TIMEOUT. Image adapters now read kwargs.timeout and reject non-positive-integer values with ArgumentError (no silent fallback). chatgpt/image.test.js updated to pass an explicit timeout when calling .func directly (the test bypasses arg coercion). * fix(runtime): reject invalid timeout ceilings * fix(timeout): normalize timeout args to integer values * fix(timeout): preserve remaining command ceilings * fix(runtime): validate timeout before browser setup |
||
|
|
4ef2cb8b1c |
enrich(toutiao): hot board (public) + bug fixes (silent column drop, partial render) (#1366)
* enrich(toutiao): hot board + bug fixes (silent column drop, partial render) Per WAWQAQ "丰富现有 adapter" pivot — Phase 2 site #3. ## New command - `toutiao hot` (Strategy.PUBLIC, browser:false) — public homepage hot board via the toutiao.com hot-event/hot-board endpoint. No login required. Returns 8 stable columns (rank/id/title/query/hot_value/label/url/image). ## Bug fixes for `toutiao articles` - **Silent column drop fixed**: `parseToutiaoArticlesText` previously did `if (title && stats) push(...)`, silently dropping any row where the stats span hadn't finished rendering by the time page.innerText was read. Slow-render bugs were invisible — adapter looked "complete" while writers saw extra rows in the dashboard. Partial rows now surface with `null` stat columns. - **Silent clamp on `--page` removed**: out-of-range / non-integer values raise `ArgumentError` with explicit bounds [1, 4]. Same validation reused by both `articles` and `hot` via `parseArticlesPage` / `parseHotLimit` in `utils.js`. - **Empty result typed**: zero-row scrape now raises `EmptyResultError` instead of returning `[]` silently (would otherwise look like a legitimate "no articles" response). ## Refactor - Parser logic extracted to `clis/toutiao/utils.js` (alongside hot-row mapping, validators, and the hot-board URL constant). - `articles.js` switches from declarative `pipeline:` to imperative `func` form so `parseArticlesPage` validation can run before the navigation step (declarative pipeline can't pre-validate args). - Strategy is now explicit: `Strategy.COOKIE, browser: true` for articles (creator dashboard is logged-in only). ## hot field map `ClusterIdStr` (or numeric `ClusterId`) → id; `Title` → title; `QueryWord` → query (falls back to title); `HotValue` → hot_value (non-negative numeric, else null); `Label`, `Url`, `Image` → respective columns. `pickImage` walks `Image.url` → first truthy `Image.url_list[]`. Empty-title rows are dropped (returns null) before ranks are densely re-assigned 1..N. ## Tests 29 contract assertions across `parseArticlesPage` / `parseHotLimit` / `parseToutiaoArticlesText` / `mapHotRow` + registry-level shape checks + `hot` adapter func behaviour (typed errors / no silent clamp / fetch failure paths / dense-rank). ## Audits - typed-error-lint: 196 = 196 (unchanged baseline) - silent-column-drop: 103 = 103 (unchanged baseline) - listing-id-pairing: hot has `id` column (round-trippable when a detail command lands later); advisory list unchanged. ## Manifest 757 → 758 entries (+1 for `hot`). ## Doc - index.md: toutiao mode 🔐 → 🌐/🔐 (hot is public, articles is logged-in) - toutiao.md: per-command mode/domain table + column docs + prerequisites * fix(toutiao): tighten hot and articles contracts |
||
|
|
69ee36f997 |
fix(linkedin): surface detail_error on --details (no silent catch / no silent empty) (#1363)
* fix(linkedin): surface detail_error on --details (no silent catch / no silent empty)
The previous --details enrichment path had two indistinguishable failure modes
that both produced `description: '', apply_url: ''`:
1. `if (!job.url)` early return — row had no jobId, so we couldn't navigate.
2. `} catch {}` — page.goto / page.evaluate threw (network, timeout, parse error).
Callers couldn't tell "upstream had no description" from "we failed to fetch",
and the catch swallowed every error without logging. For an enrichment that
costs one page navigation per row, silent failure is especially harmful — users
just see an empty cell with no way to debug.
Fix: replace empty strings with `null` for missing/failed rows, add a new
`detail_error` column (string|null) carrying a short typed reason:
- 'no url' — row had no jobId
- 'fetch failed: <msg>' — page.goto / page.evaluate threw
- 'missing description' — page loaded but body was empty
- null — success
Every failure is also logged to stderr with the offending URL so debugging is
possible. Per-row failures still don't abort the batch (the original intent),
but they're now visible.
Tests: 13 new contract assertions in clis/linkedin/search.test.js covering
parseCsvArg, mapFilterValues (ArgumentError on unknown values), decodeLinkedinRedirect,
and 5 enrichJobDetails paths (no-url / goto-throw / empty-description / success /
multi-row-mixed). Added `export const __test__` for testability.
Audits clean: typed-error-lint 196/196, silent-column-drop 103/103.
* fix(linkedin): fail fast on auth walls
|