Commit Graph

164 Commits

Author SHA1 Message Date
jakevin fa9b38cd92 feat(reddit): add whoami, home, subreddit-info read commands (#1491)
* feat(reddit): add whoami, home, subreddit-info read commands

Closes gap against jackwener/rdt-cli — three commands the existing 17 reddit
adapters were missing:

- `reddit whoami` — show the currently logged-in identity (fields:
  Username, ID, Post / Comment / Total Karma, Account Created, Gold, Mod,
  Verified Email, Has Mail, Inbox Count). Probes `/api/me.json` with
  two-pronged auth detection (401/403 OR `data.name` missing on 200 —
  Reddit returns 200 with an empty body for stale anon sessions, see PR
  #1428).

- `reddit home` — personalized Best feed (`/best.json`). Distinct from
  the public `frontpage`/`r/all` command: enforces login via the same
  two-pronged auth check rather than silently degrading to the
  unauthenticated default feed. `--limit` accepts [1, 100] — out-of-range
  raises `ArgumentError` before navigation, no silent clamp.

- `reddit subreddit-info` — subreddit metadata (Name, Title, Subscribers,
  Active Now, NSFW, Type, Description, Created, URL) from
  `/r/<X>/about.json`. Banned / private / quarantined / 404 subreddits
  raise `EmptyResultError` so the output table never holds a silent
  sentinel row.

All three use Strategy.COOKIE + siteSession:'persistent' matching the
existing reddit adapters, validate args upfront before `page.goto`, and
use the 5-kind discriminated-union pattern (kind: auth/http/missing/
exception/ok) from PR #1428 to map page.evaluate results to typed errors
on the Node side. Intermediate object keys deliberately avoid the
declared columns (`field`/`value`/`rank`/etc.) per the silent-column-drop
audit sediment from PR #1329.

Tests: 28 new (whoami 6, home 9, subreddit-info 13); full reddit suite
38/38. Audits: typed-error-lint 189/189 (0 new), silent-column-drop
103/103 (0 new). Manifest 812 → 815.

Refs: https://github.com/jackwener/rdt-cli

* fix(reddit): tighten new read command failure contracts

* fix(reddit): treat inaccessible subreddit info as empty
2026-05-12 04:10:44 +08:00
jakevin eb59b7444d feat(ctrip): add hotel-search + flight browser-mode commands (#1481) (#1489)
* feat(ctrip): add hotel-search + flight browser-mode commands

Closes #1481.

Two new browser-mode commands on top of the existing public `search` /
`hotel-suggest` pair:

- `ctrip hotel-search <city> --checkin --checkout [--limit]` reads
  `window.__NEXT_DATA__.props.pageProps.initListData.hotelList` on
  `hotels.ctrip.com/hotels/list`. SSR-rendered first page ships ~13
  entries; the server ignores `&pageSize=N` so limit caps at 30 with
  default 10. AuthRequiredError surfaces when Ctrip redirects to the
  captcha gate.

- `ctrip flight <from> <to> --date [--limit]` searches one-way flights on
  `flights.ctrip.com/online/list/oneway-…`. The post-load XHR is not
  currently captured by the daemon network buffer (per the known
  daemon_capture_pipeline_bug_2026_05_07 in agent memory), so rows are
  pulled from `.flight-list > span > div` cards via a position-anchored
  innerText parser. A generic `buildScrollUntilJs(selector, target)`
  helper mirrors the PR #1487 xiaohongshu scroll-until pattern with the
  selector parameterised. Round-trip + airline filters are out of scope
  for v1.

All argument validation (IATA / ISO date / city ID / limit range) fires
upfront before any `page.goto`, per the PR #1387 boundary standard. No
silent clamps, no sentinel rows: rows missing required fields are
dropped, and end-state checks raise `ArgumentError` /
`AuthRequiredError` / `EmptyResultError` as appropriate. The new
`mapHotelRow` / `pickHotelMapCoords` / `buildFlightExtractJs` /
`buildScrollUntilJs` helpers live in `clis/ctrip/utils.js` alongside the
existing suggest helpers.

Docs at `docs/adapters/browser/ctrip.md` now distinguish the public
suggest commands from the browser-mode commands and document each
command's columns + caveats.

Verified:
- 61/61 vitest tests in `clis/ctrip/ctrip.test.js` (including JSDOM
  exercises of `buildFlightExtractJs` and full `mapHotelRow` shape parity)
- `check:typed-error-lint` 189/189 (0 new)
- `check:silent-column-drop` 103/103 (0 new)
- `build-manifest` clean — 812 entries total (was 810)

* fix(ctrip): harden browser search failure contracts

* fix(ctrip): tighten browser empty-vs-parser failures
2026-05-12 03:43:32 +08:00
Gaurav Saxena 150551be8c feat(reddit): add reply command for replying to comments (#1428)
* feat(reddit): add reply command for replying to comments

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(reddit/reply): replace silent-sentinel rows with typed errors

reply.js originally mirror-copied comment.js's failure pattern: returning
[{ status: 'failed', message: 'HTTP 403' }] on auth/HTTP/Reddit errors and
relying on the caller to inspect the row instead of throwing. That's the
'silent-sentinel' anti-pattern from typed-errors.md — failures should
surface as typed errors so an agent can actually branch on them.

Round 21 lesson (f) — "grandfathered-not-exempt + helper-refactor boundary
is new" — applies: comment.js / upvote.js / save.js can stay grandfathered,
but a brand-new file does not inherit that exemption.

Changes:
- Throw AuthRequiredError when /api/me.json or /api/comment returns 401/403,
  or when /api/me.json returns 200 but data.name is missing (stale anon
  session — empty modhash alone isn't a strong enough signal).
- Throw CommandExecutionError for non-2xx HTTP and for non-empty
  data.json.errors (e.g. RATELIMIT, NO_TEXT, TOO_OLD).
- Drop the over-defensive `if (!page) throw ...` — registry guarantees a
  page object when browser:true.
- Intermediate result object uses `kind` discriminator + `detail` /
  `httpStatus` / `where` keys that don't overlap with columns
  ['status','message'], so the silent-column-drop audit stays quiet
  (per PR #1329 sediment).

Verified:
- npx tsc --noEmit clean
- node scripts/check-typed-error-lint.mjs → 189/189, 0 new
- node scripts/check-silent-column-drop.mjs → 103/103, 0 new
- npx vitest run clis/reddit src/convention-audit → 11/11 pass
- node ./dist/src/main.js validate → 0 errors

Success path is unchanged: still returns
[{ status: 'success', message: 'Reply posted on t1_<id>' }].

* fix(reddit): harden reply command contract

* fix(reddit): reject suffixed reply urls

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-05-12 02:33:09 +08:00
Benjamin Liu 64ac362a40 feat(rednote): add rednote.com adapter mirroring xiaohongshu read commands (#1136) (#1475)
* feat(rednote): add rednote.com adapter mirroring xiaohongshu read commands (#1136)

Implements rednote.com support as discussed in issue #1136. The mainland
xiaohongshu adapter stays in place; international users redirected to
www.rednote.com now have a CLI without a copy-pasted adapter.

Issue #1136 documents that xiaohongshu and rednote share DOM selectors,
URL paths, API paths, response schema, cookies, and the xsec_token auth
mechanism. The only material differences:

  Layer            xiaohongshu                rednote
  Web host         www.xiaohongshu.com        www.rednote.com
  API host         edith.xiaohongshu.com      webapi.rednote.com
  Security host    fe-static.xhscdn.com       as.rednote.com
  Cookie root      .xiaohongshu.com           .rednote.com
  Search gate      Inline text                Full-screen modal + text

## Architecture (minimal)

`clis/xiaohongshu/*` keep all selector / regex / extraction logic. Each
command file is touched minimally to export the IIFE or pipeline so the
sibling adapter can reuse it:

  search.js          + export const buildSearchExtractJs(webHost)
                     + export const command = cli({...})
  note.js            + export const NOTE_EXTRACT_JS
                     + export const command = cli({...})
  comments.js        + export function buildCommentsExtractJs(withReplies)
                     + export parseCommentLimit
                     + export const command = cli({...})
  download.js        + export function buildDownloadExtractJs(noteId)
                       (CDN allowlist now includes rednote alongside xhscdn)
                     + export const command = cli({...})
  user.js            + export const USER_SNAPSHOT_JS
                     + export const command = cli({...})
  feed.js            + export function buildFeedPipeline(webHost)
                     + export const command = cli({...})
  notifications.js   + export function buildNotificationsPipeline(webHost)
                     + export const command = cli({...})
  note-helpers.js    buildNoteUrl now accepts `cookieRoot` + `signedUrlHint`
                     options (defaults preserved so xhs callers and tests
                     are unchanged)
  user-helpers.js    buildXhsNoteUrl / extractXhsUserNotes accept an
                     optional `webHost` argument (default xhs)

The `export const command = cli({...})` pattern matches twitter/lists.js
and clis/discord-app/*; without it the build-manifest scanner attributes
xhs's command to whichever rednote sibling triggered the transitive
import first.

## clis/rednote/ — thin shims

Each rednote command file imports the relevant builder / constant from
its xiaohongshu sibling and calls `cli()` with the rednote host triple.
No selectors, regexes, or extraction logic are duplicated.

  search.js          imports buildSearchExtractJs + noteIdToDate
                     declares its own WAIT_FOR_CONTENT_JS (modal + text
                     login-gate variants — the one xhs behaviour that
                     genuinely differs)
  note.js            imports NOTE_EXTRACT_JS + buildNoteUrl + parseNoteId
  comments.js        imports buildCommentsExtractJs + parseCommentLimit
                     + buildNoteUrl + parseNoteId
  download.js        imports buildDownloadExtractJs + buildNoteUrl + parseNoteId
  user.js            imports USER_SNAPSHOT_JS + extractXhsUserNotes
                     + normalizeXhsUserId

## Scope (initial)

Ships the five commands verified live against the user's logged-in
rednote.com session: search / note / comments / user / download.

`feed` and `notifications` are intentionally left out. Both rely on
intercepting the xiaohongshu Pinia store at the `homefeed` / `you`
capture pattern; live verification on rednote returns `tap → dict
(error)` for the feed step, so shipping them would surface a broken
contract. The mainland xiaohongshu commands continue to work. Adding
the rednote-side feed / notifications is straightforward follow-up
work once someone with rednote access maps the network surface.

Creator-center commands (publish, creator-*) have no rednote
counterpart and stay xiaohongshu-only, per the reporter's note in #1136.

## Verification

  - clis/xiaohongshu/ + clis/rednote/: 103/103 tests green
  - npx tsc --noEmit: clean
  - npm run build: 807 manifest entries (xhs 13 + rednote 5 + everything
    else preserved)
  - silent-column-drop / typed-error-lint: 103 / 189 baseline entries,
    no new violations
  - Live verify against the user's rednote.com session:
      rednote search "travel" --limit 1 → real note row
      rednote note <signed-url>         → 7 field/value rows
      rednote comments <signed-url> --limit 3 → 3 top-level rows
      rednote user 5b21f6564eacab3b38f05c39 --limit 2 → 2 profile notes
    Spaced 15–30s between runs per the xhs/rednote rate-limit guidance;
    no write commands invoked. Regression check: xiaohongshu/feed on
    the existing mainland session still returns the standard 6-field
    rows after the refactor.

Closes #1136

* fix(rednote): tighten adapter failure boundaries

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-05-12 02:32:25 +08:00
E2ern1ty b262d8ffd5 feat(chatgpt): support local image uploads (#1476)
* feat(chatgpt): support local image uploads

* chore: refresh cli manifest

* fix(chatgpt): harden image upload flow

* fix(chatgpt): validate image uploads before navigation

* fix(chatgpt): keep send fallback click in sync

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-05-11 17:52:16 +08:00
jakevin 467fdd0b62 refactor(adapter): rename site browser reuse to persistent sessions (#1462) 2026-05-11 04:56:34 +08:00
jakevin 9c06e84c89 refactor(browser): replace workspaces with sessions (#1461) 2026-05-11 04:26:51 +08:00
Benjamin Liu 674f0e1105 feat(openreview): add author command for ID-explicit publication lookup (#1365)
* feat(openreview): add author command for ID-explicit publication lookup

Closes the missing leaf in the openreview adapter. Among the public-strategy
academic adapters, dblp and arxiv both already ship an `author` command for
ID-explicit publication lookup; openreview only had `search` (full-text),
`paper` (detail by note id), `reviews` (thread by forum id) and `venue`
(listing by invitation / venue text). There was no way to ask "give me every
submission this author put on OpenReview, newest first."

`openreview author <profile>`:
  - takes a canonical profile id (`~First_LastN`); validated by
    `requireProfileId` so a dblp PID or a bare name fails before any
    network call,
  - hits `/notes?content.authorids=~<id>&limit=<n>&sort=cdate:desc`,
  - returns rank-ordered rows with the same shape as `openreview search`
    (id / title / authors / venue / pdate / url),
  - throws `EmptyResultError` when the profile has no public submissions
    instead of returning an empty list,
  - inherits the typed-error envelope from `openreviewFetch` so network
    failure, non-200, malformed JSON, and in-band error envelopes all
    surface as `CommandExecutionError`.

Tests: 6 new `it` blocks plus 1 updated registration test in
`clis/openreview/openreview.test.js`.

  - `requireProfileId` (1 block, 9 assertions): accepts canonical
    `~First_LastN`, `~Bo_Liu17`, and a multi-segment middle-name id;
    rejects empty, whitespace, missing tilde, missing trailing number,
    embedded space, and a dblp-style PID.
  - 5 author runtime cases covering pre-network ArgumentError, empty
    result, non-200, fetch network error, and the happy path with a
    request-shape assertion (`content.authorids` filter + `cdate:desc`
    sort).
  - Registration test extended to expect five commands and lock the new
    `columns` contract.

Manifest auto-regenerated to register the new command.

Live-verified end to end against `~Yoshua_Bengio1`: the most recent ICLR
2026 workshop submissions return with the expected fields. A malformed
profile is rejected before any HTTP call. A nonexistent profile yields
`EMPTY_RESULT`.

* fix(openreview): accept real profile id slugs

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-05-11 01:53:46 +08:00
jakevin 644d45177b feat(twitter): add unlike + retweet + unretweet + quote (write-action symmetry P0) (#1400)
Round 21 P0 — Twitter write-action symmetry (4 of 4: unlike, retweet, unretweet, quote).

## Scope
Closes write-action gap with existing siblings (`like`, `bookmark`, `unbookmark`, `delete`):
- `unlike` (UI strategy, navigateBefore:true)
- `retweet` (UI strategy)
- `unretweet` (UI strategy)
- `quote` (UI strategy, `/compose/post?url=` route — same family as `reply.js` `/compose/post?in_reply_to=`)

+745/-0 in initial commit, plus 3 progressive review fixes. Final: 4 adapters + 4 tests; modified `shared.js`, `shared.test.js`, manifest, docs.

## Iteration history (4 heads, 102/102 tests on final)

- `07836783` — initial 4 adapters + 4 tests, 96/96
- `55a89776` — fix #1: shared `parseTweetUrl()` URL invariant + quote post-submit verify (102/102)
- `dc9eab66` — fix #2: article-scoping for unlike/retweet/unretweet (delete.js sibling pattern)
- `8809d2c1` — fix #3: exact status-id matching (`match?.[1] === tweetId`) + quote-card exact id guard

## 4 progressive blockers caught (codex-mini0 lead + F-P-0 aux)

1. **URL validation (silent-clamp class)**: original passed any host containing `/status/<id>`. Fixed: `parseTweetUrl()` requires `https` + Twitter/X exact host + exact `/<user|i>/status/<id>` path; host-suffix, embedded URL, path-suffix all `ArgumentError` pre-nav.

2. **Quote silent-success illusion**: original click-implies-success without composer/toast verify. Fixed: pre-submit quoted-card exact id render assertion + post-submit success toast OR composer-clear assertion, otherwise return failed row.

3. **Broad querySelector scoping (delete.js sibling pattern)**: original state probe + click + post-click verify on conversation pages picked first matching button. Fixed: scope to `article` containing requested exact status id (sibling `clis/twitter/delete.js:22-23` pattern).

4. **Substring vs exact status-id matching**: `/status/123` substring-matched `/status/1234`. Fixed: regex `/\/status\/${id}(?:\/|$)/` segment-edge anchor + `match?.[1] === tweetId` exact compare.

## Cultural sediment (Round 21)

**Audit checklist 5 rules (pre-write upstream selection net)**:
1. cross-grep sibling URL-construction patterns before adopting
2. silent-clamp class detection (any normalize-then-trust path)
3. broad querySelector → article-scoping requirement
4. missing-validation early reject before navigation/IO
5. ID-based DOM/URL matching exact-not-substring

**Augment framing**: Round 21 audit-first 是 Round 18 字面量 self-check 的 **upstream pre-write 阶段**, 两者作用阶段不同, 共存比替换稳。

**Meta-anchor "Structural exactness for identity matching"** unifying:
- URL layer (#1391 isFacebookAuthRedirectPath: `\.php` + `(/|$)` segment edge)
- URL parser layer (#1392 parseGrokSessionId: bare UUID exact / URL host-exact-or-subdomain + path-exact)
- DOM layer (#1400 article-scoping: status-id `/\/status\/${id}(?:\/|$)/` regex or pathname segment-array exact compare)

Common invariant: boundary-lock structural shape, 不 trust substring 模糊 — fuzzy match 是 silent failure 温床。

## Validation gates (final head `8809d2c1`)

Local: Twitter tests 102/102, `node --check` touched files, `npx tsc --noEmit`, `npm run build`, typed-error-lint 189/189, silent-column-drop 103/103, doc-coverage 140/140, docs:build clean, listing-id advisory unchanged 13, `git diff --check` clean, merge-tree clean.

GitHub: build×3 (ubuntu/macos/windows) SUCCESS, unit-test×2 shards SUCCESS, bun-test SUCCESS, adapter-test SUCCESS, audit SUCCESS, doc-coverage SUCCESS, docs-build SUCCESS, smoke-test skipped, PR `CLEAN/MERGEABLE`.

## Strategy/UI boundary (better-solution verdict)

UI write path acceptable for P0 symmetry (matches existing Twitter write siblings). GraphQL write migration + structured `idempotent:true` flag are cross-sibling upgrades, P5 candidate, not P0 blockers.

Round 17 race-mitigation 第 8 连续 race-free execution (this round absorbed author scope-uncertainty hold-then-retract event without producing actual race).

Reviewers:
- Lead: @codex-mini0 (4-round iteration, all blockers caught)
- Aux: @First-principles-0 (better-solution triangulation, scope-discipline verdict, regression invariants)
- Author: @opencli-user
2026-05-08 01:10:54 +08:00
jakevin bf914f20f1 fix(grok): replace sentinel rows + silent-clamp with typed errors, deliver image cmd (#1397)
fix(grok): replace sentinel rows and deliver image command
2026-05-07 22:45:45 +08:00
jakevin 3b585fb4d1 feat(grok): add browser chat baseline commands (read/history/detail/new/send/status) (#1392)
Phase 3 — Grok adapter baseline (LLM browser-chat command family, parallel to ChatGPT/Qwen/Yuanbao).

## Surface
6 commands: `status` / `history` / `read` / `detail` / `new` / `send`. Site-local `clis/grok/utils.js` justified by 6 commands sharing helpers, not over-abstraction.

## 4-head review iteration

1. **`b4e81bad`** — initial baseline (12 Grok/shared files)
2. **`0a8112fc`** — mechanical rebase (CHANGELOG conflict only, all 12 Grok files preserved business-equivalent through rebase)
3. **`481e87e2`** — security fix: `parseGrokSessionId()` SSRF-shape vulnerability close — switched from regex string match to `new URL()` parser with branch separation:
   - Bare UUID mode: only exact UUID shape (no URL/query suffix accepted)
   - URL mode: requires `https` scheme + exact `grok.com` or subdomain host + exact `/c/<uuid>` path
4. **`a082023c`** — test-only hardening: 2 additional negative anchors covering existing implementation rejections (bare UUID `?next=abc` query tail / `grok.com.evil.com` host-suffix trick)

## Negative anchor coverage (8 cases)
http / off-domain / fakegrok / host-suffix subdomain / embedded URL / path suffix / UUID-tail / bare query tail

## Better-solution evidence form
LLM browser-chat family pattern (matching ChatGPT/Qwen/Yuanbao baseline) + 5 live probes — not first-site hostile scrape. TipTap editor API send seam (`editor.commands.focus/clearContent/insertContent`) is correct boundary because Grok ignores DOM input events; isolated in `sendMessage()`. Lack of full TipTap mock = residual risk, not blocker.

## Invariants locked
- `parseGrokSessionId()` URL parser branch separation (bare UUID exact / URL exact path)
- `history --limit` rejects invalid/out-of-range
- `status` uses `null` for unknowns (no fabrication)
- Bubble extraction preserves image-only assistant turns (no silent HTML-only drop)
- Auth/empty semantics aligned with LLM browser-chat baseline family

## Verification
Local: Grok adapter tests `28/28`, typecheck, build/manifest, docs:build, typed-error-lint `189/189`, silent-column-drop `103/103`, doc coverage `140/140`, listing-id advisory `13` unchanged, diff-check clean.
GitHub: build ubuntu/macos/windows × unit-test 1/2 + 2/2, bun-test, adapter-test, audit, doc-coverage, docs-build all SUCCESS. PR CLEAN/MERGEABLE.

Lead: codex-mini1. Aux: First-principles-1. Coordination: pr-monitor.
2026-05-07 21:11:39 +08:00
jakevin b2ebe211d1 feat(yuanbao): add browser-web baseline commands (status/read/detail/history/send) (#1394)
Wire up the standard browser-LLM command surface for Yuanbao, matching the
recently shipped chatgpt + claude + qwen baselines:
- status — login + current model + (agentId, convId) + URL
- read   — render the visible conversation as User/Assistant rows
- detail — open `<agentId>/<convId>` and read its messages
- history — list sidebar conversations with stable IDs
- send   — fire-and-forget, returns once the send button has been clicked

Refactor `ask.js` to share helpers (`sendYuanbaoMessage`, `normalizeBooleanFlag`)
with the new commands via `shared.js`, keeping the public ask behavior intact.

Notable bits:
- `parseYuanbaoSessionId` accepts only full chat URLs or `<agentId>/<convId>`
  pairs — Yuanbao chat URLs encode both, and silently opening the wrong agent
  on a bare UUID is a worse failure mode than throwing. URL regex anchored
  with `(?:[/?#]|$)` so 37+ char tails reject rather than truncate.
- `sendYuanbaoMessage` polls the send button (up to 3s) for the React
  re-render that drops `style__send-btn--disabled___*` after composer input —
  a fixed wait raced the debounce and produced silent no-op clicks.
- `getYuanbaoMessageBubbles` uses `data-conv-id`/`data-conv-idx`/
  `data-conv-speaker` attributes for stable per-turn identity (was relying
  on innerHTML alone).
- Status surfaces both human label (`Yuanbao`) and `dt-model-id`
  (`hunyuan_gpt_175B_0404`) — sentinel strings would silently look like a
  real model name; null is the typed-unknown signal.

Verified: 25 unit tests pass; targeted live smoke for status/read/detail/
history/new/send + ask round-trip on yuanbao.tencent.com.
2026-05-07 20:38:17 +08:00
jakevin b9b87a5c64 refactor(facebook/notifications): pipeline→func + typed errors + 7-col contract + runtime upfront limit (Phase 3 P5, #1391)
First Facebook adapter — Pattern C HTML scrape (lead 5 + author 4 = 7 endpoint family probe matrix dual-source negative evidence: graphql×3 / m.facebook redirect / login.php / checkpoint.php / fetch-patch / Messenger relay / ajax legacy 全 unauth 不可达, DOM walk over rendered notification rows + path-anchored auth detection 是当前 reviewable boundary).

Caller-visible delta: 3 cols (index/text/time) → 7 cols (+unread/+url/+notif_id/+notif_type).

[Bug fix] — 5 silent failures resolved
- silent-bad-shape: text.substring(0,150) → full body via per-row 'Mark as read' aria-label
- silent-bad-shape: time || '-' sentinel → string|null typed unknown
- silent-column-drop: unread badge / anchor href / notif_id / notif_t 暴露
- silent-empty-row: /login(.php)? + /checkpoint(.php)? redirect 返 [] → AuthRequiredError; empty/no-recoverable-text → EmptyResultError
- silent-clamp: limit 越界 silent clamp → ArgumentError (1-100), upfront before any navigation (navigateBefore: false)

[Structural refactor]
- pipeline → cli() func form + Strategy.COOKIE + navigateBefore: false (runtime upfront invariant 与 #1387 standard 拉齐)
- module-level pure exports: normalizeNotificationsLimit, stripMarkAsReadPrefix, stripAnchorChrome, parseNotifQuery, extractNotificationRowsFromDoc, isFacebookAuthRedirectPath, buildNotificationsScript
- Live IIFE 通过 \${fn.toString()} 嵌入 (dianping #1313 / hupu #1387 / xiaoe #1388 lineage)
- Locale 表 6 prefix / 4 badge label 显式列出
- AUTH_REQUIRED: sentinel → Node-side AuthRequiredError mapper

[Typed-error hardening]
- Path-anchored auth helper: isFacebookAuthRedirectPath(/^\/(?:login|checkpoint)(?:\.php)?(?:\/|\$)/i) — domain-invariant-first encoding (FB top-level auth-only invariant), 排除 /loginhelp /help/login /account/login/identify
- Three-layer navigateBefore=false invariant lock: registration assertion + manifest absence + executeCommand runtime page.goto-zero-call (test layer 与 invariant layer 完整对齐)
- Row-level silent-empty-row defense: anchor rows with no recoverable body text 直接 skip, 不 emit text:null success row

[Doc fix]
- docs/adapters/browser/facebook.md notifications enrichment + Output table (列类型 / null vs sentinel 语义) + auth/empty error contract
- Boy Scout audit: cross-checked profile / feed / search / marketplace-listings / marketplace-inbox 例 commands 与 args 定义一致

Tests
- notifications.test.js 39/39 + src/execution.test.ts 21/21
- Anti-pattern regression guards: not.toMatch(/text\.substring\(0,\s*150\)/) + not.toMatch(/time\s*\|\|/)
- JSDOM frozen-fixture (slim 13 lines, 0 blank): header listitem skip / full text / unread badge / query parsing / null time / blank-row skip / relative href absolute / 19-case auth path matrix
- typed-error-lint baseline 192 → 191 (silent-sentinel resolved 1)

Review iterations (4 head, A 组 codex-mini0 lead + First-principles-0 aux):
1. 052d2b18 (initial 29 tests) → 376cb50f (lead gate fix: Ubuntu lint + auth path-segment + anchor.href + 5 typed-error func tests)
2. 376cb50f0d6c1340 (pr-monitor grep cross-verify catch /login.php false-negative; lead 加 \\.php 边界)
3. 0d6c13403e5a5ff0 (opencli-user 19-case 实测 + lead 抽 named helper isFacebookAuthRedirectPath domain-invariant-first encoding + 2 row-shape silent-failure 顺手 catch)
4. 3e5a5ff036e44f73 (F-P-0 aux blocker: registry-injected navigateBefore 在 limit validation 之前 fire pre-nav 违反 #1387 upfront boundary; navigateBefore:false + 三层断言 registration/manifest/runtime executeCommand)

Closes #1391
2026-05-07 17:49:57 +08:00
jakevin 381f095706 feat(qwen): add detail command + fix stale message bubble selector (#1390)
* feat(qwen): add detail command + fix stale message bubble selector

`getMessageBubbles` was matching `[data-msgid="<id>-question|answer"]` from an
older Qianwen frontend. The reshipped DOM no longer carries that attribute on
chat turns; `[data-message-id]` now lives on citation cards inside assistant
responses, so the old selector silently returned an empty list and `qwen read`
had been silently broken.

Rewire to walk `[data-chat-question-wrap]` and `[data-chat-answers-wrap]` in
DOM order (correct Q/A interleaving) and synthesize stable IDs from the
nearest sibling `data-req-id` so `waitForAnswer.seenAssistantId` and
read/ask/detail dedupe paths keep working. Verified live against an existing
conversation: 3 user turns + 3 assistant turns extracted; old selector
returned 0.

`qwen detail <id|url>`: open a specific conversation by ID or full chat URL,
poll up to 20s for the transcript to render, return Role/Text rows. Adds
`parseQianwenSessionId` (5 unit tests covering ID/URL parsing + ArgumentError
on malformed input). Reuses the same site-level browser session as `read`/
`ask` so consecutive calls continue in the same Qwen tab.

- clis/qwen/detail.js (new)
- clis/qwen/utils.js (parseQianwenSessionId + getMessageBubbles rewire)
- clis/qwen/utils.test.js (new)
- docs/adapters/browser/qwen.md (detail entry + options/columns)
- cli-manifest.json (regenerated)

* fix(qwen): anchor URL regex to reject 33+ hex tail truncation

codex-coder review on PR #1390 caught that
`https://www.qianwen.com/chat/<33+ hex>` would silently truncate to the
first 32 chars and open the wrong conversation. Adds end-of-input /
slash / query / fragment boundary to the URL match group and two new
unit-test cases (digit tail + letters tail) covering the truncation gap.
2026-05-07 17:39:23 +08:00
jakevin 99986c3101 feat(chatgpt): add browser chat baseline commands
Add ChatGPT web ask/send/read/history/detail/new/status alongside existing image support. Tighten ChatGPT web helper selectors and typed error contracts, update docs/changelog, regenerate manifest, and seed local ChatGPT verify fixtures for ask/read.
2026-05-07 17:37:18 +08:00
jakevin 6f7eb6a76a refactor(xiaoe x3): pipeline→func + typed errors + content silent-drop fix (Phase 3 P1)
Phase 3 P1 (xiaoe catalog/courses/content) — pipeline→func refactor + typed-error hardening + content silent-drop bug fix + URL upfront validation + inherited legacy doc fix。

## Tags (PR body honesty 演进 dual-nature framing 试用)

- **[Bug fix]** `xiaoe/content` silent-column-drop (caller-visible delta)
- **[Structural refactor]** `xiaoe/catalog` + `xiaoe/courses` pipeline→func 包壳 (parity by construction, IIFE 字节级保留)
- **[Typed-error hardening]** 三 func `page.goto` + `page.evaluate` failure 包成 `CommandExecutionError`; `content/catalog` URL upfront `ArgumentError` (missing/malformed/non-https/off-domain) before navigation
- **[Doc fix]** `docs/adapters/browser/xiaoe.md` `courses --limit 10` (legacy doc 错误 inherit) + `--url` wording → 实际 positional `url` (manifest aligned)

## Per-tag detail

### [Bug fix] content silent-column-drop (real caller-visible bug)
adapter 名"提取小鹅通图文页面内容为文本", IIFE 返 `{title, content, content_length, image_count, images}`, 但 columns 只声明 `[title, content_length, image_count]` → `content` (那段文本本身) 被 silent drop。**用户拿到 "1234 chars" 但拿不到那 1234 chars** — adapter 名字撒谎了。
- Fix: 公开列 `[title, content, content_length, image_count]`, `content` 真 caller-visible delta
- Choice A (vs B reshape): legacy `images` 是 `JSON.stringify(slice(0, 20))` 截断/stringified 坏合同, **不暴露成新列** (避免把 silent-bad-shape 升级成公开坏合同), 留 follow-up 另开 explicit media/images contract
- `image_count` 用 `countXiaoeImages(doc)` 全页计数, 不 slice (既有 metadata 质量修正)

### [Structural refactor] catalog + courses pipeline→func wrapper (parity by construction)
- `pipeline:[]` form → `func` form
- IIFE body 字节级保留 (Xiaoe 没 public REST, Vue 私有 runtime 是唯一稳定 hook, JSDOM 复刻不了 Vue tree)
- Pure helpers extracted: `pickContentText`, `countXiaoeImages` (content) / `typeLabel`, `buildItemUrl`, `chapterUrlPath` (catalog) / `buildCourseUrl` (courses)
- IIFE 通过 `\${fn.toString()}` 嵌同一份代码 (dianping #1313 / hupu #1387 同模式)
- No live verify acceptable: IIFE 字节级保留 + helper 全 unit-test + manifest column shape 不变 = 行为 parity by construction
- `buildScript` 反向断言 `images.slice(0, 20)` legacy anti-pattern 不出现 (anti-pattern regression guard, 同 #1387 `documentElement.outerHTML` 反向 guard)

### [Typed-error hardening] 三 func navigation + evaluate boundary
- `requireXiaoePageUrl()` for `content/catalog`: missing/malformed/non-https/off-domain URL → upfront `ArgumentError` before `page.goto` (test asserts `expect(page.goto).not.toHaveBeenCalled()`)
- `content/catalog/courses`: `page.goto` moved inside try, navigation/evaluate failures both wrap as `CommandExecutionError`, no raw CDP/browser error path leaks
- Empty shell stays `EmptyResultError` (no reliable login-wall signal to justify `AuthRequiredError`, 避免 false positive — 应用 #1384 secUid 教训)

### [Doc fix] inherited legacy doc errors
- `xiaoe courses --limit 10` example removed (no `--limit` arg in manifest, legacy doc 错误 inherit)
- positional `url` wording aligned with manifest (was incorrectly `--url`)
- 同 #1386 positional docs 教训, 但延伸到 "继承 legacy doc 错误也是新 PR 责任" (Boy Scout typed-error hardening 在 doc 层延伸)

## Tests: 46/46 green
- 3 cmd registration contract
- pure helper unit tests (selector chain / image filter / URL priority / type label fallback / no synthetic URL)
- `buildScript` invariants (`images.slice(0, 20)` 反向断言)
- wire tests: ArgumentError upfront (BEFORE page.goto), EmptyResultError empty rows + empty content, CommandExecutionError navigation/evaluate failure, rows verbatim happy path

## Lint gates
- typed-error-lint 190/190 (no new) ✓
- silent-column-drop 103/103 (no new) ✓ (注: `pipeline:[]` IIFE string template AST walker 看不进, lint follow-up scope)
- doc-coverage 140/140 ✓
- listing-id-pairing advisory unchanged 13 ✓

## GitHub checks (head a6d37d70)
build ×3 / unit-test ×2 / bun-test / adapter-test / audit / doc-coverage / docs-build SUCCESS, smoke skipped, MERGEABLE / CLEAN

## Review
B 组: @codex-mini1 lead + @First-principles-1 aux, double-green confirmed, Round 17 race-mitigation 第 4 轮 protocol clean closeout (第 4 次连续无 race 执行: #1384 / #1386 / #1387 / #1388)。

## Sediment lessons
- Silent-failure 三类 taxonomy: silent-column-drop (列没声明) / silent-bad-shape (字段在但 shape 错) / silent-empty-row (错误状态返空行而不是抛 typed error) — 三类 fix 路径不同, blast radius 不同
- PR body honesty 演进 4 链: #1384 R4 race disclosure → #1386 positional docs 教训 → #1388 silent-failure 三类分开写 + dual-nature tag 矩阵
- F-P-1 first-principles call: 不顺手暴露 legacy 坏合同 (silent-bad-shape ≠ silent-drop, fix 路径完全不同)
2026-05-07 17:02:23 +08:00
jakevin e610260705 refactor(hupu/hot): pipeline→func + querySelectorAll + 4 enrichment columns (Phase 3 P3)
Phase 3 P3 (hupu/hot) — pipeline→func refactor + 2 真 bug 修 + 4 列 enrichment + JSDOM-frozen-fixture test pattern (#1313 复用) + anti-pattern regression guard。

## Summary
- Pipeline form (`pipeline:[]` + `documentElement.outerHTML` regex) → `func` form (`querySelectorAll('.t-info')` DOM walk)
- **Bug 1 修**: outerHTML regex 静默漏行 (markup 抖动就漏, mocked test 抓不到)
- **Bug 2 修**: regex 抓所有 9-digit 锚点 → ~70 个 anchor 但页面只 render 60 个 `.t-info` row → legacy adapter 每次返 ~10 个 phantom 行 (导航链接 conflated 成 thread 行)
- **4 enrichment columns** (4→8): `lights` (亮 count int|null, 万 expanded `1.2万→12000`) / `replies` (回复 count int|null) / `forum` (per-row sub-section) / `is_hot` (bool 暴露 hupu \" hot\" marker, 不 filter 行序保持页面顺序)
- columns/manifest/docs sync: `[rank, tid, title, lights, replies, forum, is_hot, url]`,`null` vs `0` 语义清楚

## Typed errors
- `--limit` 上游 `ArgumentError` for 0/-1/>100/1.5/non-numeric (BEFORE `page.goto`,**不 silent clamp**)
- 空页 `EmptyResultError`
- `page.evaluate` failure 包成 `CommandExecutionError` (test regression locked)

## JSDOM frozen-fixture test pattern (#1313 复用)
- 抽 `extractHupuHotRowsFromDoc(doc, limit, parseCount)` 为 module-level pure export
- in-page IIFE 通过 `\${fn.toString()}` 嵌同一份代码
- JSDOM test 直接调 export against `__fixtures__/hot-home.html` (slim 6-row hand-crafted fixture)
- 17/17 tests green (contract / normalize / parseCount / extract / buildHotScript invariants / wiring / phantom-anchor exclusion / evaluate-error envelope)

## Anti-pattern regression guard (#1313 fixture pattern 延伸)
- `buildHotScript` 反向断言 `not.toContain('documentElement.outerHTML')` 锁不回退到旧 broad regex
- `buildHotScript` 反向断言 `not.toContain('regex.exec')` 同向锁
- fixture 顶部 `.t-info` 外的 9-digit phantom anchor `639999999` 反向锁: 旧 broad regex 会抓到, 新 `.t-info` extractor 不抓 — 把 fixture 反向验证从断言层升到证据层

## Better-solution check (live probe evidence-based)
DOM `.t-info` = 60 visible rows, `window.\$\$data.pageData.threads` = 70 (10 hidden/non-rendered)。对"首页可见 hot rows" 任务, DOM walk 比 bootstrap JSON 更贴 source of truth (后者会引入 hidden/不渲染条目)。这条 60 vs 70 数字是设计决策的硬 justify, 不是设计意见。

## Lint gates
- typed-error-lint 190/190 (no new) ✓
- silent-column-drop 103/103 (no new) ✓
- doc-coverage 140/140 ✓
- listing-id-pairing advisory unchanged 13 ✓

## GitHub checks (head 874d4e4e)
build ×3 / unit-test ×2 / bun-test / adapter-test / audit / doc-coverage / docs-build SUCCESS, smoke skipped, MERGEABLE / CLEAN

## Review
A 组: @codex-mini0 lead + @First-principles-0 aux, double-green confirmed, Round 17 race-mitigation 第 4 轮 protocol clean closeout.
2026-05-07 16:59:17 +08:00
jakevin 464de7059e refactor(tiktok): write commands -> button-walker Route 1 with typed errors (Phase 3 P0.5)
Phase 3 P0.5: refactor 3 TikTok write commands (comment, follow, unfollow) from time-window-wait UI flow to a button-walker + state-verification path with a typed-error boundary, sharing a parallel helper structure to the #1384 read PR.

Two-layer helper boundary (clis/tiktok/utils.js extension):
- BUTTON_WALKER_HELPERS (browser side): button-walker (locate / pre-click state read / click / state-verify post-click) + cleanText reuse + cookie/auth-secUid plumbing for write-auth + plain Error throws on contract violations
- throwButtonWalkerError() (Node side): map browser-thrown errors -> typed CommandExecutionError (button missing / state-verify fail / captcha / rate-limit / navigation/eval/empty-row defensive failures) / AuthRequiredError (cookie + viewer secUid) / ArgumentError (upfront input validation). Explicitly NO EmptyResultError mapping (button contract violation is not an empty result, per #1384 R4 lesson on auth-vs-empty classification).

Per command:
- comment <video-url> <text>: button-walker click + state-verify by checking comment-list state (not wait-2s)
- follow <username>: pre-click state read distinguishes idempotent fast path (`already-following` / `already-friends`) from post-click success (`followed`). Post-click result causality preserved (post-click never returns `already-*`).
- unfollow <username>: pre-click `already-not-following` fast path; post-click `unfollowed`.

result enums (per row):
- comment: `posted` (no idempotent path - comments cannot dedupe)
- follow: `followed` | `already-following` | `already-friends` (last two pre-click only)
- unfollow: `unfollowed` | `already-not-following` (last one pre-click only)

retryable contract (in hint string `retryable=<bool> reason=<...>`):
- comment failures: retryable=false reason=server-fan-out
- follow/unfollow failures: retryable=true reason=idempotent (server-side dedupe is safe)

Lead push iterations during review (codex-mini1 maintainer-fixes-directly):
- f5730f16: rate-limit/captcha -> CommandExecutionError + retryable hint BEFORE auth regex (auth precedence bug); follow post-click success -> `followed` (NOT `already-friends`, fixing causality misclassification); navigation/empty-row defensive failures route through throwButtonWalkerError (containing raw Error leakage).
- b683f46c: parseTikTokVideoUrl() requires canonical /@user/video/<numeric-id> with only optional trailing slash/query; malformed suffixes (e.g. /123abc, extra path) -> upfront ArgumentError.
- f5dc91d6: docs examples updated to actual positional args for write commands (was stale --url/--text/--username flag form), covering write-rewrite + sibling like/unlike/save/unsave on touched docs file (Boy Scout).

Intentionally NOT addressed (separate scope, candidate post-merge follow-ups):
- Direct /api/commit/follow/user/ or /api/comment/publish/ (would require X-Bogus signing reverse engineering, separate risk surface)
- RetryableError as core typed-error metadata (currently encoded in hint string, post-merge candidate to import into engine)
- TikTok Studio creator metrics commands (separate Phase scope)

Validation:
- clis/tiktok/ tests: 64/64 (38 read from #1384 + 22 new write contract + 4 regression for blockers caught during review)
- typed-error-lint: 190/190
- silent-column-drop: 103/103
- doc-coverage: 140/140
- listing-id advisory: 13 unchanged
- docs:build pass, manifest 764 entries
- GitHub gates on f5dc91d6: build x3 / unit x2 / bun / adapter-test / docs-build / doc-coverage / audit all SUCCESS, smoke skipped, CLEAN/MERGEABLE

Reviewers: codex-mini1 (lead, 3 contract pushes f5730f16 -> b683f46c -> f5dc91d6), First-principles-1 (aux, validated 4 contract patches + better-solution check confirming button-walker Route 1 vs /api/commit/* + X-Bogus separation).
2026-05-07 16:33:00 +08:00
jakevin 9a7dd44b3e refactor(tiktok): 6 read commands -> page-context API (Phase 3 P0, absorbs #1382)
Phase 3 P0: refactor 6 TikTok read commands (explore, following, friends, live, notifications, user) from DOM/network-intercept to TikTok web's own page-context API endpoints, sharing one helper boundary.

Helper boundary (clis/tiktok/utils.js):
- BROWSER_HELPERS: in-browser fetchJson + cleanText + asNumber (null/'' -> null preserve missing-vs-zero distinction) + cookie/msToken plumbing
- VIDEO_ITEM_NORMALIZER: normalize page-context item -> row shape
- assertTikTokApiSuccess(data, label): unify TikTok in-band envelope (status_code/statusCode != 0; code 8 or auth-looking message -> AUTH_REQUIRED; other -> upstream label API failed)
- throwTikTokPageContextError() (Node side): map browser-thrown errors -> AuthRequiredError / EmptyResultError / CommandExecutionError

Per command:
- explore: /api/recommend/item_list/ pagination, --limit upfront ArgumentError
- following: /api/user/list/ relationships
- friends: /api/user/list/ + cross-filter
- live: /api/live/discover/ feed
- notifications: /api/notice/multi/ (status 8 -> AUTH_REQUIRED)
- user (absorbed from #1382): secUid resolve via __UNIVERSAL_DATA_FOR_REHYDRATION__ -> /api/user/detail/, /api/post/item_list/ pagination, /api/search/general/full/ exact-author fallback. !secUid -> EmptyResultError (NOT AuthRequiredError; auth still covered by HTTP 401/403 + envelope status_code 8/auth-looking msg). source field = bootstrap | profile-api | search-fallback in row/columns/manifest/docs/tests.

Closes #1382 (absorbed; #1382 closed without separate merge per WAWQAQ direction).

Validation:
- clis/tiktok/ tests: 38/38
- typed-error-lint: 190/190
- silent-column-drop: 103/103
- doc-coverage: 140/140
- docs:build pass, manifest no drift
- GitHub gates: build x3 / unit x2 / bun / adapter-test / audit / doc-coverage / docs-build all SUCCESS, smoke skipped, MERGEABLE

Reviewers: codex-mini0 (lead, push 4 boundary fixes 18cdf930 -> a1f1ada4 -> 53499609 -> 276dce3b), First-principles-0 (aux, caught secUid auth-vs-empty boundary + verified 6 cmd integral helper boundary).
2026-05-07 16:15:30 +08:00
jakevin b327da5b3c feat(llm): reuse browser sessions by site (#1385) 2026-05-07 15:49:25 +08:00
jakevin fa7851bb9a feat(browser): add adapter session reuse (#1383) 2026-05-07 15:24:54 +08:00
jakevin c6d5da54ee feat(web): add exhaustive same-origin frame mode (#1373) 2026-05-07 01:15:10 +08:00
jakevin 67cde0e263 enrich(coupang): product detail cmd + replace silent clamp/sentinel/Error with typed errors (#1370)
* enrich(coupang): add product detail cmd + replace silent clamp/sentinel/Error with typed errors

Two enrichment changes plus three silent-failure fixes on top of existing
search / add-to-cart.

New cmd: coupang product
─────────────────────────
Pairs with search as the listing↔detail round-trip target. Reads a logged-in
product page and extracts a single canonical row with price, original_price,
discount_rate, rating, review_count, seller, brand, rocket, delivery_promise,
image_url, url. Three-source extractor (JSON-LD Product schema → bootstrap
globals → DOM) merged in priority order, mirroring the search.js pattern.

The columns use string|null typing — null means "upstream did not provide
this field on this product" (e.g. some items have no original_price).
Failures (login wall / page mismatch / page failed to render) raise typed
errors instead of silently returning empty rows, so callers can treat any
returned row as real data.

Search column shape: added product_id
─────────────────────────────────────
Listing must pair with detail by id. The data was already extracted by
normalizeSearchItem; only the columns array needed updating so the field
projects through to the rendered row. Per the listing-id-pairing convention
(PR #1297) the new column lets agents round-trip rows directly into
`coupang product` without re-scraping URLs.

Silent-failure fixes
────────────────────
1. search --limit silent clamp.
   Old: `Math.min(Math.max(Number(kwargs.limit||20),1),50)` silently
        rewrote `--limit 999` to 50 and `--limit 0` to 1.
   New: `parseLimitArg(raw, 20, 50)` throws ArgumentError on out-of-range
        / non-integer / negative input. Same convention as the typed-fail-fast
        memory & PR #1289.

2. search --page silent clamp.
   Old: `Math.max(Number(kwargs.page||1),1)` silently lifted negative pages.
   New: parsePageArg throws ArgumentError on non-positive input.

3. Generic `throw new Error(...)` → typed errors.
   - Empty query, unsupported --filter, missing --product-id/--url
     → ArgumentError
   - Login wall detection → AuthRequiredError('coupang.com', ...)
   - Empty result / filter-not-rendered → EmptyResultError
   - PRODUCT_MISMATCH / OPTION_REQUIRED / button-not-found / unknown
     ack failure (add-to-cart) → CommandExecutionError
   - The PRODUCT_MISMATCH and `actualProductId || 'unknown'` sentinel were
     also fixed (silent-sentinel was the audit hit there).

Coverage
────────
- 21 contract assertions in clis/coupang/coupang.test.js covering
  parseLimitArg / parsePageArg (no silent clamp), registry shape (search has
  product_id, product is read-class with expected columns, add-to-cart is
  write-class), and typed-error pre-flight rejections (empty query / bad
  filter / out-of-range limit & page / missing detail args).
- Manifest 763 → 764 (+1 entry: coupang/product).
- Audits: typed-error-lint 196 → 194 (resolved 2 silent-clamp/sentinel
  baseline entries; baseline updated). silent-column-drop 103/103 unchanged.

* fix(coupang): tighten product id and browser errors

* fix(coupang): require real product urls
2026-05-07 00:18:57 +08:00
jakevin a5a3248a77 refactor(linux-do): remove deprecated hot/category/latest compat shims (#1368)
* refactor(linux-do): remove deprecated hot/category/latest compat shims

The three shims have been pure backward-compat wrappers since linux-do/feed
became the unified entrypoint. With no stable release commitment to preserve,
they are pure surface cost: 3 manifest entries, 3 deprecated branches in help
output, and a `buildLinuxDoCompatFooter` helper that exists only to feed them.

- delete clis/linux-do/{hot,category,latest}.js
- drop now-orphaned `buildLinuxDoCompatFooter` from feed.js and unexport
  `executeLinuxDoFeed` (no external consumers remain)
- remove the Compatibility section in docs/adapters/browser/linux-do.md
- regenerate cli-manifest.json (-125 lines)

BREAKING CHANGE: `opencli linux-do hot|category|latest` are removed. Use
`opencli linux-do feed --view top --period <period>`,
`opencli linux-do feed --category <id-or-name>`, and
`opencli linux-do feed --view latest` instead.

* fix(linux-do): finish compat shim removal
2026-05-06 23:57:15 +08:00
jakevin 4ef2cb8b1c enrich(toutiao): hot board (public) + bug fixes (silent column drop, partial render) (#1366)
* enrich(toutiao): hot board + bug fixes (silent column drop, partial render)

Per WAWQAQ "丰富现有 adapter" pivot — Phase 2 site #3.

## New command
- `toutiao hot` (Strategy.PUBLIC, browser:false) — public homepage hot
  board via the toutiao.com hot-event/hot-board endpoint. No login required.
  Returns 8 stable columns (rank/id/title/query/hot_value/label/url/image).

## Bug fixes for `toutiao articles`
- **Silent column drop fixed**: `parseToutiaoArticlesText` previously
  did `if (title && stats) push(...)`, silently dropping any row where
  the stats span hadn't finished rendering by the time page.innerText
  was read. Slow-render bugs were invisible — adapter looked "complete"
  while writers saw extra rows in the dashboard. Partial rows now
  surface with `null` stat columns.
- **Silent clamp on `--page` removed**: out-of-range / non-integer
  values raise `ArgumentError` with explicit bounds [1, 4]. Same
  validation reused by both `articles` and `hot` via `parseArticlesPage`
  / `parseHotLimit` in `utils.js`.
- **Empty result typed**: zero-row scrape now raises `EmptyResultError`
  instead of returning `[]` silently (would otherwise look like a
  legitimate "no articles" response).

## Refactor
- Parser logic extracted to `clis/toutiao/utils.js` (alongside hot-row
  mapping, validators, and the hot-board URL constant).
- `articles.js` switches from declarative `pipeline:` to imperative
  `func` form so `parseArticlesPage` validation can run before the
  navigation step (declarative pipeline can't pre-validate args).
- Strategy is now explicit: `Strategy.COOKIE, browser: true` for
  articles (creator dashboard is logged-in only).

## hot field map
`ClusterIdStr` (or numeric `ClusterId`) → id; `Title` → title;
`QueryWord` → query (falls back to title); `HotValue` → hot_value
(non-negative numeric, else null); `Label`, `Url`, `Image` →
respective columns. `pickImage` walks `Image.url` → first truthy
`Image.url_list[]`. Empty-title rows are dropped (returns null) before
ranks are densely re-assigned 1..N.

## Tests
29 contract assertions across `parseArticlesPage` / `parseHotLimit` /
`parseToutiaoArticlesText` / `mapHotRow` + registry-level shape checks
+ `hot` adapter func behaviour (typed errors / no silent clamp / fetch
failure paths / dense-rank).

## Audits
- typed-error-lint: 196 = 196 (unchanged baseline)
- silent-column-drop: 103 = 103 (unchanged baseline)
- listing-id-pairing: hot has `id` column (round-trippable when a
  detail command lands later); advisory list unchanged.

## Manifest
757 → 758 entries (+1 for `hot`).

## Doc
- index.md: toutiao mode 🔐🌐/🔐 (hot is public, articles is logged-in)
- toutiao.md: per-command mode/domain table + column docs + prerequisites

* fix(toutiao): tighten hot and articles contracts
2026-05-06 23:17:32 +08:00
jakevin 69ee36f997 fix(linkedin): surface detail_error on --details (no silent catch / no silent empty) (#1363)
* fix(linkedin): surface detail_error on --details (no silent catch / no silent empty)

The previous --details enrichment path had two indistinguishable failure modes
that both produced `description: '', apply_url: ''`:

1. `if (!job.url)` early return — row had no jobId, so we couldn't navigate.
2. `} catch {}` — page.goto / page.evaluate threw (network, timeout, parse error).

Callers couldn't tell "upstream had no description" from "we failed to fetch",
and the catch swallowed every error without logging. For an enrichment that
costs one page navigation per row, silent failure is especially harmful — users
just see an empty cell with no way to debug.

Fix: replace empty strings with `null` for missing/failed rows, add a new
`detail_error` column (string|null) carrying a short typed reason:

  - 'no url'                — row had no jobId
  - 'fetch failed: <msg>'   — page.goto / page.evaluate threw
  - 'missing description'   — page loaded but body was empty
  - null                    — success

Every failure is also logged to stderr with the offending URL so debugging is
possible. Per-row failures still don't abort the batch (the original intent),
but they're now visible.

Tests: 13 new contract assertions in clis/linkedin/search.test.js covering
parseCsvArg, mapFilterValues (ArgumentError on unknown values), decodeLinkedinRedirect,
and 5 enrichJobDetails paths (no-url / goto-throw / empty-description / success /
multi-row-mixed). Added `export const __test__` for testability.

Audits clean: typed-error-lint 196/196, silent-column-drop 103/103.

* fix(linkedin): fail fast on auth walls
2026-05-06 22:57:09 +08:00
jakevin da2453cfbd enrich(reuters): article-detail + bug fixes (silent clamp, silent error envelope) (#1362)
* enrich(reuters): article-detail + bug fixes (silent clamp, silent error envelope)

Per WAWQAQ "丰富现有 adapter" pivot — Phase 2 site #2.

- `reuters article-detail` — full article body + canonical metadata for a
  Reuters URL. Pairs with `reuters search` (use the `url` column to
  round-trip).

- **Silent clamp on `--limit` removed**: out-of-range values now raise
  `ArgumentError`. Validation happens before browser navigation.
- **Silent error envelope removed**: the in-page IIFE used to swallow
  `fetch` errors with `catch(e) {}` and return `{error: ...}`, then the node
  side did `if (!Array.isArray(data)) return [];`. Now:
  - in-page IIFE returns `{ ok, status, body, error? }` raw envelope
  - node side throws typed errors:
    - `CommandExecutionError` on in-page exception
    - `CliError(FETCH_ERROR)` on non-2xx upstream
    - `CommandExecutionError` on captcha HTML (200 + non-JSON body)
    - `EmptyResultError` on empty articles array
- **Empty query**: now `ARGUMENT_INVALID` instead of triggering an empty
  upstream call.
- **Column shape enriched**: previously dropped `section_path` and
  `authors` are now stable columns.

- `docs/adapters/browser/reuters.md`: full Commands / Columns / Error
  Behaviour section (was a 3-line stub).
- `docs/adapters/index.md`: add `article-detail` to the commands cell.

27 contract assertions across `parseLimit` / `mapSearchArticles` /
`mapArticleDetail` / `buildSearchScript` / `buildArticleDetailScript`
+ registry-level checks for both commands (Strategy, ARG validation
before nav, every typed-error path, success path).

- typed-error-lint: 196 → 195 (silent-clamp resolved on
  `clis/reuters/search.js:18`); baseline updated.
- silent-column-drop: 103 = 103 (unchanged).
- listing-id-pairing: advisory only (article-detail keys off `url`).

757 → 758 entries (+1 for `article-detail`).

* fix(reuters): type auth and fetch failures

* fix(reuters): preserve search detail round trip
2026-05-06 20:00:36 +08:00
jakevin 61c4637b4c enrich(ctrip): hotel-suggest + bug fixes (silent clamp, dropped columns, fake URL) (#1361)
* enrich(ctrip): hotel-suggest + bug fixes (silent clamp, dropped columns, fake URL)

Per WAWQAQ "丰富现有 adapter" pivot — Phase 2 site #1.

## New command
- `ctrip hotel-suggest` — surfaces hotel-context suggestions (cities,
  business areas, individual hotels) via the same backing endpoint with
  searchType=H. Distinct from `ctrip search` (searchType=D) which returns
  destinations / scenic spots / railway stations.

## Bug fixes for `ctrip search`
- **Silent clamp on `--limit` removed**: out-of-range values (≤0, ≥51,
  non-integer) now raise `ArgumentError` with explicit bounds rather than
  silently snapping to [1, 50].
- **Silent column drop fixed**: previously the adapter discarded `id`,
  `cityId`, `cityName`, `provinceName`, `countryName`, `lat`, `lon`, `eName`
  and `displayType` from upstream rows. Now all are surfaced as stable
  columns.
- **Fake URL fixed**: previously `url` was always `''`. Now constructs
  canonical Ctrip URLs by `type` (City / Markland / Hotel / Zone / RailwayStation)
  and returns `null` (no silent fabrication) for unknown types.
- **In-band error envelope typed**: `Result: false` payloads now surface
  as `COMMAND_EXEC` (was previously not handled — adapter returned empty
  rows).

## Doc fix
- `Mode: 🔐 Browser` → `🌐 Public` (search uses public API, no login)
- Add `hotel-suggest` to commands table in both `docs/adapters/index.md`
  and `docs/adapters/browser/ctrip.md`.

## Coords picker
Mainland China rows ship `gdLat`/`gdLon` (gaode); international rows ship
`gLat`/`gLon` (wgs84). Adapter picks the first non-zero pair (zero is the
upstream sentinel for "missing"); returns `null` if all variants are zero.

## Tests
25 contract assertions across `parseLimit` / `pickCoords` / `buildUrl` /
`mapSuggestRow` + registry-level checks for both commands (Strategy /
shape parity / typed errors / no silent clamp).

## Audits
- typed-error-lint: 196 = 196 (unchanged baseline)
- silent-column-drop: 103 = 103 (unchanged baseline)
- listing-id-pairing: advisory only (search has `id` round-trip column)

## Manifest
757 → 758 entries (+1 for `hotel-suggest`).

* fix(ctrip): wrap suggest fetch and json failures
2026-05-06 19:34:13 +08:00
Benjamin Liu 6469a02ea6 feat(deepseek): add detail and send commands for explicit conversation control (#1344)
* feat(deepseek): add detail and send commands for explicit conversation control

doubao already ships `detail <id>` and `send` for ID-explicit conversation
read/write; deepseek had only `read` (current page only) plus the
implicit-resume `ask`. Adding both gives users a stable handle when they
know the conversation ID, without going through `ask`'s resume detection
or its full prompt-then-wait pipeline.

`deepseek detail <id>`:
  - parses a bare UUID or any URL containing `/a/chat/s/<id>`,
  - rejects malformed input via `ArgumentError` before any browser
    navigation,
  - navigates to `https://chat.deepseek.com/a/chat/s/<id>` and returns
    the visible message list,
  - throws `EmptyResultError` when the conversation has no rendered
    messages.

`deepseek send <id> <prompt>`:
  - takes the conversation id as a required positional, because the
    framework runs each browser command in an ephemeral per-command
    workspace (a fresh tab) and there is no shared "current conversation"
    across commands; the navigation must be explicit,
  - drives input through CDP `Input.insertText` via `page.nativeType`,
    mirroring the doubao adapter (#1278); `execCommand('insertText')` plus
    a synthesised input event leaves the React-controlled state desynced
    on a freshly-opened tab and the resulting click silently no-ops,
  - keeps the verification loop inside the same `page.evaluate` so the
    framework cannot close the tab mid-flight; counts user-class bubbles
    by text-match (DeepSeek virtualises the message list, so a numeric
    bubble-count check is unreliable),
  - throws `CommandExecutionError` with a specific reason when the
    textarea did not populate, the send button stayed disabled, the
    bubble never settled, or the optimistic render rolled back during
    a 3s settle window,
  - treats "Promise was collected" from the post-click eval as success,
    matching the existing pattern in `ask --file`.

Helper `parseDeepSeekConversationId` is exported from utils.js so the
same parser feeds both commands and round-trips the canonical lower-case
ID.

Tests:
  - utils.test.js: 5 cases covering bare UUID, upper-case
    normalisation, URL extraction with and without query string, empty /
    null / whitespace input, and non-UUID rejection.
  - detail.test.js: 5 cases covering registration, navigation +
    message return, URL normalisation, ArgumentError before browser
    navigation, and EmptyResultError on no-messages.
  - send.test.js: 7 cases covering registration, ArgumentError on bad
    id, full happy-path through nativeType + IIFE verification, the
    textarea-mount timeout, missing nativeType helper, focus failure,
    IIFE-reason translation to CommandExecutionError, and the
    "Promise was collected" success path.

Manifest auto-regenerated to register both commands.

Live-verified end-to-end against my own DeepSeek session:
  - `detail` returns the canonical message list for a bare UUID, parses
    a full chat URL, and rejects malformed IDs before any browser
    navigation,
  - `send` lands the prompt as the latest user message in the target
    conversation and gets an AI response back; reload of the
    conversation page in a separate tab confirms the message persisted
    server-side.

* docs(deepseek): document detail and send commands

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-05-06 19:13:54 +08:00
jakevin f033481e67 feat: 2 read adapters (wttr, openfda) + contract tests (#1355)
* feat: 2 read adapters across 2 new sites + contract tests (wttr, openfda)

Trimmed from original Round 11 per WAWQAQ feedback (msg=3899a382): drop
novelty/niche sites (timeapi / zippopotam / spacedevs / citybik) — keep
only sites with clear real-world utility:

- wttr (current, forecast) — wttr.in weather, no auth, simple text/json toggle
- openfda (drug-label, food-recall) — FDA drug labels + food recall enforcement

13 contract tests across 2 sites cover Lucene operator query construction
(openfda +AND+ literal handling), [string] 1-elem array unwrap, brand-OR-
generic match, wttr [{value:"..."}] array-of-objects 1-elem unwrap.

Manifest 757→759 (+2). Audits clean: typed-error-lint=196 baseline.

* fix(openfda): use brand or generic label search
2026-05-06 17:47:05 +08:00
jakevin 39943c05e9 feat: 12 read adapters across 6 new sites + contract tests (round 7) (#1350)
* feat: 12 read adapters across 6 new sites + contract tests (round 7)

New sites: wikidata, lichess, rest-countries, nuget, flathub, oeis
- wikidata: search (wbsearchentities) + entity (Special:EntityData) — Q/P/L ids,
  localised label/description with English fallback
- lichess: user + top (perf rankings) — closed accounts → EmptyResultError, no
  silent disabled rows; 13 perf types validated
- rest-countries: country (substring) + region — population-sorted by default,
  flattened languages/currencies/capitals
- nuget: search + package (full version history) — registration page-walking
  for 100+ version histories; case-insensitive id with strict shape gate
- flathub: search + app — appId reverse-DNS, dual-shape timestamp coercion
  (search is unix-seconds int, /appstream is ISO string)
- oeis: search (paginated) + sequence — A-id zero-padded, 12-term preview with
  (+N) suffix, surfaces commentCount/formulaCount/etc instead of full graphs

Contract tests: 6 files × 7 assertions = 42 contract assertions, all green.
Live verified all 12 commands against real APIs.

Audits clean: typed-error 196=baseline, silent-column 103=baseline, 0 Round-7
listing-id-pairing violations (all 6 listings carry round-trip ids).

* fix(nuget): fail fast on malformed registration pages
2026-05-06 16:54:21 +08:00
jakevin 9ae44c228a fix(round5): strengthen contract tests (#1348) 2026-05-06 14:19:55 +08:00
jakevin 498ad3930c feat: 13 read adapters across 6 new sites (round 4) (#1347)
Six new public-API sites — package registries + Docker images + OpenAlex
scholarly works — all unauthenticated, no browser required.

  dockerhub  search image
  rubygems   search gem
  homebrew   formula cask popular
  packagist  search package
  maven      search artifact
  openalex   search work

Conventions held:
  - access: 'read' on every command
  - typed errors (ArgumentError / EmptyResultError / CommandExecutionError)
    instead of generic CliError or silent fallback
  - input validators per site (image slugs, gem names, Composer names,
    Maven coordinates, OpenAlex work-id / DOI normalization)
  - listing rows carry an id-shaped column (image / gem / token / package /
    coordinate / id) that round-trips into the corresponding detail command
  - HTTP 429 surfaces with retry hint, 404 → EmptyResultError

Audits:
  - check:typed-error-lint   → no new violations (baseline 196)
  - check:silent-column-drop → no new violations (baseline 103)
  - advise:listing-id-pairing → unchanged at 13
2026-05-06 13:38:43 +08:00
jakevin 55088bbb28 feat: 13 read adapters across 5 new sites + 4 extensions (round 3) (#1346)
New sites (8 commands):
- npm    : search / package / downloads (registry.npmjs.org + api.npmjs.org)
- pypi   : package / downloads (pypi.org + pypistats.org)
- crates : search / crate (crates.io)
- mdn    : search (developer.mozilla.org)
- nvd    : cve (services.nvd.nist.gov)

Extensions (5 commands; +1 dblp/author surfaced in index):
- hf            : spaces (Hugging Face Spaces by likes / created_at / last_modified)
- dblp          : venue (search dblp's venue registry by acronym/topic)
- coingecko     : derivatives (perpetual / futures markets, 24h volume)
- stackoverflow : related (related questions for a given question id)

All commands hit public unauthenticated endpoints (Strategy.PUBLIC, browser:false),
typed-fail-fast on bad inputs (no silent fallback / clamp), and round-trip listing
ids into their detail commands where applicable.

Audits (all green vs baseline):
- typed-error-lint        : 196 = 196 baseline, no new
- silent-column-drop      : 103 = 103 baseline, no new
- listing-id-pairing      : 13 advisory (was 12; +1 = dblp/venue with no
                            corresponding venue-detail command)

Doc coverage : 120/120 adapter dirs documented (+5 new doc pages, +4 updated)
Manifest     : 722 entries (was 709; +13 commands)

Live verified:
- npm search react / npm package react / npm downloads react --period last-week
- npm downloads react --period 2025-01-01:2025-01-05
- pypi package requests / pypi downloads requests --period recent / overall
- crates search tokio / crates crate serde
- mdn search fetch
- nvd cve CVE-2021-44228
- hf spaces --limit 3
- dblp venue ICLR
- coingecko derivatives --limit 3
- stackoverflow related 79935770 --limit 3
- typed-error sanity: invalid CVE id, bad npm name, bad --period
2026-05-06 13:14:41 +08:00
jakevin a78ceb1602 feat: 11 read adapters across 8 sites (round 2) (#1345)
* feat: 11 read adapters across 8 sites (dblp / steam / bbc / devto / lobsters / medium / coingecko / hf)

Round 2 of the adapter expansion sweep. All 11 commands hit public APIs (no
browser, no auth), follow the post-#1332 typed-error / no-silent-failure
discipline, and were live-verified against real endpoints.

New adapters:
- dblp/author      : recent publications for one author (resolve PID by name, or pass --pid)
- steam/search     : storefront name search (storesearch API)
- steam/app        : single app detail (appdetails API; HTML entities decoded)
- bbc/topic        : per-topic RSS (8 canonical BBC News feeds)
- devto/latest     : /api/articles/latest with --page pagination
- lobsters/domain  : stories from a specific source domain (/domains/<d>.json)
- medium/tag       : tag RSS (description full-length, no silent truncation)
- coingecko/exchanges  : trust score + 24h BTC volume leaderboard
- coingecko/categories : sector buckets with 6 sort options
- coingecko/global     : aggregate market totals + BTC/ETH dominance
- hf/paper         : single-paper detail by arXiv id (summary, ai_summary, ai_keywords, upvotes)

Also adds clis/steam/utils.js + clis/bbc/utils.js as shared helpers (HTML entity
decode, RSS parsing). All listings carry a round-trippable id where a detail
sibling exists; advise:listing-id-pairing reports zero new violations. typed-
error-lint and silent-column-drop gates both unchanged from baseline.

Manifest: 698 → 709 (+11 entries).

* fix: tighten adapter round2 contracts
2026-05-06 12:46:14 +08:00
jakevin 6f597a2a4b feat: 8 read adapters across 5 sites (arxiv / SO / coingecko / wikipedia / hf) (#1338)
* feat: add 13 read adapters across 6 sites (github / arxiv / SO / coingecko / wikipedia / hf)

New site:
- github: user, repo, search-repos, user-repos, releases (unauth REST API; 60 req/h IP limit)

Existing sites — gap-fill for high-traffic verticals:
- arxiv author (papers by author, newest first; au:"name" phrase match on the public Atom API)
- stackoverflow user / tag (Stack Exchange API 2.3, with HTML-entity decode for display names / titles)
- coingecko coin / trending (single-coin market detail; 24h trending search-volume)
- wikipedia page (full plain-text article extract; opt-in --paragraphs cap, no silent truncation)
- hf models / datasets (downloads/likes/trending/freshness sorted lists)

All adapters use Node-side func + typed errors per the post-#1332 convention:
- ArgumentError for invalid limit / bad enum / empty positional / malformed owner-repo
- EmptyResultError for genuinely-empty results (no silent return [])
- CommandExecutionError for upstream HTTP/JSON failures (rate limit / 5xx / parse)
- AuthRequiredError reserved for endpoints that genuinely refuse anonymous traffic
- No silent clamp on --limit; no sentinel rows; no scalar 'unknown' / '-' fallbacks

Audit gates locally green:
- check:typed-error-lint        196/196 (no new)
- check:silent-column-drop      103/103 (no new)
- check:doc-coverage --strict   113/113 (added github.md, extended 5 existing pages)
- advise:listing-id-pairing     advisory only (+2 wikipedia entries: title is the
                                round-trippable key into wikipedia/page; not a gate)

* chore: drop github adapter set per WAWQAQ directive

WAWQAQ (#opencli-pr-review): "我们不需要GitHub的adapter,因为已经有GH了"

Removes the 5 github commands + utils + docs added in 664ed1aa
(github/user, github/repo, github/search-repos, github/releases,
github/user-repos). The remaining 8 read commands across 5 sites
(arxiv author, stackoverflow user/tag, coingecko coin/trending,
wikipedia page, hf models/datasets) are unaffected.

Audit gates re-checked:
- check:typed-error-lint: 196/196 (baseline unchanged)
- check:silent-column-drop: 103/103 (baseline unchanged)
- doc-coverage: 112/112 (one less site documented)
- advise:listing-id-pairing: 12 advisory (unchanged)

* fix(adapter-expansion): tighten id and currency contracts
2026-05-06 02:24:16 +08:00
SnakeEye-sudo (Er. Sangam Krishna) 9c9c8f976d feat: add uisdc and aibase news adapters (closes #1201) (#1249)
* Add uisdc news adapter for CLI

Implements a CLI adapter for fetching the latest AI/design news from uisdc.com. Allows specifying the number of news items to return.

* feat(aibase): add aibase daily news adapter

This file implements a news adapter for AIbase that fetches the latest AI industry news and allows for configurable limits on the number of news items returned.

* fix(news): harden uisdc and aibase adapters

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-05-06 02:14:12 +08:00
Greatkai d0b1b6a89e feat(pubmed): revive public eutils adapter (#819)
Co-authored-by: jackwener <jakevingoo@gmail.com>
Co-authored-by: Greatkai <4587517+Greatkai@users.noreply.github.com>
2026-05-06 01:54:29 +08:00
Luke 1794933d06 feat: add tiktok creator-videos command (#1335)
* feat: add tiktok creator-videos command

TikTok Studio creator content list with views/likes/comments/saves/shares.

Hits the Studio item_list endpoint
(https://www.tiktok.com/tiktok/creator/manage/item_list/v1/?aid=1988) from a
logged-in /tiktokstudio/content session and pages with cursor until limit is
satisfied (server caps size at 50). Username for the resulting video URL is
extracted from the user_text= query param on play_addr / download_info entries,
falling back to scraping a[href*="/video/<id>"] from the Studio page DOM.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tiktok): regen manifest + replace silent-clamp with ArgumentError

- Regenerate cli-manifest.json (CI gate: must match `npm run build` output)
- Replace `Math.max(1, Number(args.limit) || 20)` and
  `Math.min(Math.max(limit, 1), 50)` with an explicit positive-integer
  guard + a server-cap-only ternary, per the silent-clamp guidance in
  references/typed-errors.md (typed-error-lint baseline is unchanged)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tiktok): tighten creator videos contract

---------

Co-authored-by: root <root@example.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-05-06 00:58:38 +08:00
Carson dbf1f6afa1 feat(weixin): add Sogou article search (#1250)
- 新增 `clis/weixin/search.js`:通过 Sogou 微信搜索做公众号文章发现,`access: read`
- typed fail-fast: bad query/page/limit upfront ArgumentError;captcha/频控/goto/wait/evaluate/unreadable payload/selector drift/partial card extraction → CommandExecutionError;no-result 页面 → EmptyResultError;--limit 不 silent-clamp >10 直接拒绝
- maintainer-fixes-directly 闭环:rebase 到 latest main `f4637486`、补 access:read、删除 silent-clamp pattern
- F-P-1 first-principles 评估:Sogou 是 public 搜索页 ≠ 微信官方 API,可接受边界 = fail-fast + 清晰字段契约,不做字段猜测/partial success
- CI 全绿(smoke-test SKIPPED);B 组 codex-mini1 lead green + First-principles-1 aux green on `31c2b035`
- 非阻塞残留:search.url 是 Sogou redirect link,串联 download 需后续单独支持 redirect resolution 或暴露 resolved mp URL

Round 15a B 组对位收口。
2026-05-06 00:52:05 +08:00
Jack He b485e96d90 feat(xianyu): add publish command for listing items (#1282)
- 新增 `clis/xianyu/publish.js`:发布闲鱼商品(标题/描述/分类/价格/图片/condition),`access: write`
- 参数 upfront `ArgumentError`:空 title/description/category、非法 price/original_price、未知 condition、图片格式/数量/文件不存在
- DOM/UI fail-fast:表单缺失、分类选择失败、必填字段未填、file input 缺失/上传失败、submit 失败、发布失败或超时未确认 → `CommandExecutionError`;登录墙 → `AuthRequiredError`
- 删除 `status=failed` success-row anti-pattern:失败/未知发布结果不再作为 success 表格返回
- F-P-1 aux catch real-runtime blocker:`page.url()` 在 IPage/BasePage 无定义,test mock `url` 字段遮住;改 `page.getCurrentUrl()` + fallback publish URL,加 IPage-shape 回归锁住
- xianyu publish JSDOM 回归 + docs/README/index 同步
- B 组 codex-mini1 lead green + First-principles-1 aux green on `2d78144d`

Round 14 (B 组对位)。
2026-05-05 23:35:23 +08:00
jakevin a5d70466ba feat: add qwen / 1point3acres / coingecko adapters (#1329)
3 new sites / 17 adapters,+2471 LOC。

**adapters**:
- coingecko (PUBLIC): `top` 全球加密货币市值排行
- 1point3acres (Discuz, GBK/UTF-8 mixed): `digest/forum/hot/latest/search/notifications/thread/user`
- qwen (browser, COOKIE chat): `ask/send/image/history/status`

**typed fail-fast 全闭环 (A 组三轮迭代后)**:
- silent-column-drop heuristic key collision 修法:rename intermediate keys 避开 columns 名字
- 所有外部参数 (limit/page/contentLimit/timeout/page_size) 越界/非法 → `ArgumentError`
- fetch / non-2xx / malformed JSON / API error → `CommandExecutionError`
- empty / not-found → `EmptyResultError`
- qwen prompt 缺失 → `ArgumentError` (不是 CommandExecutionError)
- qwen/status 未知 model/session 用 typed `null` (不是 `'-'` sentinel)
- success-row 永远不塞 failure/empty 业务行
- `normalizeLimit(value, default, max, label)` 共享 helper for 1point3acres 5 adapters

Author: @opencli-user (jackwener)
A 组 review (三轮):
  - codex-mini0 lead: 第三轮 maintainer-fixes-directly 直接 push `c40daf7` 收掉 4 类深一层 contract 漏洞
  - First-principles-0 aux: 第二轮 catch typed-error-lint 9 条 + 第三轮 catch generic CliError / silent-clamp on page/timeout / success-row failure / qwen prompt class 4 类 hard blocker
2026-05-05 18:00:33 +08:00
jakevin 0f806e9473 feat(dianping): browser adapter — search + shop on www.dianping.com (#1309)
* feat(dianping): browser adapter — search + shop on www.dianping.com

Adds two browser-mode adapters for the dianping (大众点评) PC site:

- `dianping search "<keyword>" --city <name|id> --limit <n>`: keyword
  shop/restaurant search. Returns rank, shop_id, name, rating, reviews,
  price, cuisine, district, url. shop_id round-trips into `dianping shop`.
- `dianping shop <shop_id>` (alias `detail`): shop detail sheet
  (field/value rows: name, rating, breakdown 口味/环境/服务/食材, reviews,
  price, rank, hours, address, subway, features, url).

Both use Strategy.COOKIE on www.dianping.com (the PC site renders search
SSR and does not require JS hydration). m.dianping.com is intentionally
crippled for non-mobile UAs, so it's not used.

Auth detection (utils.detectAuthOrEmpty) inspects both response text and
final URL for the Meituan Yoda captcha redirect (verify.meituan.com) and
the dianping login redirect; raises AuthRequiredError with the captcha
URL embedded so the user can clear it manually in the same profile.

Listing↔detail id pairing: search.shop_id → shop.<id>. Adds 'shop' to
DETAIL_NAMES in scripts/check-listing-id-pairing.mjs so the convention
gate scans this site (35 sites / 78 listings now covered).

* fix(dianping): harden browser failure classification

* fix(dianping): fail on partial missing shop ids
2026-05-04 23:00:38 +08:00
jakevin 2b9af38db4 feat(pixiv): surface user_id + url on listings, url on user/illusts (#1300)
While auditing instagram/facebook/pixiv coverage gaps, found that pixiv
listings already extract `user_id` and construct `url` per row but drop
both fields from the table view (`columns` doesn't list them). The data
is in the row object — only the column projection was missing.

Per the listing↔detail id pairing convention (#1297), surface them so:
- `user_id` round-trips from `ranking` / `search` → `user` / `illusts`
- `url` is the canonical share link for every illust / user record

Changes:
- `ranking`: + user_id, + url
- `search`:  + user_id, + url
- `illusts`: + url (user_id is the arg, no need to repeat per row)
- `user`:    + url

No behavior change beyond the table view — JSON output already had these
fields, so existing scripts that consume `-f json` keep working.
2026-05-04 21:09:01 +08:00
jakevin 545f91a2d2 feat(dblp): public bibliography adapter — search + paper (#1299)
* feat(dblp): public bibliography adapter — search + paper

Wraps the dblp.org public API:
- `dblp search <query>` → /search/publ/api JSON, projected into one row per hit
- `dblp paper <key>` → /rec/<key>.xml, parsed into a one-row record

Why dblp on top of arxiv/openreview: dblp is the largest, oldest CS
bibliography (7M+ entries) and the only one of the three that consistently
indexes pre-arXiv literature, journal articles, books, and theses. The
canonical record key (e.g. `conf/nips/VaswaniSPUJGKP17`) round-trips
cleanly between the two commands per the listing↔detail convention.

Implementation notes:
- No deps beyond the registry — XML parsed with conservative regexes,
  same approach as the arxiv adapter.
- Polite User-Agent per dblp's API guidance; HTTP 429 mapped to a
  CommandExecutionError with a "lower --limit" hint.
- Author homonym suffixes (`"Smith 0001"`) trimmed for clean output.
- 39 unit tests cover validators, XML extraction, both commands.

* fix(dblp): fail fast on API status envelopes
2026-05-04 21:04:50 +08:00
jakevin 0a85e73aa5 feat(convention): listing↔detail id pairing rule + CI gate (#1297)
* feat(convention): listing↔detail id pairing rule + CI gate

Adds a hard convention: when a site exposes both a listing-class command
(search / hot / top / recent / ...) and a detail-class command (read /
paper / article / view / ...), every listing row MUST surface an id-shaped
column whose value round-trips into the detail command. Without that, an
agent has no way to follow up on a listing row except re-searching by
title or scraping URLs out of band — both of which break the agent-native
contract.

What's in this PR

- docs/conventions/listing-detail-id-pairing.md — full rule, examples
  table, why-it-matters, what counts as id-shaped, exemption taxonomy,
  how to add an id column to a listing.
- scripts/check-listing-id-pairing.mjs — validator that reads
  cli-manifest.json, classifies each entry as listing / detail / other,
  and fails when a listing on a site that also has a read-detail command
  is missing an id-shaped column. Exemption allowlist records WHY each
  pair is exempt so future maintainers know what to verify.
- npm run check:listing-id-pairing — strict-mode wrapper.
- CI: new step in build job runs the validator after the manifest
  freshness check on Linux.
- docs/developer/ts-adapter.md — cross-link from the adapter authoring
  guide.
- docs/.vitepress/config.mts — sidebar entries for the new conventions
  section.

Fixes brought to zero violations

- 1688/search: add offer_id (already extracted, just surfaced)
- bluesky/user: add uri (AT URI round-trips into bluesky/thread)
- tieba/search: add id + url (thread_id already extracted)
- tieba/hot: add url (rows are topics, not threads — url is the
  best-effort round-trip handle, doc'd as such)

Exemptions (intentional, doc'd in EXEMPT map with rationale)

- nowcoder/hot, bluesky/trending, twitter/trending — listing rows are
  topic strings, not posts.
- lesswrong/user, reddit/user — rows are profile-attribute key/value
  pairs, addressed by the username arg.
- discord-app/search — desktop UI session, message ids not extractable.
- notion/search — Strategy.UI Quick Find, page ids not exposed in DOM.

Validator output after this PR: 32 sites scanned, 75 listings checked,
7 exempted, 0 violations.

* fix(convention): tighten listing id gate

* fix(convention): close url-derived id loophole
2026-05-04 20:54:14 +08:00
jakevin 29b4869efd feat(indeed): add search and job adapters (US site) (#1298)
* feat(indeed): add `search` and `job` adapters (US site)

Adds an Indeed adapter that fills the US job-search gap (alongside
existing 51job / boss-zhipin / linkedin coverage). Both commands run
through a real browser session because Indeed sits behind Cloudflare
and answers bare HTTP fetches with `403` + `cf-mitigated: challenge`.

## Commands

- `indeed search <query>` — keyword job search
  - args: `query`, `--location`, `--fromage`, `--sort`, `--start`, `--limit`
  - columns: `rank, id, title, company, location, salary, tags, url`
- `indeed job <jk>` (alias `detail`, `view`) — full job posting
  - args: `id` (positional, the 16-char hex `jk` from `search`)
  - columns: `id, title, company, location, salary, job_type, description, url`

## Listing↔detail id pairing

`search.id` is the Indeed `jk` (job key, 16-char lowercase hex). It feeds
directly into `indeed job <jk>`. Conforms to the listing↔detail id
pairing convention proposed in #1297.

## CF challenge handling

The adapter polls the result selectors for up to 15s after navigation,
giving the browser time to clear the Cloudflare interstitial. If the
challenge is still up after the wait, the adapter throws a
`CommandExecutionError` with a hint pointing the user at the connected
browser to clear it once. Subsequent calls reuse the warmed cookies via
`Strategy.COOKIE`, mirroring the v2ex / boss / linkedin patterns.

## Validation

`utils.js` keeps argument validation pure and unit-testable:

- `requireJobKey` rejects anything that isn't a 16-char lowercase hex
- `requireFromage` only accepts `1` / `3` / `7` / `14` (Indeed's enum)
- `requireSort` only accepts `relevance` / `date`
- `requireBoundedInt(limit, default=15, max=25)` — Indeed serves at most
  one page (10 jobs/page); ArgumentError on out-of-range, no silent
  clamping, per the typed-error feedback in #1289.

## Tests

18 unit tests in `clis/indeed/indeed.test.js` cover registration,
validators, URL builders, and DOM-card normalizers. Browser-driven
verification stays out of CI by design (CF challenge is interactive).

## Docs

- `docs/adapters/browser/indeed.md` — full adapter doc with prerequisite
  CF-challenge notes and listing↔detail id pairing callout.
- Sidebar entry + adapter index row.

* fix(indeed): tighten timeout fail-fast and runtime tests

* fix(indeed): align readiness with search parser
2026-05-04 20:52:16 +08:00
jakevin 328140966e feat(openreview): public adapter — search/venue/paper/reviews (#1294)
* feat(openreview): add public adapter — search/venue/paper/reviews

OpenReview is the open peer-review platform used by ICLR / TMLR / COLM
and ML workshops. Its v2 API exposes everyone-readable submissions,
reviews, and decisions without auth, so all four commands run with
`browser: false`.

Commands:
- `openreview search <query>` — full-text search
- `openreview venue <venue>` — list submissions; accepts either a venue
  display name (matched against `content.venue`, e.g. "ICLR 2024 oral")
  or a full invitation id (e.g. "ICLR.cc/2025/Conference/-/Submission")
  via `/-/` heuristic; supports offset pagination
- `openreview paper <id>` — single-paper detail with full abstract
- `openreview reviews <forum>` — paper + threaded reviews/decisions/
  comments, ordered chronologically with paper lifted to row 0;
  classifies notes via invitation tail (REVIEW / DECISION / REBUTTAL /
  COMMENT / META_REVIEW / WITHDRAWAL); per-row truncation via
  `--max-length` (min 200)

Listing IDs round-trip into `paper`/`reviews`. PDF URLs normalized to
absolute `https://openreview.net/pdf/...`. `pdate` falls back to
`cdate` when missing, formatted as `YYYY-MM-DD`.

All limits/offsets/ids fail-fast with typed errors (`ArgumentError`,
`EmptyResultError`, `CommandExecutionError`) — no silent clamping, no
empty-array fallbacks. fetch + json + non-2xx + 404 are wrapped so
network/API failures never look like empty results.

Tests: 23 unit tests covering the column contract, content extraction,
date/PDF normalization, invitation-vs-venue dispatch, error paths
(network/JSON/HTTP), pagination offset accounting, and the reviews
classifier + section joiner + truncation.

Live-verified against api2.openreview.net for search ("diffusion
model"), venue ("ICLR 2024 oral"), paper (KS8mIvetg2), and reviews on
that paper's full thread.

* fix(openreview): tighten error and review typing

* fix(openreview): stabilize review contracts
2026-05-04 19:23:55 +08:00
jakevin ed0b2acc82 docs(stackoverflow): clarify read fetches answers up to --answers-limit (not 'all') (#1295)
Follow-up from PR #1293 review: 'all answers' was misleading because
the implementation is limit-bounded (default 10, max 100) rather than
unbounded pagination. Spell out the actual contract — including the
accepted-answer-outside-page fallback path — so users don't expect
infinite-scroll behaviour.

Non-blocking docs-only change flagged by codex-mini1 + First-principles-1
during #1293 review.
2026-05-04 19:20:43 +08:00
jakevin c1a4bd3b7e feat(stackoverflow): surface question_id on listings + new read <id> (#1293)
* feat(stackoverflow): surface question_id + metadata on listings, add `read <id>`

Agent-native gap: all 4 stackoverflow listings (`hot`, `search`,
`unanswered`, `bounties`) only emitted `[title, score, answers, url]`,
which means an agent could see a hot question but had no `id` to round-
trip into a body read, no `tags` to filter by topic, no `views` to gauge
demand, and no `is_answered` / `creation_date` / `author` to triage.
There also wasn't a `read` adapter, so reading a SO question through
opencli was impossible.

Listings (`hot` / `search` / `bounties` / `unanswered`):
- Add `rank`, `id` (question_id), `views`, `is_answered` (skipped on
  `unanswered` since always false), `tags` (joined), `author`
  (owner.display_name), `creation_date` columns.
- Pass `pagesize` to the upstream API instead of fetching the default
  page and trimming locally.

New `stackoverflow read <id>`:
- 4-call fan-out against the public Stack Exchange API
  (`/questions/{id}` + `/questions/{id}/comments` +
  `/questions/{id}/answers` + batched `/answers/a;b;c/comments`).
- Returns `POST` + `Q-COMMENT` + `ANSWER` + `A-COMMENT` rows mirroring
  the `hackernews read` and `lobsters read` shape.
- Accepted answer is always surfaced first and tagged `accepted='true'`;
  remaining answers follow in descending vote order, capped by
  `--answers-limit`.
- HTML body cleanup: tags stripped, `<pre><code>` preserved, `<code>`
  inline-fenced, `<li>` rendered as `- `, comments indented with `> `.
- Entity decoding: a shared `decodeEntities` handles named (incl.
  `&hellip;`/`&copy;`/etc), decimal (`&#246;`), and hex (`&#x27;`)
  forms, applied to both bodies AND `display_name` (otherwise users
  like `Jonas K&#246;lker` come through mojibaked).
- Typed fail-fast: `ArgumentError` for non-numeric id and
  `--max-length < 100` (with no-fetch assertion); `EmptyResultError`
  when `items` is empty; `CommandExecutionError` for HTTP non-2xx and
  for Stack Exchange's in-band `error_id` envelopes (throttle / quota).
  No silent clamps anywhere.

Tests: 14 vitest assertions
- 4 listing column-shape (incl. `unanswered` skipping `is_answered` and
  `bounties` keeping its `bounty` column position)
- 10 read-adapter cases: registration / args / strategy + 3 typed-error
  fail-fast paths (with no-fetch assertion on the pre-fetch ones) + the
  full POST/Q-COMMENT/ANSWER/A-COMMENT row order with accepted-first +
  the answer-comments fetch verified to batch ids semicolon-joined +
  HTML entity decoding (named/decimal/hex) on both body and display_name
  + answers-limit honored when there are more answers than the cap.

Live verification:
- `stackoverflow hot --limit 2` → `id`/`tags`/`views`/`is_answered`/
  `author` populated.
- `stackoverflow search "async await" --limit 1`,
  `stackoverflow unanswered --limit 1` → same shape.
- `stackoverflow read 79935770` and the very-long classic question
  `stackoverflow read 11227809 --answers-limit 1 --comments-limit 2`
  → produces the threaded POST/Q-COMMENT/ANSWER/A-COMMENT structure
  with proper entity decoding (`Jonas Kölker` reads correctly).
- `stackoverflow read not-numeric` → exits with `ARGUMENT`.
- `stackoverflow read 999999999` → exits with `EMPTY_RESULT`.

* fix(stackoverflow): wrap fetch/json/coerce paths in typed errors

Apply the 3 lessons from PR #1292 (devto) review at merge time, before
B-group hits this PR:

1. CLI args may arrive as strings (e.g. `--max-length 50` → `'50'`).
   The bare `Number.isInteger(value)` in `requirePositiveInt` /
   `requireMinInt` would accept negative-but-coerced numbers and reject
   string-form integers. Now the helpers `coerceInt` first then validate,
   and the rejection message echoes the raw input via `JSON.stringify`.

2. `await fetch(url)` and `await res.json()` were not wrapped — a network
   blip would surface as a raw `TypeError` and a maintenance HTML page
   would surface as a raw `SyntaxError`. Both are now caught and rethrown
   as `CommandExecutionError` with hints, matching the in-band error_id
   path.

Tests: +3 cases (17 total)
- fetch network failure → CommandExecutionError
- malformed JSON body → CommandExecutionError
- string-form max-length "50" / "abc" rejected with ArgumentError before
  fetching

* fix(stackoverflow): avoid partial read fanout
2026-05-04 19:10:50 +08:00