mirror of
https://github.com/jackwener/OpenCLI.git
synced 2026-09-14 18:25:42 +08:00
docs/drop-for-developers
1254 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
913fb80aad |
docs(readme): drop For Developers section
Per WAWQAQ: from-source install instructions are infrastructure detail that don't belong in a public-facing README. Contributors finding themselves in this repo will already know `npm install / build / link` patterns; users who reach the README from npm don't need them. Removed in both EN and ZH. |
||
|
|
ce432c2428 |
chore(release): 1.8.0 (#1682)
* chore(release): 1.8.0 Substantial release: weread-official adapter, wider LinkedIn / Twitter / Reddit / Zhihu coverage, 12306 / Suno / Xianyu additions, security and reliability fixes, plus a 20% README shrink. * chore: remove orphan docs/adapters-doc/ones.md The file was a leftover from PR #386 (2026-04-10) and has been superseded by docs/adapters/browser/ones.md. Bundled into the 1.8.0 release commit chain so the release doesn't ship with a dead docs file alongside the new docs. Skipped from opus-reviewer's audit (Tier 1 #1-#4) for this release: - #1 smart-search dead refs (10 spots) — owned by @codex-coder's skill-deletion PR; release PR will rebase on top of it. - #3 clis/test-utils.js relocation — touches 19 importers, separate refactor PR. - #4 clis/slock/ orphan — needs WAWQAQ design call. - #6 opencli-usage:161 wording — current "Commands that used to exist" framing is already clear enough. - #7 docs/adapters/index.md sync (8-20 missing sites) — broader docs PR, not release-time bundling. * chore: remove clis/slock + sync docs/adapters/index.md (audit #4 + #7) Per WAWQAQ post-audit directive on #OpenCLI:f046ece7: - `clis/slock/` was a half-finished orphan with only `_utils.js` and no command entry points. Removed. - `docs/adapters/index.md` was missing 11 browser adapters: 12306, suno, weread-official, qwen, 1point3acres, brave, duckduckgo, cnki, flomo, jianyu, taobao. Added all with commands sourced from cli-manifest.json. Desktop section already covered all 7 desktop adapters (Cursor / Codex / Antigravity / ChatGPT App / ChatWise / Discord / Doubao App).v1.8.0 |
||
|
|
7ee16aa087 |
feat(booking): add search adapter for Booking.com hotel listings (#1680)
* feat(booking): add search adapter for Booking.com hotel listings
New `opencli booking search <destination> --checkin --checkout` adapter
scrapes the server-rendered hotel cards on www.booking.com via stable
`[data-testid=property-card]` selectors. No login required (Strategy.PUBLIC
+ browser:true).
Highlights
- 12 columns: rank, name, country, slug, star_rating, review_score,
review_count, price_amount, price_currency, distance, recommended_room,
url. `slug` + URL stay stable across locales (better round-trip key
than `name`, which Booking sometimes localizes from session cookies).
- Score parser anchors on `(\d{1,2})\.(\d)` so the duplicated "8.68.6" /
"评分8.68.6很棒" rendering doesn't mis-parse to 8.68.
- Currency symbol → ISO 4217 map (US$/€/£/¥/¥/₹/₩/HK$/A$/NT$/S$/CN¥);
honor `--currency` URL param for stable codes.
- Pagination via `--offset` (Booking pages 25/request); `rank` includes
the offset so paginated calls stay sortable.
- Captcha-page detection short-circuits to CommandExecutionError instead
of silent empty rows.
Typed errors (no silent clamp / fallback)
- Bad date / out-of-range adults/rooms/children/limit/offset / unknown lang
/ malformed currency → ArgumentError up front (before any navigation).
- Browser nav failure → CommandExecutionError.
- Zero cards rendered → EmptyResultError with a hint.
- Captcha page → CommandExecutionError.
29 unit tests cover the helpers, the registry shape, every typed-error
path, the {session,data} CDP envelope unwrap, and offset-aware rank
numbering. Silent-column-drop + typed-error-lint audits unchanged.
Live-verified against Tokyo + Paris.
* fix(booking): harden search parser boundaries
* fix(booking): separate no-card drift from empty
|
||
|
|
7a2ab47bf8 | chore(skills): remove smart-search (#1683) | ||
|
|
2c8b50c4fd |
docs(readme): shrink CLI Hub + Core Concepts + merge Update into Install (#1681)
Per WAWQAQ:
1. **CLI Hub**: drop the 13-row 3-column table; enumerate just the
names inline ("gh · docker · vercel · wrangler · ntn · obsidian · …")
plus one-liner register / list commands. Removes "Manual install"
ntn note (search lives in external-clis.yaml / ntn's own docs).
Compresses the 7-row Desktop App Adapters table to a single inline
line pointing at docs/adapters/desktop/.
2. **Core Concepts** section dissolved: its four subsections
("browser", "Built-in adapters", "Writing a new adapter",
"CLI Hub and desktop adapters") duplicated the intro 3-bullet
+ later dedicated sections. Kept the substantive "Writing a new
adapter" callout as its own top-level section. The "For AI Agents
(Developer Guide)" tail block at the bottom was a third copy of
the same recipe — removed.
3. **Update** merged with **Install skills**: install header now
reads "Install skills (also refreshes existing installs)", and
the standalone Update section collapses to a single command
(`npm install -g @jackwener/opencli@latest && npx skills add ...`).
Net: EN 410 → 326 (-20%), ZH 455 → 366 (-20%). Same coverage; just
less repetition.
|
||
|
|
51a9456305 |
feat(linkedin): add people-search command (#1649)
* feat(linkedin): add people-search command (#1621) Closes #1621. Adds opencli linkedin people-search <keywords> for finding people on standard LinkedIn (not Sales Navigator). Architecture note. Standard LinkedIn moved its people search results page to Server-Driven UI / React Server Components on the /flagship-web/rsc-action/... path stack. The legacy Voyager REST endpoint /voyager/api/search/dash/clusters returns HTTP 500 from a web context; its modern camelCase rename voyagerSearchDashClusters returns the same. The result list is rendered server-side and the page HTML IS the result payload; Voyager calls from the page are sidebar / notification concerns, not search results. Extraction strategy. LinkedIn SSR uses obfuscated CSS class hashes (e.g. _997b7c77) that rotate on every deploy AND display:contents wrappers that flatten the DOM tree. Class-based selectors, walk-up- to-card logic, and anchor-pair element ranges all fail because no element boundary matches a person's card. Working approach: extract main.innerText once, split by newline, slice between consecutive person names. The names come from the aria-hidden spans of /in/<handle> anchors. LinkedIn's SSR emits a card as a name line followed by degree badge / headline / location / action labels before the next card's name line - a layout that has been stable through several DOM refactors. Critical filter: /in/<handle> anchors over-count because LinkedIn renders each mutual connection as a /in/ anchor inside another card's result. The skip() predicate during name-line lookup drops mutual-connection lines ("X, Y and N other mutual connections"), so anchors that don't have a real name line are filtered out. CUL caveat. LinkedIn imposes a monthly Commercial Use Limit on people search against the standard site. Burst behaviour is irrelevant - the limit is a calendar-month counter. The adapter runs one navigation per invocation (no pagination) so a single call costs exactly one CUL query. --limit is capped at 10 to keep a single call's information density high without surfacing the "reached commercial use limit" yellow banner faster. Schema: rank, name, headline, location, profile_url Live verified against kyfw 12306-style throttled cadence (sleep 60s between dev iterations to keep CUL consumption visible): 5/5 rows populated with name + headline + location + profile_url for the keyword "reinforcement learning". Mutual-connection anchors correctly filtered out so the row order matches LinkedIn's own ranking. Tests: 10 unit tests covering URL construction, limit validation, extraction-script invariants (anchor enumeration, text-slice approach, mutual-connection filter, aria-hidden span as name source), limit slicing, AuthRequiredError on missing JSESSIONID, CUL- flavoured CommandExecutionError on redirect, EmptyResultError on zero rows, ArgumentError on empty keywords, and registry shape. * fix(linkedin): harden people search typed boundaries * fix(linkedin): fail people search candidate parser drift --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
62592547a4 |
fix(adapters): migrate empty-data throws to EmptyResultError across 5 commands (#1674 follow-up) (#1678)
* fix(adapters): migrate empty-data throws to EmptyResultError across 5 commands (#1674 follow-up) Continues the structured-error migration owner started in #1674 (fix(xhs,youtube): 把合法空数据语义切到 EmptyResultError). Same motivation: callers need to distinguish "the platform legitimately has no data for this target" from "fetch infrastructure is broken, retry me", because downstream automation pipelines that batch over seed lists conflate the two and trip soft-rate-limit heuristics. Sites converted (5 commands, 6 throw sites): powerchina/search.js (2 sites): - "[taxonomy=empty_result] ... extracted only navigation/portal rows" - "[taxonomy=empty_result] ... api/dom yielded no result" Both already self-labelled with the empty_result taxonomy tag, making this the canonical fix. xiaohongshu/creator-notes.js, creator-notes-summary.js (both): - "No notes found. Are you logged into creator.xiaohongshu.com?" The "is logged in" hint is preserved in the empty message so users can self-diagnose, while the error type is now structured. xiaohongshu/creator-stats.js: - "No data for period <X>. Available: <a, b, c>" Empty-data condition: requested period exists in the API surface but has zero numeric data; available periods are still surfaced in the message. xiaohongshu/creator-note-detail.js: - "No note detail data found. Check note_id and login status..." Shape: exit code 66, stderr code: EMPTY_RESULT, matching bilibili/subtitle, xhs/user, youtube/transcript precedent. Out of scope: - tiktok/{user,notifications,explore}.js: throws live inside page.evaluate template strings and run in browser context; the Node-side caller already regex-routes them via throwTikTokPageContextError({emptyPattern: /No videos found/, ...}) to EmptyResultError. The existing design is correct. - eastmoney/_secid.js / antigravity/serve.js / instagram/collection-*: input-validation throws, ArgumentError territory not EmptyResultError. * test(adapters): cover empty-result migrations --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
4a3a55634a |
docs(readme): curate built-in commands to popular sites + add wrangler (#1679)
Per WAWQAQ: 1. Built-in Commands table cut from 30 EN rows / 86 ZH rows down to a curated 11-site list (xiaohongshu, bilibili, zhihu, hackernews, linkedin, reddit, twitter, claude, gemini, notebooklm, amazon). The README is meant to surface high-traffic / well-known sites; the long-tail (100+ adapters) is one click away via docs/adapters/index.md. linkedin (full) replaces linkedin-learning in the curated set per the spec. 2. Add Cloudflare Wrangler as a new external CLI passthrough: - src/external-clis.yaml entry (binary: wrangler, npm -g) - CLI Hub table row in EN + ZH READMEs - cli-manifest.json regen reflects the new entry (857 entries) |
||
|
|
da497f0b02 |
feat(zhihu): add answer comments reader
* feat(zhihu): add answer comments reader * fix(zhihu): harden answer-comments boundaries * fix(zhihu): keep answer comments flat --------- Co-authored-by: lihaidong <lihaidong@kingsoft.com> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
2590278f43 | fix(chatgpt): detect generated image surfaces (#1677) | ||
|
|
488e407a65 |
feat(twitter): add device-follow notification stream command
* feat(twitter): add device-follow command for /i/timeline notification stream (#1628) Closes #1628. Adds the twitter device-follow command, which reads the curated tweet list aggregated under a bell-icon "new posts from @userA and N others" notification. Direct GET /i/timeline redirects to /home, so the data is only reachable via the legacy v1.1 REST endpoint /i/api/2/notifications/device_follow.json , none of the existing twitter commands cover this stream: - twitter timeline home for-you / following feed (different endpoint) - twitter notifications the notification list itself, not aggregated tweets inside any one notification - twitter search search-based, can't reproduce the aggregation Endpoint discovery + field-mapping originally proposed by @traddo in #1628; this PR upstreams a clean implementation that: - Strategy.COOKIE + ct0 from CDP cookie jar + the public web bearer token from clis/twitter/utils.js (same auth path as twitter timeline) - Hits /i/api/2/notifications/device_follow.json directly via page.evaluate fetch on the x.com origin so SameSite=Lax cookies are preserved - Joins each entry.content.item.content.tweet.id to globalObjects.tweets[id] and resolves the author via globalObjects.users[tweet.user_id_str] - Returns the canonical twitter row columns (id, author, text, likes, retweets, replies, views, created_at, url), matching twitter timeline minus has_media / media_urls / card / quoted_tweet which the legacy v1.1 endpoint does not surface - Sets views: null rather than a 0 sentinel; the legacy endpoint does not return view counts even with include_ext_views=true, and the GraphQL TweetResultByRestId round-trip per tweet was judged too expensive for a list command (typed-errors §3: no scalar sentinels that lie about real engagement) - parseLimit enforces strict 1-200 integer validation with no silent clamping; the only baseline addition is the silent-sentinel on the "unknown" author fallback, which matches the exact precedent in twitter/timeline.js:76 that is already baselined Tests: 17 unit tests in device-follow.test.js cover parseLimit strict validation, URL parameter shape, entry/tweet join, user-resolution fallback, dedup via the seen set, empty-stream shape, the canonical column registration, AuthRequiredError on missing ct0, and CommandExecutionError on non-2xx fetch. Live verified the endpoint shape end-to-end against the logged-in session: HTTP 200 with the expected {globalObjects: {tweets, users}, timeline: {id: 'tweet_notifications', instructions: [{addEntries: {entries: []}}]}} envelope. The tester account has no bell-notification follows enabled, so entries is empty, but the shape and auth path are confirmed against the documented spec. * fix(twitter): harden device-follow typed boundaries * fix(twitter): fail fast on device-follow drift --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
e682c1c30a | fix(deps): restore Node 20 runtime compatibility (#1673) | ||
|
|
86f57c0846 |
feat(reddit): 在 listing 命令上暴露 post_hint / url / preview / gallery 4 个媒体路由列 (#1676)
* feat(reddit): 在 listing 命令上暴露 post_hint / url / preview / gallery 4 个媒体路由列 5 个 reddit listing 命令(popular / hot / frontpage / search / subreddit) 每行新增 4 列,下游消费者不用 scrape selftext 也能区分 image / gallery / hosted:video / link / self 五类内容: - `post_hint` — Reddit 自报的内容类型(image | hosted:video | link | self 等) - `url_overridden_by_dest` — 外链帖的原始 URL(image/link 类型才有) - `preview_image_url` — 缩略图地址(HTML-decoded,Reddit 即便 raw_json=1 也会在 preview URL 里返回 `&`) - `gallery_urls` — 多图相册数组(HTML-decoded) ## 实现 每个 adapter 的 evaluate 块内嵌两个 helper: - `decodeHtml(s)` — 6 个 HTML entity 替换(& / < / > / " / ' / ') - `extractRedditMedia(d)` — 从 post `data` 中抽 4 个字段,gallery_urls 从 `gallery_data.items[].media_id` × `media_metadata[id].s.u` 组合得到 helper 在每个 adapter 里 inline 复制(reddit 没有 shared 文件,模式跟现有 adapter 一致)。`clis/reddit/extract-media.test.js` 把 helper 行为锁在 8 个 fixture(plain / image / gallery / hosted-video / link / html-decode / 缺字段 / nullish input);每个 adapter 的 .test.js 额外 grep 自己源码里有 `function extractRedditMedia` 和 `...extractRedditMedia(c.data)` 两处接入痕迹,并断言 columns 数组形状。 frontpage 和 subreddit 之前 evaluate 返回原始 `children`、map 块按 `item.data.title` 索引;为了让 `gallery_urls` 这种数组字段能被 map 块的 模板字符串渲染,refactor 成和 popular/hot/search 一致的"evaluate 内部 就 map 成中间对象、map 块按 `item.title` 索引"模式。 ## 范围 只覆盖 5 个 listing 命令。**`read` 不在本 PR 内**:它的 evaluate 块在 post-#1651 时代已经是 error-kind-discriminated 的富结构(`kind: 'inaccessible'` / `kind: 'http'` / `kind: 'malformed'`),原始 commit 的"POST 行带 media、 comment 行空"模式和当前结构冲突太深,单独的 read 接入留作后续 PR。 完全 additive:既有字段名、顺序、值都不变;新字段加在每行末尾。 ## 验证 - `npx vitest run clis/reddit/ --project adapter` → 84/84 通过 - `node scripts/check-silent-column-drop.mjs` → current=97, baseline=97, new=0 - `npx tsc --noEmit` 干净 - `npm run build` 干净 * feat(reddit): expose home media route columns --------- Co-authored-by: huanghe <he.huang@extremevision.mo> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
4a6cfe8060 |
fix(xhs,youtube): 把合法空数据语义切到 EmptyResultError (#1674)
* fix(xhs,youtube): 把"合法空数据"语义切到 EmptyResultError 把 xiaohongshu/user 和 youtube/transcript 跟 bilibili/subtitle 已有的 structured-error 模式对齐,让下游能区分"用户/视频没内容"和"fetch 真的挂了"。 ## xiaohongshu user 返回 `No public notes found for this Xiaohongshu user` 的场景——目标用户 零公开笔记(销号 / 私密 / 全删)——原来抛 plain `Error`,下游无法和 "真的 fetch 失败 / cookie 死"区分。 改抛 `EmptyResultError`,exit code 变 66,stderr 携带 `code: EMPTY_RESULT`, 跟 `bilibili subtitle` empty 同 shape。 ## youtube transcript `No captions available for this video`(作者没开 CC、YouTube 也没自动生成) 是数据条件,不是基础设施失败。原来跟 HTTP/解析错误一样抛 `CommandExecutionError`,造成调用方反复重试。 改这个特定 case 抛 `EmptyResultError`;其他 caption 错误(HTTP / parse / empty response)继续走 `CommandExecutionError` 触发重试。 ## 为什么 downstream 需要这个 调用方(如自动化采集流水线)通常对 "data.length === 0 && exitCode !== 0" 做 soft-rate-limit 启发式:N 次连续 soft fail 触发 24h 平台跳过。当 seed 列表 里有变质条目(XHS 账号销号 / YouTube 视频丢失字幕),"empty" 响应堆积会 误触跳过——cookies 和平台本身都健康。EmptyResultError 让调用方能区分 "这个用户没内容"和"API 挂了"。 ## 测试 - `npx vitest run clis/xiaohongshu clis/youtube` —— 全过 - `npx tsc --noEmit` 干净 * fix(xhs): distinguish empty user notes from parser drift * fix(empty): tighten legal empty evidence --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
bcb0fb362f |
feat(twitter): expose bio on read command
* feat(twitter): 在 read 命令上暴露 bio(用户简介) `list-tweets` / `timeline` / `search` 三个读命令现在每行多一列 `bio`,从 `user.legacy.description` 抽。匹配 `profile` 命令已有的 `bio` 字段,让下游 消费者展示作者画像时省去"读了推文还要再读作者主页"的 roundtrip。 bio 在 user 对象缺失或没 description 时回落到 `''`。columns 数组同步更新, `--format columns` 会渲染 bio。完全 additive:既有字段名、顺序、值都不变。 延续 #1660 (card binding_values) 和 #1667 (quoted_tweet) 的同一类 read-side enrichment 模式。 ## 验证 - `clis/twitter/list-tweets.test.js` / `clis/twitter/search.test.js` 已有 shape assertion 补上 `bio: ''` 行 - `timeline.test.js` 用 `toMatchObject`(子集匹配),新增 bio 不会破断言 - `npx vitest run clis/twitter/list-tweets.test.js clis/twitter/timeline.test.js clis/twitter/search.test.js --project adapter` → 48/48 通过 - `npx tsc --noEmit` 干净 - `npm run build` 干净 - `node scripts/check-silent-column-drop.mjs` → current=97, baseline=97, new=0 * test(twitter): cover inline bio extraction * feat(twitter): expose thread author bio --------- Co-authored-by: huanghe <he.huang@extremevision.mo> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
85b1c07ba9 |
feat(zhihu): include answer links in question results
* feat(zhihu): include answer links in question results * fix(zhihu): avoid fake answer links for malformed ids * fix(zhihu): dedupe answers by trusted id --------- Co-authored-by: lihaidong <lihaidong@kingsoft.com> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
6fbaf0d5b8 |
feat(twitter): 在 read 命令上暴露 quoted_tweet(被引用的推文) (#1667)
* feat(twitter): expose quoted_tweet on read commands When a tweet quotes another tweet (embedded preview with commentary), the quoted tweet's content is in `tweet.quoted_status_result.result` — same `legacy / core / card / note_tweet` shape as the outer tweet. Until now none of the 5 read commands (list-tweets / timeline / thread / tweets / search) surfaced this nested object, so downstream consumers couldn't render the quoted preview card. Adds `extractQuotedTweet(tw)` in shared.js (mirrors the `extractMedia` / `extractCard` helper pattern) and threads it through all 5 read commands plus their CLI `columns:` declarations. Output shape is a deliberately small subset of the main tweet (id/author/name/text/created_at/url + media + card). Counts and full author bio are intentionally omitted to keep timeline payloads from ballooning 2-3x; consumers needing those can re-fetch `twitter thread <quoted_id>`. Notable edge cases tested in shared.test.js: - plain tweets (no `is_quote_status`) -> null - tombstoned / unavailable quoted tweets (deleted / privacy-restricted) -> null - TweetWithVisibilityResults `result.tweet` shim unwrap - long-form note_tweet text preferred over truncated full_text - quote-of-a-quote does NOT recurse (avoids payload explosion on threads where every reply re-quotes the root) * fix(twitter): require quoted tweet render evidence * fix(twitter): validate quoted tweet author shape --------- Co-authored-by: ml-scout <ml-scout@anthropic.com> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
34f351e59f |
feat(reddit/subscribed): 接入 LoginWallError 嗅探(#1650 的第一个 caller) (#1668)
* feat(errors,utils): 添加 LoginWallError 与 HTML-as-JSON 响应嗅探器
部分 adapter(twitter list-tweets/thread、reddit search/subreddit 等)历史上
直接 `JSON.parse(await r.text())` 或 `await r.json()` 解析响应。当服务端返回
登录墙、限流页或 WAF 拦截页(而不是 JSON)时,body 以 `<!DOCTYPE html>` 或
`<html ...>` 开头,解析直接抛出晦涩的
`SyntaxError: Unexpected token '<', "<!DOCTYPE "... is not valid JSON`,
调用方无法把它和真正的 JSON 解析失败区分开。
本 PR 添加三块共享基础设施,让 adapter 能识别 HTML 情况并抛出结构化的
`LoginWallError`(带 status / url / body 预览),而不是裸的解析栈:
- `LoginWallError`(src/errors.ts):新的 `CliError` 子类,含 `status`、
`url`、`bodyPreview` 字段,hint 提示"重新登录或等待限流过期",
退出码映射到 `EXIT_CODES.NOPERM`。
- `parseJsonOrThrowLoginWall(response)`(src/utils.ts):Node 端 helper,
供从 daemon 侧 fetch 的 adapter 使用(接收 Fetch Response)。
- `BROWSER_JSON_SNIFF_FN` + `throwIfLoginWall(value)`(src/utils.ts):
browser 端等价物。字符串片段嵌入到 `page.evaluate` 里,返回值是
解析后的 JSON,或 `{ error: status }` HTTP 形状,或 `LoginWallSignal`
哨兵对象(`{ __loginWall: true, status, url, ... }`),Node 侧拿到后
转成 `LoginWallError`。
行为是 opt-in:现有 adapter 不调用这些 helper 就完全不受影响。后续 PR
会把 reddit / twitter adapter 接入这些 helper。`src/utils.test.ts` 新增
13 个单测,覆盖 Node 端 + browser 端两条路径以及 body 预览的 100 字截断。
* feat(reddit/subscribed): 接入 LoginWallError 嗅探(#1650 的第一个 caller)
`reddit subscribed` 是从 daemon-backed browser session 调 Reddit 的 JSON API
(`/api/me.json` + `/subreddits/mine/subscriptions.json`)。原本两处 `await res.json()`
在 Reddit 返回登录墙 / WAF / over-18 拦截页(HTML body + 200 OK)时会抛
`SyntaxError: Unexpected token '<'`,被外层 try/catch 兜底成 `kind: 'exception'`
→ Node 侧报成 `CommandExecutionError: subscribed failed: SyntaxError ...`,
看不出 root cause。
接入 #1650 的 helper 后:
- **Browser 侧**:用 `BROWSER_JSON_SNIFF_FN` 提供的 `fetchJsonOrLoginWall(url, init)`
替换裸 `fetch + .json()`。helper 内部 sniff `Content-Type: text/html` 或
`<!DOCTYPE` / `<html` body 前缀,返回 `{ __loginWall: true, status, url,
contentType, bodyPreview }` 哨兵(不抛,交给调用者)。
- **Cross-boundary**:两处 fetch 站点(me.json + subscriptions.json)发现哨兵后
返回 `{ kind: 'login-wall', sentinel, where }` 透传给 Node。
- **Node 侧**:`throwIfLoginWall(result.sentinel, { url: result.where })` 把
哨兵转成结构化 `LoginWallError`(含 `status` / `url` / `bodyPreview` 字段,
exit code `EXIT_CODES.NOPERM`,hint 提示重新登录或等限流过期)。
这是 #1650 的第一个真实 caller,覆盖 3 块 export 全部(`BROWSER_JSON_SNIFF_FN`
+ `throwIfLoginWall` + `LoginWallError`)。其它 adapter 后续按这个模板逐个接入。
回归测试新增 1 个:mock evaluate 返回 `{ kind: 'login-wall', sentinel }`,
断言 Node 端抛 `LoginWallError` 且 `status` / `url` / `bodyPreview` 字段正确。
原有 12 个测试不动,全部通过。
|
||
|
|
acc18be999 |
docs(readme): tighten skill attribution + remove redundant Highlights (#1666)
Per WAWQAQ T1 + T2 review: T1 — skill attribution carries the same intent PR #1654 started but hadn't fully cleaned up: - Skill table row for `opencli-adapter-author` no longer claims it "operate[s] a site in real time" (SKILL.md explicitly says ad-hoc driving lives in `opencli-browser`). Browser-op example ("Help me check my Xiaohongshu notifications") moved to the `opencli-browser` row where it belongs. - "How it works" section's 5 browser primitives (navigate / read / interact / extract / wait) now point to `opencli-browser` instead of `opencli-adapter-author`. - Skill references list re-orders to surface `opencli-browser` first with a concrete description, and `opencli-adapter-author` no longer claims to cover "browser operation". T2 — drop the Highlights section. Pre-Quick-Start had four parallel summary blocks (3-line tagline / 3-bullet automation intro / CLI-hub + desktop line / 5-bullet Highlights) that all said the same thing. Highlights was the most-recent and most-redundant of the four; the remaining three carry the value props cleanly: tagline → three usage modes → CLI-hub + desktop scope. EN + ZH READMEs synced. |
||
|
|
67ed9e9c81 |
feat(linkedin-learning): add search / trending / course read commands (#1657)
* feat(linkedin-learning): add search / trending / course read commands (#1021) Closes #1021. Adds a new linkedin-learning site adapter with three read-only commands against LinkedIn Learning's public learning-api REST surface. Shares cookie session with linkedin.com; Learning queries are not subject to the people-search CUL. Commands: - linkedin-learning search <keywords> searchV2?q=keywords - linkedin-learning trending feedRecommendationGroups?q=learner - linkedin-learning course <slug> courses?q=slug Endpoints were discovered via browser network capture on /learning/search and /learning/<slug> pages: searchV2 returns a flat list of courses/videos/paths keyed by entityType, headline.title.text holds the canonical title, length is a TimeSpan in seconds, and rating is averaged from ratingSum/ratingCount when averageRating is missing. trending walks the carousels array on each recommendation group, flattens cards across them, dedups by slug, and respects --limit. Group is labeled with the carousel title (e.g. "Top picks for you") or the upstream annotation tag (TOP_PICKS). course accepts either a bare slug or a full /learning/<slug> URL, then hits /learning-api/courses?q=slug. The detail endpoint omits rating fields even when search reports them; this is documented in the adapter doc rather than fixed via a second /reviews fetch to keep the PR scoped to one endpoint per command. CUL caveat: Learning's API has no per-month limit, so dev iterations can be much more aggressive than the people-search adapter (#1649). Three commands were live-verified against a logged-in account with 60s sleeps between calls (conservative for first-pass safety). Tests: 28 unit tests across search.test.js (12), trending.test.js (6), course.test.js (10) cover URL construction, limit validation, author join, duration / rating coercion, row mapping, carousel flattening and dedup, slug parsing from URL forms, and the standard auth / empty / fetch-failure error paths. Live verified: - search "AI agent" --limit 3: 3 rows with title/instructor/rating - trending --limit 3: 3 personalized course picks - course agentic-ai-build-your-first-agentic-ai-system: title, 3932s duration, 18 videos, release date 2026-03-27 * fix(linkedin-learning): harden read result boundaries * fix(linkedin-learning): require course title evidence --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
8577d88ee7 |
fix(cli): escape leading-dash positional values via argv preprocessor (#1658)
* fix(cli): escape leading-dash positional values via argv preprocessor (#1160) Closes #1160. `opencli boss detail -abc123def` failed with `error: unknown option '-abc123def'` because commander treats any argv token starting with `-` as an option. BOSS 直聘 securityId tokens are opaque base64-ish strings that can legitimately start with `-`, and the same shape is possible for any adapter that takes an opaque-id positional. Adds escapeLeadingDashPositional() to src/cli-argv-preprocess.ts, called from main.ts after the existing rewriteBrowserArgv pass. The preprocessor: - Reads cli-manifest.json (the same manifest the registry uses) and builds a set of `<site>/<cmd>` keys whose first positional is required. - Walks past root flags (matching the existing rewriteBrowserArgv walker) to find the site + command tokens. - If the next argv token starts with `-`, is not the recognised short flags `-f` / `-v` / `-h`, is not `--*`, and is not the pre-escaped `--` separator, inserts `--` before it. Tests: 12 new unit tests in cli-argv-preprocess.test.ts cover the basic insertion, trailing-flag preservation, non-touched cases (normal values, recognised short flags, long flags, already-escaped, non-positional commands, unknown commands, short argv, and the `--profile work boss detail -abc` form that walks past a root value flag). Live verified: `node ./dist/src/main.js boss detail -abc123def` no longer raises 'unknown option'. The adapter now receives the dash-leading value and proceeds to fetch, where it correctly surfaces an upstream "missing required parameter" error for the fake id used in this smoke test. * fix(cli): preserve options around dash positionals * fix(cli): preserve attached short option values --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
cd731bd2aa |
feat(twitter): 在 read 命令上暴露 card binding_values(链接预览卡片) (#1660)
* feat(twitter): expose card binding_values on read commands Surface tweet link-preview cards (title, description, image, domain, landing URL) on `search`, `list-tweets`, `thread`, and `timeline` so downstream renderers can build native-style link cards without re-fetching. Pure GraphQL-response extractor — no query strategy, interceptor, or network changes. extractCard returns null when the tweet has no card or when the card is structurally empty (no url AND no title/description). Missing fields are omitted from the output to keep JSON consumers clean. * fix(twitter): bind cards to matching URL entity --------- Co-authored-by: ml-scout <ml-scout@anthropic.com> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
d1c714ecd3 |
feat(twitter): 新增 list-create 命令(GraphQL CreateList mutation) (#1656)
* feat(twitter): add list-create command Adds a new `twitter list-create` command so users can create Twitter/X lists from the CLI (the existing list commands only covered reading, adding, and removing members). Uses the GraphQL CreateList mutation with the same cookie + CSRF pattern as list-add, no UI clicks needed. Args: name (positional, max 25), --description (max 100), --mode (public|private). QueryId resolved at runtime via resolveTwitterQueryId, with a known fallback for offline / bundle-scan misses. * fix(twitter): pin list-create queryId + features to a working pair Twitter's GraphQL rejects CreateList when queryId and the features schema drift apart (DecodeException). Stop resolving the queryId dynamically (which would pull a newer schema), hardcode a known-good queryId, and trim features to the minimal set the real web client sends. Also: Twitter sometimes returns a non-fatal errors array from a side-effect serializer while still creating the list. Check for a valid list payload first and only treat errors as fatal when no list came back. * fix(twitter): add missing access:'write' on list-create (#9) `twitter/list-create` was missing the required `access` field, which made manifest validation fail on every opencli invocation and spam stderr with: ⚠ Failed to load manifest .../cli-manifest.json: Command twitter/list-create must declare access: 'read' | 'write' Per docs/conventions/convention-audit.md (rule missing-access-metadata), every adapter command must declare access. Since list-create is a create action, set access: 'write'. Also rebuilds cli-manifest.json — picks up missing `quoted_tweet` columns on list-tweets / search / list-tweets-username from PR #8 (which didn't rebuild the manifest). * fix(twitter): harden list-create mutation contract * fix(twitter): verify created list name --------- Co-authored-by: huanghe <he.huang@extremevision.mo> Co-authored-by: Kary <karyhe1019@gmail.com> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
8182ffbe89 |
chore(deps): bump tsx from 4.21.0 to 4.22.2 (#1663)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.21.0 to 4.22.2. - [Release notes](https://github.com/privatenumber/tsx/releases) - [Changelog](https://github.com/privatenumber/tsx/blob/master/release.config.cjs) - [Commits](https://github.com/privatenumber/tsx/compare/v4.21.0...v4.22.2) --- updated-dependencies: - dependency-name: tsx dependency-version: 4.22.2 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
5a4984789d |
chore(deps): bump @types/node from 25.6.0 to 25.9.0 (#1664)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 25.6.0 to 25.9.0. - [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases) - [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node) --- updated-dependencies: - dependency-name: "@types/node" dependency-version: 25.9.0 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
e1185da882 |
chore(deps): bump ws from 8.20.0 to 8.20.1 (#1662)
Bumps [ws](https://github.com/websockets/ws) from 8.20.0 to 8.20.1. - [Release notes](https://github.com/websockets/ws/releases) - [Commits](https://github.com/websockets/ws/compare/8.20.0...8.20.1) --- updated-dependencies: - dependency-name: ws dependency-version: 8.20.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
9446bddb60 |
chore(deps): bump undici from 6.25.0 to 8.3.0 (#1661)
Bumps [undici](https://github.com/nodejs/undici) from 6.25.0 to 8.3.0. - [Release notes](https://github.com/nodejs/undici/releases) - [Commits](https://github.com/nodejs/undici/compare/v6.25.0...v8.3.0) --- updated-dependencies: - dependency-name: undici dependency-version: 8.3.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
0c4bcdbb86 |
feat(reddit): 新增 subscribed 命令 + 在 listing 命令上暴露 id / created_utc / selftext (#1651)
* feat(reddit): subscribed command + expose id/created_utc/selftext on listing commands Adds `opencli reddit subscribed` to list the user's subscribed subreddits, mirroring `saved.js`'s cookie auth + AuthRequiredError pattern. Auto-paginates via `/subreddits/mine/subscriptions.json` (max 1000 subs, default 100). Also extends the JSON output of `popular` / `search` / `subreddit` with `id`, `created_utc`, `selftext` (and `author` on popular) — the table view stays clean (columns: unchanged), but `--format json` now surfaces fields needed for downstream content-recommendation tooling that filters by post age, dedupes by post id, or uses self-post bodies for embeddings. Tests: 4 new vitest cases for subscribed.js (happy / auth fail / HTTP / --limit truncation). All existing reddit tests still pass. Note on cli-manifest.json diff: the rebuild on fork/main drops 13 entries whose source files import lowercase `selectorError` from `@jackwener/opencli/errors` (the actual export is `SelectorError` — casing bug pre-existing in fork/main). Not introduced by this PR. * fix(reddit): harden subscribed listing contract * fix(reddit): require subreddit identity for subscriptions --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
ec3b7dadf3 |
fix(zhihu): decode numeric entities in answer detail (#1629)
Co-authored-by: lihaidong <lihaidong@kingsoft.com> |
||
|
|
4de04c43ad |
feat(suno): add suno.com music-generation adapter (#1638)
* fix(suno): harden generation adapter contracts * fix(suno): separate session auth and API failures --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
40592ea5fd |
fix(adapters): drop silent-sentinel row fallbacks across Apple Podcasts, Reddit, and Gitee (#1634)
* fix(adapters): drop silent-sentinel row fallbacks across Apple Podcasts, Reddit, and Gitee Continues the audit-baseline cleanup from #1611 (lesswrong) and #1631 (wikipedia / 36kr / xiaoyuzhou / zhihu), and follows the direction set by |
||
|
|
942539a695 |
fix(twitter/lists): 跳过 "Discover new Lists" 推荐区块,避免被当成用户的 list 抓取 (#1652)
* fix(twitter): skip "Discover new Lists" recommendations in lists adapter The X.com /<user>/lists page powers two sections from a single ListsManagementPageTimeline GraphQL response: "Discover new Lists" (algorithmic recommendations) and "Your Lists" (owned + subscribed). The previous parser ignored entry.entryId entirely and returned every list it found, so recommendations leaked through and downstream consumers treated them as the user's own lists. X distinguishes the sections by entry.entryId prefix: owned-subscribed-list-module-* → owned + subscribed (keep) list-to-follow-module-* → Discover recommendations (drop) cursor-* → pagination cursor (no list payload) Filter on the owned-subscribed prefix in parseListsManagement and expose isOwnedSubscribedEntry for testing. The existing test fixture used a fictional entryId shape that no longer matches real responses; update it to the nested-module shape Twitter actually returns and add two new tests: one proving Discover entries are skipped, and one for the entryId classifier. Verified end-to-end against a live account: 10 raw entries (3 Discover + 7 owned/subscribed) now correctly return 7 owned/subscribed lists with zero leakage. * fix(twitter): harden lists parser boundary * fix(twitter): require list-remove postcondition evidence --------- Co-authored-by: huanghe <he.huang@extremevision.mo> Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
dc645a5bcf |
fix(youtube/transcript): 把 timedtext URL 匹配限定到当前 videoId,修跨视频字幕串台 (#1655)
* fix(youtube/transcript): scope timedtext URL match to current videoId
YouTube watch-page is an SPA — page.goto between watch URLs preserves
performance.getEntriesByType('resource') entries from prior videos.
findTimedtextUrl filtered only by lang, so a previously-viewed
same-language video's timedtext URL could be picked up by the polling
loop before the current video's fetch hook captured a fresh one,
returning the wrong video's captions to the caller.
Fix: require URLs to contain v=<currentVideoId> across all three paths:
- in-page findTimedtextUrl (resource-buffer scan)
- in-page isJson3TimedtextUrl (fetch/XHR hook)
- Node-side extractSegmentsFromNetworkCapture (CDP capture)
Most likely to hit callers that reuse a single daemon tab to fetch
many transcripts back-to-back (e.g. ml-scout). Confirmed in the wild:
a Fox News Ukraine clip got Whisper Flow promo captions written to
its row when the prior call on the same tab pulled an English
Whisper Flow video.
Adds one source-contract assertion (both in-page sites use a shared
videoIdMarker) and one behavioral test (CDP capture buffer with a
stale 'v=prev' entry alongside the current 'v=abc' returns only the
current video's captions).
* fix(youtube): exact-match transcript timedtext video id
---------
Co-authored-by: ml-scout <ml-scout@anthropic.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
|
||
|
|
f0d9aa187c |
docs(readme): fix skill attribution for "operate any website" use case (#1654)
Per WAWQAQ feedback: the intro section's "Let AI Agents operate any website" bullet mistakenly references `opencli-adapter-author`, which is the skill for **writing** adapters (correctly referenced in the adjacent "Write new adapters" bullet). The skill for ad-hoc browser driving is `opencli-browser` — its own SKILL.md frontmatter explicitly says "Not for writing adapters — see opencli-adapter-author for that", and `opencli-adapter-author` says "For ad-hoc browser driving (no adapter), see opencli-browser instead". Two locations affected with the same error: the intro bullet and the Core Concepts > `browser` section. Both EN and ZH READMEs updated. |
||
|
|
e82e32abc6 |
feat(12306): add full read adapter (stations / trains / train / price / me / passengers / orders) (#1637)
* feat(12306): add stations / trains / train read commands (no login required) Adds a first-pass 12306 (中国铁路) adapter for the public anonymous query endpoints. Closes the no-login slice of #1589. The authenticated `me / passengers / orders` commands the issue proposes are explicitly left as a follow-up. Commands: - 12306 stations <keyword> search station bundle - 12306 trains <from> <to> --date YYYY-MM-DD availability between stations - 12306 train <train-no> --from <s> --to <s> --date stop list All three use Strategy.PUBLIC + browser: false, anonymous, no cookie storage, no CAPTCHA bypass. Sensitive behaviors the issue rules out (ticket sniping, order submission, payment, anti-abuse circumvention, password storage) are not implemented. Notes worth flagging for review: - 12306 rejects anonymous query endpoints with HTTP 302 to /mormhweb/logFiles/error.html. The adapter first hits /otn/leftTicket/init to mint JSESSIONID / route / BIGipServerotn cookies, then attaches them to subsequent queries. No CAPTCHA path. - 12306 rotates the train-query endpoint name (queryO / queryZ / queryA / queryG) every few weeks. When the wrong name is hit the server returns `{c_url: "leftTicket/queryX", status: false}` pointing to the current correct name. The adapter walks a list of known names, captures the rotation hint, and retries; the runtime list is also mutated so subsequent calls in the same process skip the warm-up round trip. - The `|`-separated train wire format includes a booking-handshake `secret` field at position 0. Since this PR is read-only and the issue explicitly rules out booking, that field is parsed but not surfaced in the returned row, and a unit test asserts it cannot leak via the public adapter contract. - Station resolution accepts Chinese name (`上海虹桥`), telecode (`AOH`), full pinyin (`shanghaihongqiao`), or short alias (`shhq`). Anything else raises ArgumentError with a hint. - `limit` arguments use a tight validator that throws ArgumentError on non-integer / out-of-range input rather than silently clamping, matching the typed-error pattern used in #1397 (grok) and #1370 (coupang). Live verified anonymously against kyfw.12306.cn: - `12306 stations 上海 --limit 5` returns 5 stations including 上海 (SHH) / 上海南 (SNH) / 上海虹桥 (AOH). - `12306 trains 北京 上海 --date 2026-05-22 --limit 1` returns G547 06:18 -> 12:11 with first / second / business / no-seat availability columns populated. - `12306 train 24000000G10L --from 北京南 --to 上海虹桥 --date 2026-05-22` returns the 7-stop G1 route from 北京南 through 沧州西 / 德州东 / 曲阜东 / 南京南 / 苏州北 to 上海虹桥, with arrival / departure / stopover times. Tests: 18 unit tests covering parseStationBundle, resolveStation (including ambiguous / case-insensitive cases), validateDate, buildCookieHeader, parseTrainRecord (including a regression test asserting the `secret` field cannot leak into the row). Deliberately deferred to a follow-up: `12306 price`. The queryTicketPrice endpoint needs train_no + per-stop station_no + per-train seat-type letters, so an ergonomic `12306 price <code>` would cascade three API calls (trains -> stops -> price) per invocation. Wanted to keep this PR's blast radius small. If the maintainer prefers a Phase 1 that includes price even with the cascading-call cost, happy to add it. * feat(12306): add me / passengers / orders / price authenticated + price read commands Completes the #1589 12306 (中国铁路) adapter on top of the stations / trains / train slice landed in the prior commit of this branch. The full command set is now: Anonymous (no login): 12306 stations search station bundle by Chinese / telecode / pinyin 12306 trains list trains between two stations on a date 12306 train list stops of one train 12306 price ticket prices for one train segment + date Authenticated (cookie session): 12306 me account summary (sensitive fields masked by default) 12306 passengers saved-passenger list (sensitive fields masked) 12306 orders in-progress orders (not yet ridden / refunded) Notes worth flagging for review: - 12306 sets the auth cookie `tk` and the session cookie `JSESSIONID` with `Path=/otn`. CDP `Network.getCookies` filters by URL path, so `page.getCookies({ url: 'https://kyfw.12306.cn' })` returns 7 cookies without `tk` / `JSESSIONID`, even on a freshly-navigated logged-in tab. Switched the login check to read `document.cookie` via `page.evaluate`, which the current navigated page exposes regardless of cookie path. Centralized as `require12306Login` in utils.js so all three authenticated commands share the same check. - All authenticated commands mask sensitive fields by default: - `me`: real name (Chinese mask), email, mobile (12306 already masks server-side), birth date (year only). - `passengers`: name + birth year by default; 12306 already masks ID number and mobile server-side and this adapter never decodes those. - Both expose `--include-sensitive` to opt back into the unmasked fields the user is entitled to see on their own account. - `orders` returns the `queryMyOrderNoComplete` slice (orders that have not yet been ridden / refunded / completed). The historical `queryMyOrderApi` endpoint requires extra page-state handshakes that proved fragile when probed; left as a follow-up so this command can ship reliably for the immediate "what's still on my account" use case. - `price` cascades three anonymous API calls per invocation: init -> queryByTrainNo (to resolve segment station_no within the train route) -> queryTicketPrice. 12306 returns prices keyed by one-or-two-letter seat codes (`A9` 商务座 / `M` 一等座 / `O` 二等座 / `WZ` 无座 / etc.) and additionally doubles some up as bare numeric codes (e.g. `"9": "21580"` mirrors `"A9": "¥2158.0"`); the bare-numeric duplicates are filtered out so the row set is one-per-seat-class. - Strictly anonymous queries; no CAPTCHA / slider / SMS bypass, no credential storage, no ticket sniping, no order submission, no payment - per the issue's Non-goals list. Live verified anonymously and authenticated against kyfw.12306.cn, sleeping 15-25 seconds between hits to keep 12306's anti-abuse throttle gentle: - 12306 me: account summary returned with real_name / email / mobile / birth date all masked at the adapter level, on top of 12306's own server-side mobile mask. - 12306 passengers: every saved passenger returned with name masked to `<surname>*<...>` and 12306-side ID/mobile masks preserved verbatim. - 12306 orders: empty for this test account (no in-progress orders), correct EmptyResultError surface. - 12306 price G1 北京南 -> 上海虹桥 2026-05-22: returns 商务座 ¥2158 / 特等座 ¥1163 / 一等座 ¥1035 / 二等座 ¥626 / 无座 ¥626, sorted desc. Tests: 23 unit tests (5 new beyond the prior commit's 18) cover the mask helpers (email / mobile / Chinese name) plus the parsePriceData filter that drops the bare-numeric duplicates and sorts by descending price. * fix(12306): harden browser auth boundaries * fix(12306): tighten API drift boundaries --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
254d51835f |
fix: keep media filenames in output directory (#1642)
* fix: keep media filenames in output directory * fix(download): sanitize media filename segments --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
87dfb68e74 |
fix(browser): goto 重试时回收陈旧 page identity + 把 -32000 "Cannot find default execution context" 归类为可重试 (#1645)
* fix(browser): recover from stale page identity on goto retry (#5) When a chrome-backed adapter pre-navigates after its cached `_page` targetId has been invalidated (tab closed externally, identity evicted), the extension throws `Page not found: <id> — stale page identity` and the failure cascades — every subsequent persistent-site session call in the same process keeps re-sending the same dead targetId. Observed in a downstream parallel multi-platform recall: a single dead page handle got reused across 4+ calls (twitter thread / twitter search / reddit search) because there was no detection or recovery. The same hash appeared in adapter pre-navigations to youtube, twitter, reddit, xhs back-to-back in seconds, suggesting the cached `_page` was shared via persistent site session leases (`site:youtube` etc) and never cleared after the first "stale page identity" response. Page.goto() now catches that specific error, drops `_page`, and retries once without the stale id. The retry navigates via session-lease resolution in the extension (resolveTab → preferredTabId / new owned tab), which already handles tab eviction correctly. No effect on the happy path. Three regression tests in src/browser/page.test.ts cover: - recovery: stale id dropped, retry succeeds with new identity - no-cache safety: fresh page with no _page → error propagates unchanged (nothing to drop, retrying would loop) - error scoping: unrelated extension errors (e.g. disconnected) still surface immediately — no implicit retry * fix(errors): classify -32000 "Cannot find default execution context" as retryable (#6) classifyBrowserError previously only matched CDP -32000 errors when the message contained "target" (e.g., "target closed"). It missed "Cannot find default execution context", a CDP protocol error that also indicates the inspected target went away — observed in a downstream parallel adapter recall against youtube channels. Widening the secondary check to `/target|context/i` lets the existing target-navigation retry path (200ms delay + re-attach) recover instead of surfacing the error as non-retryable. * fix(browser): tighten stale page recovery notes --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
000c867f3a |
fix(zhihu): harden search pagination (#1615)
Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
24f643af16 |
feat(xianyu): add inbox, messages, and reply commands (#1639)
* fix: tighten internal callback types * feat(xianyu): add private message commands * fix(xianyu): harden IM command contracts --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
1f30a9027b |
feat(linkedin): consolidate messaging and Sales Navigator commands (#1647)
* fix(linkedin): harden sales navigator commands * fix(linkedin): harden salesnav message boundaries --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
72f2b020de |
feat(weread-official): add official gateway CLI
Add the WeRead official Agent Gateway as an in-tree pure HTTP adapter with 8 commands, typed errors, tests, and docs. |
||
|
|
261b8bfbb5 |
build: restore +x on dist/src/main.js after tsc rebuild (#1644)
clean-dist deletes dist/ and tsc --build re-emits files without preserving the executable bit on the bin entry. Symlinked global install then hits EACCES on spawn until manually chmod'd. Chain a chmodSync into the existing prebuild-manifest hook so any future rebuild self-heals. node -e instead of bare `chmod +x` to keep the script portable (npm runs on Windows via Git Bash where chmod is a no-op, but fs.chmodSync still silently no-ops there too — no extra branching needed). Co-authored-by: Kary <karyhe1019@gmail.com> |
||
|
|
1e7ebe7f27 |
feat(twitter): rewrite download profile path on GraphQL UserMedia with cursor pagination (#1636)
* fix(twitter): harden profile media download * fix(twitter): fail closed on repeated media cursor --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
7e44e71150 |
fix(lesswrong): drop "Unknown" silent sentinel in author column (#1611)
* fix(lesswrong): drop "Unknown" silent sentinel in author column
Twelve lesswrong commands had `author: item.user?.displayName ?? 'Unknown'`
which masks the missing-author signal: an agent reading the result row
cannot distinguish "post has no associated user" from "author is literally
named Unknown". The repo's typed-error lint flags this pattern
(silent-sentinel rule, see scripts/check-typed-error-lint.mjs:323).
Replace `?? 'Unknown'` with `?? ''` so the missing-author case stays
visible as an empty string. Consistent with `clis/lesswrong/_helpers.js:68`
which was already using the empty-signal form.
Shrinks scripts/typed-error-lint-baseline.json from 173 to 161 entries.
Follows the same direction as #1603 (fix(adapters): surface silent empty
fallbacks).
Verified live: `opencli lesswrong frontpage --limit 2 -f json` returns
real posts with non-empty author values; empty-author rows would now
show `"author": ""` instead of fabricating `"Unknown"`.
* test(lesswrong): add empty-signal coverage for the author sentinel swap
Per owner's pattern in
|
||
|
|
76a9c78261 |
feat(weibo): add delete command to remove user's own posts (#1620)
* feat(weibo): add delete command to remove user's own posts
Adds `opencli weibo delete <id>` so the same workflow that creates a
post can also remove one without leaving the CLI. The id positional
accepts either the numeric `idstr` (e.g. `5299336218674412`) or the
base62 `mblogid` (e.g. `QFGbHAoBS`) found in any weibo URL or in the
output of `weibo me` / `weibo feed` / `weibo post`.
Implementation lives in a single `page.evaluate` IIFE so cookies +
the XSRF-TOKEN double-submit token stay first-party:
1. Resolve mblogid / idstr via `GET /ajax/statuses/show?id=<input>`,
which returns the canonical `idstr`. Empty result -> 404 path.
2. Read the `XSRF-TOKEN` cookie via `document.cookie`.
3. `POST /ajax/statuses/destroy` with `id=<idstr>` body and the
`X-Xsrf-Token` header.
4. Return `[{ status: 'deleted', id, mblogid }]`.
Typed errors:
- 401 / 403 from either show or destroy -> `AuthRequiredError`
- `show` returning no `idstr` -> `EmptyResultError`
- Non-2xx HTTP on either call -> `CommandExecutionError` with status
- API response `ok !== 1` -> `CommandExecutionError` with the API msg
Closes #1619.
Verified live on macOS / opencli v1.7.22, weibo cookie session:
- Deleted the lingering test post from #1602 verification
(idstr=5299336218674412, mblogid=QFGbHAoBS):
`weibo delete QFGbHAoBS` returned
`[{ status: 'deleted', id: '5299336218674412', mblogid: 'QFGbHAoBS' }]`
- `weibo me` shows `statuses: 3` (was 4 before the delete)
- `weibo post QFGbHAoBS` now throws "Post not found"
Unit tests: 8 / 8 in `clis/weibo/delete.test.js` (happy path,
empty-id, auth, not-found, show-http, destroy-http, api-msg, envelope
unwrap). Full weibo suite: 38 / 38 pass.
* fix(weibo): require delete postcondition evidence
---------
Co-authored-by: jackwener <jakevingoo@gmail.com>
|
||
|
|
030a0ad885 |
feat(xiaohongshu): add delete-note command to remove published notes (#1624)
* fix(xiaohongshu/publish): invoke shadow-DOM publish handler directly XHS creator center now wraps the publish/save-draft button in an `<xhs-publish-btn>` web component backed by a CLOSED shadow root. Calling `.click()` on the host element does not dispatch into the internal handler, and CDP coordinate clicks cannot penetrate the shadow boundary. The previous text-match `button.click()` loop hit the host element, returned `ok`, and yet the note silently stayed on the publish page as a draft, so the adapter reported the soft `⚠️ 操作完成,请在浏览器中确认` status while nothing was actually posted. Invoke the publish/save method directly on the `<xhs-publish-btn>` host (`_onPublish` / `_onSave` and a few candidate names XHS has shipped historically). Fall back to the legacy `<button>`/`[role="button"]` text-match click for older creator-center variants that still expose plain buttons. Patch shape suggested by the OpenCLI autofix report in #1606 from @chcc-funny (who verified an end-to-end real publish locally). Closes #1606. Verified live on macOS / opencli v1.7.22 / extension v1.0.15, with creator center logged in: - `opencli xiaohongshu publish ... --draft` -> `✅ 暂存成功`, creator home shows "草稿箱中有未发布的作品" - `opencli xiaohongshu publish ...` (real publish) -> `✅ 发布成功`, note appeared on the account feed (visible from mobile app); test note deleted after verification Unit tests: 12 / 12 in `clis/xiaohongshu/publish.test.js` pass (mocks updated to reflect the new `{ ok, via, name|text }` invoke result shape). * feat(xiaohongshu): add delete-note command to remove published notes Adds `opencli xiaohongshu delete-note <note-id>` so the workflow that creates a note can also remove one without leaving the CLI, mirroring `weibo delete` (#1619 / #1620). The creator-center HTTP delete API requires the `X-S-Common` signature header that `publish.js` deliberately avoids, so this follows the same UI automation route. Flow: 1. Navigate to creator note-manager 2. Switch to "已发布" tab (delete entry only appears there; "审核中" and "未通过" rows have no web delete action, mobile app only) 3. Locate the `.note` row whose `data-impression` JSON contains the target noteId (exact JSON-parsed match, not substring, so values that happen to share the noteId prefix in other fields cannot match the wrong row) 4. Click the inline `<span class="control data-del">` action 5. Click "确定" in the `.d-modal-footer` confirmation modal 6. Poll for the row disappearing (iteration-bounded so tests with mocked `page.wait` exhaust the loop quickly) Typed errors: - /login redirect after navigation: AuthRequiredError - 已发布 tab not found / not clickable: CommandExecutionError (UI drift) - target noteId not present in the rendered list: EmptyResultError with a hint about review-state limitation - row found but no delete action visible: CommandExecutionError - confirmation modal missing / no 确定 button: CommandExecutionError - row still visible after the configured poll window: CommandExecutionError Closes #1623. Verified live: published a test note, deleted via this adapter, follow-up `xiaohongshu creator-notes` confirms it is gone. Unit tests: 8 / 8 cover happy path, empty-id ArgumentError, login redirect AuthRequiredError, tab-not-found CommandExecutionError, row-not-found EmptyResultError, no-delete-action / no-modal / unverified-delete CommandExecutionError paths. Built on top of #1613 (xiaohongshu publish shadow-DOM fix) so the live verify could exercise publish-then-delete end to end. Will rebase onto main once #1613 lands. * fix(xhs): make delete-note fail closed * fix(xiaohongshu): harden delete-note boundary --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
e29150bab5 |
fix(weibo/publish): replace brittle CSS-module hash with placeholder selector (#1625)
* fix(weibo/publish): replace brittle CSS-module hash with placeholder selector `clis/weibo/publish.js` matched the compose textarea via `textarea._input_13iqr_8`, where `_input_13iqr_8` is the Vite CSS-module hash Weibo rebuilds on every frontend deploy. The hash drifted (current build emits `_input_1f5hn_8`), so step 4 of the publish flow throws "Weibo compose editor did not appear" before anything else can run. Reported in #1602. Replace the single hashed selector with a placeholder-text-based chain that survives Weibo's CSS-module rebuilds: textarea[placeholder*="有什么新鲜事"] textarea[placeholder*="新鲜事"] textarea._input_13iqr_8 // legacy hash kept last for older variants Two visible textareas can match on the home feed (the always-rendered "home-strip" prompt + the post-click modal compose). Pick the LAST visible candidate: the modal opens on top and is appended to DOM later, so the last-visible textarea is the modal. Both the editor-visibility poll (Step 4) and the text-insertion step (Step 6) use the same chain. Also drops `evaluateWithArgs` from Step 8 success polling. The IIFE there does not reference any outer args, but `evaluateWithArgs` injects its `const`-bound parameter names into the page context, and re-running on each iteration of the success-poll loop threw `Identifier 'maxIterations' has already been declared` after the first iteration. This was masked previously because Step 4 always failed first; with the selector fixed, the latent Step 8 bug surfaces. Switched to plain `page.evaluate` to avoid re-declaring per loop. Closes #1602. Verified live on macOS / opencli built locally / extension v1.0.15, weibo cookie session: - `opencli weibo publish "明洞那家店真不错"` returned `status: success, message: 发布成功, text: 明洞那家店真不错` - Confirmed via `/ajax/statuses/mymblog`: the post landed at `idstr=5299403716821218`, `mblogid=QFHWzsCvE`, text matches what was typed (proves selector chain picks the right textarea and the text insertion path works end-to-end) - Cleaned up: deleted via the same `/ajax/statuses/destroy` path that PR #1620 exposes as `weibo delete` Unit tests: 8 / 8 in `clis/weibo/publish.test.js` pass (mocks updated to reflect the new `evaluate`-vs-`evaluateWithArgs` split for Step 8 and the longer poll window). * test(weibo): lock publish placeholder selector path --------- Co-authored-by: jackwener <jakevingoo@gmail.com> |
||
|
|
a50074d684 |
fix(adapters): drop silent-sentinel row fallbacks across 6 read commands (#1631)
* fix(adapters): drop silent-sentinel row fallbacks across 6 read commands
Continues the audit-baseline cleanup started in #1611 (lesswrong) and
the direction set by #1599 / #1603 / #1604. Replaces the
`silent-sentinel` row-data fallbacks (`'Unknown'` / `'-'` / `'unknown'`
that mask missing fields) with the empty-string signal so agents can
tell apart "field really has the value Unknown" from "upstream returned
no value".
Touched 6 read adapters, 10 baseline entries:
- wikipedia/trending: title, description
- 36kr/article: author, date, body
- xiaoyuzhou/download: podcast
- xiaoyuzhou/transcript: podcast
- zhihu/collection: dedup key + type field (the empty prefix still
produces a unique-per-content dedup key, just without the `unknown:`
noise)
- zhihu/download: author
Intentionally skipped (line-by-line audited):
- v2ex/me.js: `'Unknown'` is an in-band control-flow sentinel. Line 35
initialises `let username = 'Unknown';`, line 41 uses
`if (username === 'Unknown')` to trigger the profileEl fallback
selector, line 75 uses the same check to raise the auth error.
Empty would silently bypass both checks and return a row with an
empty username as if auth succeeded.
- v2ex/daily.js: `'未知'` is user-facing 签到 success text in the
rendered status message, not a row field. Empty would render a
broken sentence.
- weibo/comments.js, weibo/feed.js: the sentinel sits inside an in-IIFE
error-message string composition (`'API error: ' + (data.msg || 'unknown')`),
not in a returned row. Empty would silently truncate diagnostic
output. Both stay on baseline.
Verified live: `opencli wikipedia trending --limit 3` and `opencli 36kr
hot --limit 2` both return populated rows; the empty-string signal only
kicks in when the upstream value is actually missing.
* test(adapters): add empty-signal coverage for the cluster-2 sentinel swap
Per owner's pattern in
|
||
|
|
368581ea4d |
fix(electron-apps): move codex CDP port off 9222 to avoid browser-bridge collision (#1630)
* fix(electron-apps): move codex CDP port off 9222 to avoid browser-bridge collision
`src/electron-apps.ts` had `codex: { port: 9222 }`, but `9222` is the
default Chrome DevTools port that opencli's own browser-bridge Chrome
binds whenever `opencli doctor` is OK. On every normal opencli install
the bridge owns 9222 first, so Codex Desktop can never bind it, and
`opencli codex status` (plus every other codex command) fails with:
App launched but CDP not available on port 9222 after 15s
`~/.opencli/apps.yaml` is documented as "additive only, does not
override builtins", so users have no supported way to relocate the
port from the user side.
Reported in #1626 with full repro (Codex Desktop + active opencli
browser-bridge Chrome) and root-cause pointer at
`dist/src/electron-apps.js:13`. Every other electron app in the
builtin registry already uses a distinct port in the 9224-9236
band (cursor 9226, doubao-app 9225, chatwise 9228, discord-app 9232,
antigravity 9234, chatgpt-app 9236); codex was the only one that
collided with the browser bridge.
Move codex to 9238 (the next free slot in that band, also the value
the reporter recommended). Update the test that asserts the port and
the two docs references that mention codex=9222. The pitfall entry
in `docs/advanced/electron.md` is also annotated to explicitly call
out 9222 as the bridge's port to avoid future collisions.
Closes #1626.
Verified live: `opencli codex status -v` now emits
`[verbose] [launcher] Probing CDP on port 9238...` (was 9222 before
the fix), confirming the code path picks up the new port. Full
end-to-end with a real Codex Desktop install is left to the reporter
and reviewer; the change here is a single-value config update plus
docs/tests sync.
Unit tests: 7 / 7 in `src/electron-apps.test.ts` pass (the codex-port
assertion updated to 9238). Both audit gates pass.
* docs(electron): sync codex CDP port guidance
---------
Co-authored-by: jackwener <jakevingoo@gmail.com>
|
||
|
|
0c488bbf51 |
docs(readme): simplify Highlights from 9 to 5 bullets (#1605)
Per WAWQAQ feedback: the previous Highlights list was bloated with hollow marketing phrases and overlapping bullets (e.g. "Browser Automation for AI Agents" + "AI Agent ready" said the same thing twice, "Pipeable, scriptable, CI-friendly" is generic CLI filler). Cut "AI Agent ready", "Account-safe" (folded into Live Browser Automation), "Deterministic"'s second sentence (folded into Zero LLM cost), and merged "Website → CLI" with "CLI Hub" into "100+ adapters + CLI Hub". Result is 5 concrete capability bullets instead of 9, each tied to a real feature. EN and ZH READMEs kept in sync. |