Commit Graph

454 Commits

Author SHA1 Message Date
jakevin 12176de1a5 refactor: simplify core modules (#784)
* refactor: simplify core modules — remove root shims, consolidate error classification, streamline cascade/interceptor, clean up synthesize

1. Remove root-level shim files (errors.ts, logger.ts, registry.ts, types.ts, utils.ts, launcher.ts) — update all ~840 adapter imports to reference src/ directly
2. Consolidate interceptor: reuse shared DISGUISE_FN in tap interceptor instead of reimplementing
3. Unify error classification: single ClassifiedError type with icon/exitCode/hint lookup table, eliminating duplicated pattern matching between resolveExitCode and renderError
4. Simplify cascade probe: replace repetitive switch cases with PROBE_OPTIONS lookup map
12. Clean up synthesize.ts: remove deprecated snake_case field aliases (recommended_args, recommended_columns, recommendedColumnsLegacy) and unnecessary constant aliases

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: update import paths in contributor docs and skill templates

Update all documentation and skill files to reference src/ directly,
matching the shim removal in the previous commit.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: update new Gemini adapter imports to use src/ paths

Fix imports in newly added deep-research adapter files that were
still referencing the deleted root shim files.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: update xiaohongshu tests for /search_result/ URL change

Tests now expect /search_result/<id> for bare note IDs (matching
the note-helpers.ts change from PR #774) and updated empty-shell
hint assertion.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: update LessWrong and hupu adapter imports to use src/ paths

Fix imports in newly merged LessWrong and hupu/mentions adapter
files that were still referencing the deleted root shim files.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test(xueqiu): mock logger via src path

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-05 03:35:12 +08:00
Inori333 97ae87ccee fix(cli): make operate verify work in source checkouts (#777)
* fix(cli): make operate verify work in source checkouts

* fix(cli): resolve operate verify entry from package metadata

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-05 02:56:10 +08:00
jakevin 639a31fc84 fix: review follow-ups for monorepo adapter separation (#783)
* fix: review follow-ups — better first-run log, OPENCLI_FETCH=1 skips version check

- Clarify first-run log message: "copying adapters (one-time setup)"
- Add comment explaining why scriptPath uses two levels of ../
- OPENCLI_FETCH=1 now bypasses version-skip to allow forced refresh

* fix: update doc-coverage script path after clis/ move

check-doc-coverage.sh still referenced src/clis/ after PR #782 moved
adapters to root clis/. This caused CI to fail with "0/1 documented".

* fix: resolve package root dynamically for symlink and first-run paths

The symlink at ~/.opencli/node_modules/@jackwener/opencli pointed to
dist/ instead of the package root in prod mode, breaking user TS CLIs
that import from '@jackwener/opencli/registry'.

The first-run scriptPath also resolved incorrectly in dev mode.

Extract findPackageRoot() that walks up to find package.json, fixing
both paths for dev (src/) and prod (dist/src/) layouts.
2026-04-05 02:25:53 +08:00
jakevin 80eef46b4e refactor: monorepo adapter separation (clis/ at root) (#782)
* refactor: move adapters from src/clis/ to root clis/ for monorepo separation

Separates CLI adapters from the core runtime to prepare for independent
adapter distribution via postinstall fetch.

Key changes:
- Move src/clis/ → clis/ (adapters at repo root)
- Change tsconfig rootDir from "src" to "." so tsc compiles both
- Create root-level shim files (registry.ts, errors.ts, etc.) so adapter
  relative imports (../../registry.js) resolve correctly
- Update build-manifest.ts, main.ts paths for new dist/src/ structure
- Expand ensureUserCliCompatShims() to cover all adapter import targets
  (types, utils, logger, launcher, browser/*, download/*, pipeline/*)
- Add scripts/fetch-adapters.js postinstall for ~/.opencli/clis/ sync
- Update vitest.config.ts adapter test paths
- Add package.json files field to exclude adapters from npm package

Official adapter files are unconditionally overwritten on update;
user-created files not in the manifest are preserved.

* fix: add dist/clis/ and cli-manifest.json to npm files, harden fetch-adapters

- Add dist/clis/ and dist/cli-manifest.json to package.json files field
  so built-in adapters and manifest ship with the npm package
- Replace execSync with execFileSync to prevent command injection
- Add version check to skip redundant adapter fetches
- Track tmpRoot explicitly for reliable cleanup

* fix: address review blockers — manifest-based updates, global-only fetch, first-run fallback

1. Manifest-based update strategy:
   - Read old manifest to identify previously-official files
   - Clean up files removed upstream (in old manifest but not new)
   - User-created files (never in any manifest) remain untouched

2. Only run fetch-adapters on global install (npm_config_global=true)
   or explicit OPENCLI_FETCH=1, preventing heavy side effects for
   local/dev installs

3. First-run fallback in discovery.ts:
   - ensureUserAdapters() checks for adapter-manifest.json
   - If missing and ~/.opencli/clis/ is empty, spawns fetch-adapters.js
   - Guarantees adapters are available even with --ignore-scripts

* fix: remove OPENCLI_FETCH env var, use internal _OPENCLI_FIRST_RUN instead

* feat: also support OPENCLI_FETCH=1 for explicit adapter fetch trigger

* simplify: replace git clone with local copy from dist/clis/

Adapters already ship in the npm package (dist/clis/), so there's no
need to clone from GitHub. Copy directly from the installed package:

- Eliminates git, curl, tar dependencies
- No network calls in postinstall
- No timeout/offline issues
- Version always matches the installed CLI
- ~65 lines of clone/download code replaced by one cpSync loop
2026-04-05 01:46:36 +08:00
TennyZhuang 60bee91650 fix: match the requested tweet before deleting on X (#781)
* fix(twitter): match target tweet before deleting

* review: normalize invalid twitter delete URLs

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 23:57:01 +08:00
Inori333 d818b5bed8 fix(completion): sync top-level command suggestions (#588) 2026-04-04 16:26:31 +08:00
Ray e82649379b feat(instagram): add post, reel, story, and note publishing (#671)
* Add draft Instagram posting flow

* Refine Instagram post flow

* Add dynamic Instagram posting routes

* Retry transient Instagram private setup failures

* Add Instagram reel posting command

* Add Instagram mixed-media carousel posting

* Unify Instagram post media input

* Add Instagram story posting command

* Add Instagram note publishing command

* fix(instagram): use JSON.stringify for constants in note evaluate string

Replace template literal interpolation of Node-side constants with
JSON.stringify() for consistency with codebase evaluate patterns.
Use bracket notation for dynamic property access instead of template
interpolation into a property chain.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 16:22:22 +08:00
kevin.zhang e0a66af0f0 feat(hupu): add Hupu adapter (#751)
* feat(hupu): add hupu cli adapter

* fix(hupu): prevent detail from returning the wrong thread

* refactor: deduplicate shared utilities in hupu adapter

- Merge postHupuJson and postHupuReplyJson into single function with mode parameter
- Move stripHtml and decodeHtmlEntities to utils.ts, remove duplicate definitions

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 16:03:14 +08:00
YoungCan-Wang c86e677b78 推特支持回复图片 (#756)
* feat: 推特新增回复图片能力支持本地路径和网络路径

* fix(twitter/reply): fix image upload fallback, restore execCommand, add size limit

- Fix attachReplyImage fallback: use uploaded flag instead of checking
  page.setFileInput existence, so base64 fallback actually runs when
  CDP setFileInput throws "Unknown action"
- Restore execCommand('insertText') as primary text input method for
  Twitter's Draft.js editor, with paste event as fallback
- Add 20MB size limit for remote image downloads to prevent OOM
- Remove unsafe buttons[0] fallback that could click invisible buttons

* fix(twitter/reply): add local image size check and base64 fallback warning

Local images were not validated for size — a 100MB file would fail only
at upload time. Remote images already had MAX_IMAGE_SIZE_BYTES checks.
Also add a console.warn when using the base64 fallback with large
payloads, consistent with xiaohongshu/publish.ts behavior.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 16:00:35 +08:00
yulin7645 6364934423 feat(xiaoe): add 小鹅通 (Xiaoe-tech) student platform adapter (#617)
* feat(xiaoe): add 小鹅通 (Xiaoe-tech) student platform adapter

Add 5 YAML adapters for 小鹅通 (xiaoe-tech.com), the leading Chinese
online education platform:

- courses: list purchased courses with URLs and shop names
- detail: course info (name, price, user count, shop)
- catalog: full course outline supporting normal courses (type 50),
  columns (type 6), and big columns (type 8)
- play-url: get M3U8 play URL via direct API for video courses,
  and Vue component tree search + Performance API polling for
  live replay courses
- content: extract rich-text page content as plain text

Technical notes:
- Strategy: cookie (reuses Chrome login session)
- Framework: Vue 2 + Vuex Store (SPA)
- Video courses use a two-step API chain:
  detail_info.get → play_sign → getPlayUrl → M3U8
- Live replays use Performance API + Vue data tree polling
- Catalog expands chapters via Vue component method getSecitonList()
- Supports multiple stores (cross-domain cookie sharing via
  study.xiaoe-tech.com)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* review: stop truncating xiaoe content

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 15:54:14 +08:00
AstroHan ef78aaf3a2 fix: add -v/--verbose to built-in browser commands (#719)
* fix: add -v/--verbose to explore, record, generate, cascade

Built-in browser commands were registered directly in cli.ts and
missed the -v/--verbose flag that commanderAdapter.ts wires up for
adapter commands. Also switch explore's lone log.debug() call to
log.verbose() so the flag has visible effect.

Closes #716

* refactor(cli): make builtin command wiring testable

* refactor(cli): simplify verbose wiring, use normal Commander pattern

Replace registerVerboseAction wrapper with simple applyVerbose() helper.
The wrapper broke Commander's builder chain and created awkward
indentation. Now each command uses standard .option().action() with
applyVerbose(opts) as the first line — easier to read and maintain.

* fix(cli): add -v/--verbose to doctor and synthesize commands

These commands were also missing verbose support, same root cause as
explore/record/generate/cascade — registered directly in cli.ts,
bypassing commanderAdapter's automatic -v wiring.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 15:44:23 +08:00
artshooter 292b12d9b1 feat(twitter): add --images flag to post command (#666)
* feat(twitter): add --images flag to post command

Support attaching up to 4 images when posting tweets via
`opencli twitter post "text" --images /path/a.png,/path/b.jpg`.

Uses the existing CDP DOM.setFileInputFiles mechanism (page.setFileInput)
to inject files into Twitter's file input. Includes proper file validation,
graceful error handling for older extensions, and polling-based upload
readiness detection instead of fixed delays.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(twitter): use attachments DOM signal for upload detection, add tests

Replace unreliable tweet-button-only polling with dual-condition check:
wait for [data-testid="attachments"] with correct [role="group"] count
AND button enabled. Increase timeout to 30s. Add 8 unit tests covering
image upload flow, file validation, and error paths.

Addresses PR #666 review feedback.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(twitter): use top-level imports, fix test mocks, faster upload poll

- Use top-level fs/path imports instead of dynamic imports inside func
- Fix test statSync mock to return undefined (not null) for missing files
- Fix test path mock to preserve other exports via importOriginal
- Fix null type error in no-browser-session test
- Reduce upload poll interval from 1s to 500ms for faster detection
- Use JSON.stringify for imageCount interpolation for consistency

* refactor(twitter): extract validation, fail-fast, reduce duplication

- Extract validateImagePaths() with extension validation (jpg/png/gif/webp)
  matching xiaohongshu publish pattern
- Validate images before browser navigation (fail-fast on bad input)
- Remove try/catch wrapper around setFileInput — let errors propagate
  naturally instead of masking the original error
- Deduplicate tweetButton/tweetButtonInline lookups using fallback OR
- Use constants for MAX_IMAGES, UPLOAD_POLL_MS, UPLOAD_TIMEOUT_MS
- Add tests: unsupported format, validates-before-navigating

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 15:31:01 +08:00
AstroHan 2c5066d1f4 fix(gemini): stabilize ask reply state handling (#735)
* fix(gemini): stabilize ask reply state handling

* fix: use CommandExecutionError for composer failures and clean up formatting

- Replace raw Error with CommandExecutionError for Node-side composer
  failures (prepareComposer, insertText) to match adapter error conventions
- Remove extra blank lines after __test__ export

* refactor: remove dead code and add Chinese sign-in label

- Remove unused areGeminiTurnsEqual and areGeminiLinesEqual functions
- Add Chinese sign-in label (登录) to sign-in detection for consistency
  with other Chinese labels already added in this PR

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 15:26:58 +08:00
yichuanzhao99-ctrl 8eefa3b1c9 增加新浪财经热搜股票榜 (#736)
* 增加新浪财经热搜股票榜

* fix: address review issues in stock-rank adapter

- Fix string interpolation injection: use JSON.stringify for market param
- Add choices validation for market arg (cn/hk/us/wh/ft)
- Normalize column names to lowercase (rank/name/symbol/market/price/change/url)
- Add navigateBefore: false to avoid redundant navigation
- Add null safety on tabEl with optional chaining
- Remove unused waitForElement helper and unnecessary await on querySelectorAll
- Remove unrelated ?from=opencli tracking change from rolling-news.ts
- Remove ?from=opencli tracking from stock-rank URLs

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 15:16:05 +08:00
jakevin 052bf8bbf7 refactor(zhihu): simplify question evaluate to follow pixivFetch pattern (#754)
Move data processing (HTML stripping, answer mapping) from browser-side
evaluate to Node-side, keeping the evaluate minimal: just fetch + status
check. Uses __httpError sentinel consistent with pixivFetch convention.
2026-04-04 15:06:16 +08:00
jakevin c3c3abbbff fix(1688): remove MOQ extraction from price field, rename firstLine to firstWord, fix sales regex (#755)
- Remove hover_price_text as MOQ source in search normalizeSearchCandidate
  to prevent price fields from being misinterpreted as MOQ data
- Rename firstLine() to firstWord() to match its actual behavior (splits
  by whitespace, not newlines)
- Add missing "单" unit to item.ts extractSalesText regex
- Add test case verifying hover_price_text is not used for MOQ
2026-04-04 14:53:42 +08:00
GanFanNewOrder 81de69be3a feat(1688): add browser adapter and docs (#650)
* feat(1688): add browser adapter and docs

* fix(1688): retry alternate store seed offers

* feat(1688): harden adapter contracts and search pagination

---------

Co-authored-by: 泽加武 <zejiawu@zejiawudeMac-mini.local>
2026-04-04 14:45:32 +08:00
Kyrie Cai 855eaee04e fix(zhihu): make question runtime-compatible (#732)
* fix(zhihu): make question runtime-compatible

* fix: validate questionId is numeric to prevent interpolation issues

* refactor: simplify evaluate string and harden against injection

- Build URL in Node.js, embed via JSON.stringify for safety-by-design
- Remove unnecessary (page as any) cast — IPage already has evaluate
- Simplify error message construction (no nested ternaries)
- Replace implementation-detail test with numeric ID validation test

* refactor: simplify zhihu question — move stripHtml into evaluate, return clean data

* fix: add colon separator in fetch error message for readability

"request failed Failed to fetch" → "request failed: Failed to fetch"

---------

Co-authored-by: Kyrie <kyrie@mallab.world>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 14:36:27 +08:00
ykfnxx 1bcd96f38a fix(douban): fix marks pagination and improve subject data extraction (#752)
1. marks: correct pageSize from 30 to 15 — douban grid mode shows 15
   items per page, causing pagination to stop after the first page.

2. subject: split title/originalTitle correctly — v:itemreviewed contains
   both Chinese and original titles concatenated.

3. subject: extract country/region from #info as list, split by "/".

4. subject: extract duration as pure number (min) from v:runtime or #info.

5. subject: return casts as list instead of comma-joined string.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-04 14:32:51 +08:00
jakevin 748b09261d fix: handle missing electron executable gracefully (#747)
* fix: handle missing electron executable gracefully

* fix: support antigravity electron executable fallback
2026-04-04 02:03:30 +08:00
jakevin 7aafd4af59 fix(tests): update mocks for resolveBvid and Windows platform guards (#749)
- bilibili subtitle/comments tests: use importOriginal to include
  resolveBvid in utils mock
- comments test: use valid BV ID format for aid-resolution error test
- launcher test: skip pgrep test on win32 (detectProcess early-returns)
2026-04-04 01:43:42 +08:00
deepziyu a5abd3769f fix(launcher): graceful degradation and manual CDP override for Windows (#744)
* fix(windows): graceful degradation and manual CDP override for Electron apps

* fix: validate OPENCLI_CDP_ENDPOINT with probeCDP before use

Fail-fast with a clear error if the manual CDP endpoint is not reachable,
instead of passing a bad URL downstream and getting a confusing error.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 01:32:55 +08:00
sa1ka 8070960444 feat(bilibili): support b23.tv short URL/short code resolution (#740)
* feat(bilibili): support b23.tv short URL/short code resolution

Add resolveBvid() in utils.ts to automatically resolve b23.tv short URLs
and short codes to BV IDs. Supports all input formats:
- BV ID: BV1MV9NBtENN (pass through)
- Short code: XYzsqGa
- Short URL: https://b23.tv/XYzsqGa, b23.tv/XYzsqGa

Uses Node.js https.get with 302 redirect only (no body download),
typically ~100-250ms resolution time.

Applied to: subtitle, comments, download commands.

* fix: add timeout, input coercion, and tests for resolveBvid

- 5s timeout on https.get to prevent hanging on unresponsive b23.tv
- Accept unknown input type with String() coercion
- Simplify callers (remove redundant String().trim() wrappers)
- Add unit tests for BV ID passthrough and edge cases

---------

Co-authored-by: chenruinian <chenruinian@Sa1kas-MacBookPro.local>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-04 01:30:23 +08:00
jakevin b1c0bcb464 feat(autoresearch): add Layer 4 Save-as-CLI eval with zhihu/xhs coverage (#741)
* feat(autoresearch): add Layer 4 "Save as CLI" eval + fix operate verify

- New eval-save.ts: tests full init → write → verify pipeline (14 tasks)
- 8 PUBLIC strategy tasks (httpbin, jsonplaceholder, HN, wiki, lobsters, devto)
- 6 COOKIE strategy tasks (zhihu hot/search/question, xhs feed/search/note)
- New save-reliability preset for autoresearch engine iteration
- Fix: operate verify no longer hardcodes --limit 3 for adapters without limit arg
- Rename sediment → save throughout

* experiment(operate): 两个新任务都基于已有通过任务使用的同一 API,期望 pass_count 从 14 → 16。

* experiment(operate): Added 2 new tasks ( and ) that use the exact same APIs already proven to pass in exis

* experiment(operate): Both new tasks pass. The change adds 2 more  tasks ( and ) using the same proven API, i

* fix(autoresearch): rename SedimentTask → SaveTask, fix bracket indent, gitignore results.tsv

* refactor(autoresearch): complex multi-step save tasks + adapterFile support

- Replace simple COOKIE tasks with 6 complex multi-step chains:
  - zhihu: hot+top-answer (6-step), search+question-stats (7-step), question+answers+related (8-step)
  - xhs: search+scroll+dedup (6-step), note+comments (7-step), explore+scroll+sort (8-step)
- Move complex adapter code to save-adapters/*.ts files (avoids JSON escape issues)
- eval-save.ts: support adapterFile field to read adapter from file
- Preset scope now includes skills/opencli-operate/SKILL.md for skill improvement
- All 20/20 tasks passing

* experiment(save): add hn-best and hn-jobs tasks using proven Firebase API pattern, pass_count 20→22

* fix(autoresearch): increase Claude Code timeout 180s → 300s to reduce ETIMEDOUT failures

* experiment(save): add restcountries and nager-holidays tasks using stable public APIs, pass_count 22→24
2026-04-04 01:27:50 +08:00
Josh e18e0ed7a4 fix(browser): mention Chromium in Browser Bridge hints (#738) 2026-04-03 22:28:16 +08:00
jakevin c161f0f9f0 feat: auto-downgrade output to YAML in non-TTY (#737)
* feat: auto-downgrade table output to YAML in non-TTY environments

When stdout is not a TTY (pipes, AI agents, subprocesses), automatically
output YAML instead of table with ANSI colors and box-drawing characters.
This makes opencli output parseable by downstream tools and AI agents.

Behavior:
- TTY: table (default, unchanged)
- Non-TTY: yaml (auto-detected)
- OUTPUT env var: overrides auto-detection (yaml/json/table/etc)
- Explicit -f flag: always respected

* fix: TTY detection now works with commanderAdapter default fmt

- fmt='table' from commanderAdapter now correctly triggers non-TTY downgrade
- Priority: explicit -f (non-table) > OUTPUT env var > TTY auto-detect
- Added test for explicit -f precedence over OUTPUT env var

* fix: explicit -f flag now takes precedence over TTY auto-detection

Use Commander's getOptionValueSource to distinguish explicit -f from
default. Explicit -f table in non-TTY keeps table output. Only auto-
downgrade when user didn't pass -f.

Priority: explicit -f > OUTPUT env var > TTY auto-detect > table default

* fix: explicit -f also skips command defaultFormat override

When user passes -f explicitly, command-level defaultFormat (e.g.
gemini/ask defaultFormat:'plain') no longer overrides their choice.
2026-04-03 22:26:51 +08:00
GanFanNewOrder dcad060230 feat(amazon): unify ranking commands for bestsellers/new-releases/movers-shakers (#724)
* feat(amazon): unify ranking adapters for three signal boards

* refactor: simplify bestsellers wrapper and fix pagination detection for all ranking types

1. Remove unnecessary __test__ wrapper from bestsellers.ts — the test
   now uses normalizeRankingCandidate directly from rankings.ts,
   eliminating a needless indirection layer.

2. Fix isRankingPaginationUrl to detect pagination refs for all ranking
   types: zg_bs_pg_ (bestsellers), zg_bsnr_pg_ (new releases),
   zg_bsms_pg_ (movers & shakers). Previously only matched the
   bestsellers-specific ref pattern.

---------

Co-authored-by: 泽加武 <zejiawu@zejiawudeMac-mini.local>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-03 19:12:59 +08:00
jakevin ff84d19ded fix: SVG className crash + viewport expansion + test suites (#733)
* feat: AutoResearch framework + V2EX test suite (40 tasks)

AutoResearch framework (Karpathy-style autonomous iteration):
- engine.ts: 8-phase loop (review → modify → commit → verify → guard → decide → log)
- config.ts: typed config + CLI parser + metric extraction
- logger.ts: TSV append-only results log
- commands/run.ts: main loop spawning Claude Code per iteration
- commands/plan.ts: interactive config wizard
- commands/fix.ts: auto-detect broken state, iteratively fix
- commands/debug.ts: hypothesis-driven debugging for failing tasks

V2EX test suite (5 layers, 40 tasks):
- L1 Atomic (10): open, state, click, scroll, eval, back, wait
- L2 Single Page (10): hot topics, node list, topic meta, pagination
- L3 Multi-Step (10): click-read, navigate-node, tab-then-topic, pagination
- L4 Write Ops (5): reply typing, favorite detection, form detection
- L5 Complex Chain (5): cross-page collect, multi-node compare, full workflow

Presets: operate-reliability, skill-quality, v2ex-reliability

* test: V2EX test suite 60/60 — fix selectors, add harder tasks

- Fix v2ex-collect-hot-authors selector (pathname-based member link detection)
- Fix v2ex-wait-text judge (accept "appeared")
- Fix trailing commas in eval step strings
- Add 20 harder tasks: state+click interaction + long chain workflows
- Baseline: 60/60 across all layers

* feat: Zhihu test suite — 60 tasks across 8 layers, 60/60 passing

Knowledge-intensive Chinese Q&A site (React SPA, lazy loading, complex DOM):

- L1 Atomic (10): open, state, title, url, scroll, tab, back, wait, keys, screenshot
- L2 Feed (8): feed titles, hot list, metrics, tabs, authors, content types, avatar, search
- L3 Question (8): title, meta, answer, votes, buttons, descriptions, answer count
- L4 Navigation (8): hot→question, feed→question, author profile, search, topic, user, back
- L5 Write (6): upvote/follow/comment/bookmark/write-answer/share button detection
- L6 Chain (8): read-answer-author, author-profile, multi-hot, search-then-read, scroll-answers
- L7 Search (6): basic, people, topic, click-result, filter, back
- L8 Complex (6): full workflow, deep author chain, cross-question, search-read, 3-page, scroll-deep

Key fixes during development:
- Zhihu search page needs 5s+ wait (SPA lazy loading)
- Back navigation goes to about:blank (daemon init page), fixed with direct navigate
- User profile answers page needs 4s wait for content
- Broader selectors needed (h2 a instead of specific class names)

* feat: combined eval-all runner + combined-reliability preset

* experiment(operate): fix extract-npm-description + nav-click-link-example

Round 1: Fix 2 remaining browse-tasks failures:
- extract-npm-description: use generic <p> selector instead of class-based
- nav-click-link-example: include URL in output (title is 'Example Domains', not 'IANA')

* experiment(operate): fix bench-imdb-matrix — use broader selectors for year/rating

Round 2: IMDB page selectors were too specific (data-testid changed).
Use generic h1 for title, link text match for year, broader class match for rating.

* experiment(operate): add edge cases + fix SPA navigation timing

Round 3: Add 10 edge case tasks (5 V2EX + 5 Zhihu):
- rapid-navigate: 3 consecutive opens
- eval-after-click: verify URL changes after SPA click
- scroll-and-extract: extract after deep scroll
- structured extraction: multi-field JSON from dynamic content
- lazy-load answers: scroll triggers more content

Key finding: Zhihu SPA click() doesn't update location.pathname
immediately. Use window.location.href = a.href for reliable navigation.

V2EX: 65/65, Zhihu: 65/65, Browse: 59/59 = 189/189

* experiment(operate): add agent-style tasks using state+click+type (no eval for interaction)

Round 4-5: Add 5 tasks that test the actual agent workflow:
- agent-click-first-topic: find topic index via data-opencli-ref
- agent-type-search: type into search using state index
- agent-click-navigate-back: click by ref, verify navigation
- agent-state-has-interactive: verify state output format
- agent-state-after-scroll: verify scroll position in state

V2EX: 70/70 tasks

* fix: review fixes — extractVerdict, stderr, dead code

- eval-skill.ts: remove dead TASKS_FILE variable (skill-tasks.yaml never existed)
- eval-skill.ts: rewrite extractVerdict to use brace-counting JSON.parse
  instead of regex (handles escaped quotes in explanation)
- eval-browse.ts: include stderr in runCommand error output for debuggability

* fix: SVG className crash in dom-snapshot + viewport expansion

Critical bug: isSearchElement() called el.className.toLowerCase() which
crashes on SVG elements where className is SVGAnimatedString (not a string).
This caused the entire DOM snapshot to fail and fall back to the basic
accessibility tree, losing ALL interactive element indices.

Fix: use typeof check + baseVal fallback for SVG className.

Also:
- Increase viewportExpand from 800 to 2000 (covers ~3 screens)
- Add DEBUG_SNAPSHOT env var for snapshot failure debugging

Impact on Zhihu hot page:
- Before: 50 interactive elements (accessibility tree fallback), 1/30 hot links indexed
- After: 597 interactive elements (proper DOM snapshot), 19/30 hot links indexed
2026-04-03 19:01:10 +08:00
Ted Li f2a3ee6ee4 fix(doubao): preserve image URLs in read output (#708)
* fix doubao image urls in read output

* fix(doubao): derive image selector from messageTextSelectors

Hardcoded image selector only covered the first two text selectors,
so images inside class-based message containers would be missed.
Generate from the shared selector list for consistency.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-03 17:08:33 +08:00
tiaot33 f377ec000c feat(元宝): add browser adapter and docs (#693)
* feat(yuanbao): add browser adapter and docs

* refactor(yuanbao): normalize adapter failures to CliError

* refactor: extract shared yuanbao helpers to reduce duplication

Move isOnYuanbao, ensureYuanbaoPage, hasLoginGate, authRequired,
and IS_VISIBLE_JS to shared.ts. This eliminates identical copies
across ask.ts and new.ts, reducing correctness risk when modifying
shared logic.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-03 17:03:48 +08:00
jakevin 988908f348 refactor(xiaohongshu): replace blind retry with MutationObserver wait (#730)
* refactor(xiaohongshu): replace blind retry with MutationObserver wait

Instead of retrying the entire navigation when search results are empty,
use a MutationObserver to wait for `section.note-item` elements (or login
wall text) to appear in the DOM, with a 5s timeout. This is faster (resolves
as soon as content renders) and more correct (addresses the root cause of
delayed hydration rather than working around it with a full re-navigation).

* simplify: merge login-wall detection into MutationObserver wait

WAIT_FOR_CONTENT_JS now returns 'content', 'login_wall', or 'timeout'
instead of just true/false. This eliminates the separate login-wall
evaluate call and the redundant loginWall field in the extraction payload.
Two evaluate calls total (wait + extract) instead of three.
2026-04-03 16:40:33 +08:00
GanFanNewOrder 2b623b35b6 fix(xiaohongshu): retry once on intermittent empty first paint (#681)
Co-authored-by: 泽加武 <zejiawu@zejiawudeMac-mini.local>
2026-04-03 16:26:26 +08:00
BruceLoveDecimal 835c146fb7 fix(doubao-app): connect to correct CDP target instead of background … (#674)
* fix(doubao-app): connect to correct CDP target instead of background page

Doubao desktop app exposes multiple CDP targets. The scoring logic picked
the background page (doubao-background) over the actual chat page because
its URL-as-title contained "doubao", boosting its score above the real
chat page (title "豆包"). This caused all commands (send, ask, read) to
fail with "No textarea found".

- Add `targetFilter` field to ElectronAppEntry for per-app preferred target
- Set doubao-app targetFilter to 'doubao-chat/chat'
- Penalize background/new-tab-page URLs and URL-like titles in scoring
- Thread cdpTargetFilter through execution → runtime → CDPBridge

Closes #634, closes #506

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(cdp): exclude background targets instead of targetFilter

Replace the targetFilter plumbing (4 files, new interface field) with
a single-line fix: exclude `background_page` and `service_worker`
type targets from CDP selection entirely.

Background pages should never be connection targets — they have no
visible DOM and all selectors will fail. This is the root cause of
#506/#634 (doubao-app connecting to empty background page).

Simpler fix: 1 line added vs 4 files modified. No new interface
fields, no per-app configuration needed.

---------

Co-authored-by: 刘启灏 <liuqihao@liuqihaodeMacBook-Pro.local>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-03 12:52:31 +08:00
jakevin fc818b3c2c fix: classify xianyu item auth and blocked states (#726)
* fix: classify xianyu item auth and blocked states

* fix: classify xianyu item auth and blocked states
2026-04-03 12:50:48 +08:00
BruceLoveDecimal 0ce46b15bb feat:add xianyu (#696)
* feat:add xianyu

feat:add xianyu

feat:add xianyu

* chore:add xianyu docs

* fix:update xianyu after review

---------

Co-authored-by: 刘启灏 <liuqihao@liuqihaodeMacBook-Pro.local>
2026-04-03 12:30:28 +08:00
jakevin 1708626731 fix: update BrowserBridge test to mock fetchDaemonStatus instead of isDaemonRunning (#714)
PR #712 refactored _ensureDaemon to use a single fetchDaemonStatus() call
instead of separate isDaemonRunning(). The test was still mocking the old
function, causing it to fall through to the spawn-daemon path and throw
the wrong error message.
2026-04-03 03:44:56 +08:00
jakevin 5fe081b28c perf: optimize browser pipeline — tab query dedup, parallel stealth, incremental snapshots (#713)
* perf: optimize browser pipeline — tab query dedup, parallel stealth, incremental snapshots

- resolveTab() now returns { tabId, tab } so handleNavigate skips redundant chrome.tabs.get()
- goto() fires stealth injection in parallel with navigation instead of sequentially
- snapshot() passes previousHashes to enable incremental diff marking on consecutive calls

* revert: remove stealth parallelization — simplicity over performance
2026-04-03 03:32:24 +08:00
jakevin 0c75ab3f7e perf: reduce round-trips in browser command hot path (#712)
1. eval retry delay: 1000ms → 200ms for SPA navigation errors, 500ms
   for debugger detach. SPA navigations recover within ~100ms, the old
   1000ms delay was unnecessarily long.

2. Window creation: replace fixed 200ms sleep with tab-load poll.
   Listens for chrome.tabs.onUpdated status=complete with 500ms
   fallback cap. about:blank loads in ~20ms, saving ~180ms.

3. bridge.ts _ensureDaemon: single fetchDaemonStatus() call instead of
   two sequential calls (isExtensionConnected + isDaemonRunning both
   called fetchDaemonStatus independently). Saves one HTTP round-trip.

4. goto() post-navigation: coalesce stealth injection + DOM settle into
   a single exec call. Previously two sequential round-trips
   (Node→daemon→WS→extension→CDP each). Saves ~60-160ms per goto().
2026-04-03 03:24:54 +08:00
jakevin cd5da59187 perf: skip blank page on first browser command (#710)
Two changes that eliminate the about:blank → target-domain navigation
on first command execution:

1. Extension: getAutomationWindow() accepts an optional initialUrl.
   When creating a new window, uses the target URL directly instead
   of about:blank. handleNavigate() passes cmd.url through so the
   window starts on the correct domain.

2. CLI: Remove isAlreadyOnDomain() check before pre-nav. Instead,
   always call page.goto(preNavUrl) — the extension's handleNavigate
   already has a fast-path that skips navigation when the tab is
   already at the target URL. This avoids an extra exec round-trip
   (getCurrentUrl eval) on first command.

Net effect: first command saves ~1-3s (one fewer page load),
subsequent commands behave the same (navigate fast-path handles
domain matching efficiently via chrome.tabs.get).
2026-04-03 02:59:15 +08:00
jakevin d7d5211fde refactor: remove unused newTab() and closeTab() from IPage interface (#709)
Both methods had zero production callers — only test mocks referenced them.
newTab() created about:blank pages via CDP Target.createTarget, but no
adapter or pipeline step ever invoked it. closeTab() was similarly unused.

selectTab() and tabs() are kept as they have active production usage
(e.g. doubao adapter). The scoreTarget about:blank penalty is retained
as a defensive measure against user-opened blank tabs.
2026-04-03 02:42:03 +08:00
jakevin de817730ca feat: Browser Use best practices — click/type/state improvements (#707)
* docs: improve operate skill with Browser Use best practices

- Add Critical Rules section (state over screenshot, verify with get value)
- Add Command Cost Guide (free/instant vs expensive vision tokens)
- Add Action Chaining Rules (safe to chain vs page-changing)
- Add Tips section
- Fix Core Workflow to use state/get value for verification, not screenshot
- Mark screenshot as "ONLY for user deliverables"

Inspired by Browser Use's design: DOM-first state representation,
action cost awareness, and multi-action chaining patterns.

* docs: fix operate skill — eval read-only, IIFE, interaction rules

- Add rule: NEVER use eval to click/type — use click/type/select commands
  (eval bypasses scrollIntoView + CDP pipeline, fails on off-screen elements)
- Add rule: eval is read-only, always wrap in IIFE to avoid variable conflicts
- Reorder Critical Rules for priority
- Add IIFE example in Extract section

Root cause: Claude Code was using eval("el.click()") instead of
click <index>, and hitting "already declared" errors from repeated
eval calls in the same page context.

* feat: Browser Use best practices — click/type/state improvements

Inspired by deep analysis of Browser Use's design patterns:

1. Framework listener detection (React/Vue/Angular)
   - Detect __reactProps$ onClick, Vue _vei, Angular ng-reflect-click
   - Catches <div onClick> elements that pure ARIA/tag heuristics miss

2. Click CDP fallback
   - clickJs() now returns coordinates on failure
   - BasePage.click() falls back to CDP Input.dispatchMouseEvent
   - Page.clickWithQuads() uses DOM.getContentQuads for inline elements

3. Type improvements
   - React-compatible: use native HTMLInputElement.prototype.value setter
   - Contenteditable: selectAll + execCommand('insertText') for rich editors
   - Autocomplete: detect role=combobox, wait 400ms for dropdown suggestions

4. getContentQuads precise click
   - Page.clickWithQuads() for multi-line inline elements (e.g. wrapped <a>)
   - Falls back through getContentQuads → getBoxModel → JS click

* fix: address code review — injection, silent failure, setter prototype

1. clickWithQuads: escape ref with JSON.stringify before inserting into
   JS strings and CSS selectors (injection risk)
2. base-page click: throw error when both JS click and CDP fallback fail
   instead of silently succeeding
3. typeTextJs: use matching prototype for native setter
   (HTMLTextAreaElement for textarea, HTMLInputElement for input)
2026-04-03 02:38:49 +08:00
sline b8f1abc3a1 fix(twitter): use search input for SPA navigation instead of pushState (#695)
* fix(twitter): add search input fallback for intermittent SPA navigation failures

The pushState + popstate approach works in most environments but fails
intermittently for some users (see #690), likely due to Twitter A/B
tests or timing race conditions where the pathname hasn't updated when
checked.

This commit adds a fallback strategy: when pushState fails after 2
retries, we type the query into the search input on /explore and press
Enter. This triggers Twitter's own form handler, performing SPA
navigation without a full page reload (keeping the fetch interceptor
alive).

Both strategies use selector-based waiting ([data-testid="primaryColumn"])
rather than fixed delays, with graceful fallthrough on timeout.

Fixes #690

* test(twitter): update search test for fallback evaluate call

The search input fallback adds one extra evaluate() call when pushState
fails. Update the mock chain and assertion count accordingly.

* fix(twitter): guard nativeSetter and add fallback success test

- Add optional chaining on getOwnPropertyDescriptor().set to handle
  edge cases where Twitter's sandbox overrides the HTMLInputElement
  prototype.
- Add test case covering the full fallback path: pushState fails twice,
  search input fallback succeeds, results are returned correctly.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-02 22:27:48 +08:00
jakevin bb137ce901 feat: add opencli operate — browser control commands for Claude Code skill (#614)
Add `opencli operate` subcommand group with 15+ commands for
step-by-step browser control, designed as a Claude Code skill.
No LLM API key needed — Claude Code IS the LLM.

Commands:
  Navigation: open, back, scroll
  Inspect: state, screenshot, get (title/url/text/value/html/attributes)
  Interact: click, type, select, keys
  Wait: wait selector/text/time
  Extract: eval (execute JS in page context)
  API Discovery: network (auto-captured since last open, --detail N)
  Sedimentation: init (generate adapter scaffold), verify (test adapter)
  Session: close

Infrastructure:
  - CDP passthrough with 22-method allowlist
  - Two-layer retry for extension interference (aggressive for operate:*)
  - Network interceptor auto-injected on operate open
  - node_modules symlink for user TS adapter imports

Skill: skills/opencli-operate/SKILL.md
  - Complete command reference
  - Sedimentation workflow guide (explore → network → init → verify)
  - Adapter strategy guide (PUBLIC/COOKIE/UI)
  - Dual quickstart (AI Agent 1 step / Human 3 steps)
2026-04-02 19:30:35 +08:00
gucasbrg 7d7203891f fix(twitter): resolve article ID to tweet ID before GraphQL query (#688)
* fix(twitter): resolve article ID to tweet ID before GraphQL query

Article URLs (x.com/i/article/{articleId}) use a different ID than
tweet status URLs. The GraphQL TweetResultByRestId endpoint requires
the parent tweet ID, not the article ID.

Fix: navigate to the article page first, extract the associated tweet
ID from DOM links, then use that for the GraphQL query.

Fixes article fetching returning "Article not found" for all article URLs.

* fix: distinguish article URLs from status URLs, add explicit error handling

The previous commit routed all inputs through the article page, breaking
status URL and bare ID flows. Now only article URLs trigger the
article→tweet ID resolution. Status URLs and bare IDs keep the original
behavior. Also throws an explicit error if resolution fails instead of
silently falling back to the article ID.

---------

Co-authored-by: buruguo <buruguo@lambdafintech.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-02 18:33:51 +08:00
AstroHan 081efe37f7 fix(xiaohongshu): clarify empty note shell hint (#686)
* fix(xiaohongshu): clarify empty note shell hint

* fix(xiaohongshu): simplify empty shell detection to title+author check

The 7-field conjunction was overly strict — a note that rendered only
placeholder metrics but no title/author was still a valid empty shell.
Since title and author are always present on real notes, checking just
those two fields is a more reliable and simpler signal.

---------

Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-02 18:33:36 +08:00
jakevin 777b882040 refactor: centralize daemon transport client (#692) 2026-04-02 18:28:50 +08:00
fii6 abd46bccba feat(gemini): add Gemini web adapter with minimal output (#619)
* feat(gemini): add web adapter with minimal output

* fix(gemini): use defaultFormat for minimal output

* fix(gemini): preserve full transcript responses

* docs(gemini): add browser adapter guide

* review: wire gemini into adapter indexes

---------

Co-authored-by: fii6 <246637913+fii6@users.noreply.github.com>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-02 13:39:35 +08:00
GanFanNewOrder d721eb6c6c feat(amazon): add browser adapter and docs (#659)
* feat(amazon): add browser adapter and docs

* review: wire amazon into discovery docs

---------

Co-authored-by: 泽加武 <zejiawu@zejiawudeMac-mini.local>
Co-authored-by: jackwener <jakevingoo@gmail.com>
2026-04-02 13:32:23 +08:00
ajia1206 341bb87e09 test(xiaohongshu): redact creator fixture data (#647) 2026-04-02 01:21:53 +08:00
AstroHan 127dc3edea feat: add minimal record write candidates (#665) 2026-04-02 01:21:11 +08:00