mirror of
https://github.com/jackwener/OpenCLI.git
synced 2026-09-14 18:25:42 +08:00
4475d4efe3
Round 21 follow-up to #1400 (P0 write-action symmetry, merged `644d4517`). 5 features + help docs unified into one PR per WAWQAQ "全部合成一个 PR" directive. ## Scope - **P1** (`cf10c098`): `twitter search` `--from / --has / --exclude / --product` filters, mapping to X `from:` / `filter:` / `-filter:` / `f=` operators; legacy `--filter top|live` preserved (--product win on conflict) - **P2** (`a484a69a`): new `twitter bookmark-folders` + `bookmark-folder <id>`; X Premium GraphQL `bookmarkFoldersSlice` + `BookmarkFolderTimeline`; queryId 三层 fallback (placeholder.json → client-web bundle → pinned constants) - **P3** (`f209f914`): `--top-by-engagement N` to 7 tweet-shaped read commands (search/timeline/likes/bookmarks/list-tweets/tweets/thread); single helper in `utils.js`; formula `likes×1 + retweets×3 + replies×2 + bookmarks×5 + log10(views+1)×0.5`; **N=0 reference equality no-op** → existing 157 twitter tests 0 churn - **P4** (`89283fa0`): `TWITTER_BEARER_TOKEN` + composer image helpers extracted to `utils.js` (12 GraphQL adapter dedup); reply hardening; quote adds `--image` - **P5** (`a3d10a48`): sibling article-scope helper extracted to `shared.js` (9 write commands reuse, dedup with #1400 P0 invariant) - **docs** (`2a358d80`): help-doc precision (positional-omitted defaults + download/bookmarks/notifications/timeline/lists description thicken; concurrent #1401/#1403 wording preserved) 47 files / +2594/-470. Tests **96 → 216 (+120)**, manifest 798 → 801 (+3), typed-error-lint 190 → 189 (resolved 1 grandfathered sentinel). ## Iteration history (3 review fix commits on top of 6 author commits) - `7f93779b` — codex-mini1 lead fix1: 3 blocker bundle (P5 host invariant + P2 safe-id + sentinel removal + P1 fallback fail-fast) - `2a29ecc6` — codex-mini1 lead fix2: P3 help formula consistency (doc/help text matches actual `log10(views+1)×0.5`) - `df4dcd76` — codex-mini1 lead fix3 (F-P-1 aux catch): P2 `bookmark-folder --limit` upfront validation (`Number(kwargs.limit ?? 20)` + reject non-positive/non-integer + regression `0/negative/fractional/NaN` + `page.goto` zero-call assert) ## 4 progressive blockers caught (codex-mini1 lead 3 rounds + F-P-1 aux 1 round) 1. **P5 host invariant gap** (lead): article-scope helper preserved exact `/status/<id>` path but ignored link host → off-domain `https://evil.com/alice/status/<target>` would satisfy `__twHasLinkToTarget`. Fixed: `https` + X/Twitter host or subdomain + exact `/status/<id>` or `/i/status/<id>` path; query/hash allowed; off-domain/host-suffix/non-https/path-suffix/substring-id rejected; JSDOM positive + 5 negative anchors. 2. **P2 listing→detail round-trip + sentinel** (lead): `bookmark-folders` accepted opaque IDs but `bookmark-folder <id>` only accepted numeric → round-trip broken; new `author: 'unknown'` sentinel created fabricated author URL. Fixed: `[A-Za-z0-9_-]+` opaque safe-id (rejects `/`, `?`, `%`, spaces) + `resolveTwitterQueryId()` sanitization for queryId resolution; sentinel removed → empty author + canonical `/i/status/<id>` URL. 3. **P1 fallback silent tab miss** (lead): pushState fail → fallback typing into search box, `clickProductTabIfNeeded()` silent return on tab not found → user `--product photos` silently degraded to Top results. Fixed: throw `CommandExecutionError` when requested `--product` tab cannot be selected + invalid `--from` / `--limit` upfront pre-nav reject + double-direction tests. 4. **P2 limit silent normalize** (aux): `const limit = kwargs.limit || 20` → `--limit 0` silent → 20; negative/non-integer pre-IO unchecked. Fixed: `Number(kwargs.limit ?? 20)` + require positive integer before `page.goto` + regression covers `0/negative/fractional/NaN` + `page.goto` zero-call. ## Cultural sediment (Round 21 audit checklist 7 rules / 6 dimensions) This PR **immediately validated 4 of 7 rules** in review pipeline: - (b) silent-clamp class — P1 fallback silent tab miss (silent semantic-downgrade) + P2 `|| 20` silent normalize - (e) ID exact-not-substring — P5 host invariant (was only path-exact, not host-exact) - (f) grandfathered-not-exempt — P5 helper-refactor boundary lost host invariant + P2 new adapter inherited grandfathered `'unknown'` sentinel - (g) fallback-must-have-success-criterion — P1 fallback path missing post-condition assertion 7 rules / 6 dimensions: - (a) cross-grep sibling URL pattern — structural - (b) silent-clamp class — failure mode (input) - (c) broad querySelector → article-scoping — scope - (d) missing-validation early reject — boundary - (e) ID exact-not-substring — identity - (f) grandfathered-not-exempt (corollary: applies to new file + new helper-refactor boundary; not original-file line-edit) — time-axis - (g) fallback-must-have-success-criterion (sub-rule g': fallback unit test must include post-condition assertion, not just "doesn't throw") — failure mode (output) **Cross-PR validation 4-chain on meta-anchor "Structural exactness for identity matching"**: - #1391 URL layer (`isFacebookAuthRedirectPath`: top-level anchor + `\.php` + `(/|$)` segment edge) - #1392 URL parser layer (`parseGrokSessionId`: bare UUID exact / URL host-exact-or-subdomain + path-exact) - #1400 DOM layer (article-scoping: status-id `/\/status\/${id}(?:\/|$)/` regex / segment-array exact) - #1406 P5 helper-refactor boundary (full URL invariant in shared helper: host+path re-anchored after extraction) - Common invariant: boundary-lock structural shape; **fuzzy match is silent-failure 温床**; lesson lifecycle = surface-shift not add-and-forget. **Audit framework self-discipline**: each rule must have grep-able detection signal, otherwise rule degenerates to mantra. Framework is "7 rules + sub-instance pattern in new surface", not frozen 7 rules. **Round 17 race-mitigation 第 9 连续 race-free execution**: standard alternation cadence (#1400 A 组 → #1406 B 组), lead final + aux final + `@pr-monitor squash?` trigger, pr-monitor proactive ack + serial squash, lead silent on closeout. ## Validation gates (final head `df4dcd76`) Local: Twitter adapter tests `25 files / 216 tests`, focused P1/P2/P3/P5 tests `99/99`, `node --check` touched runtime, `npx tsc --noEmit`, `npm run build`, manifest 801 entries, typed-error-lint `189/189`, silent-column-drop `103/103`, doc-coverage `140/140`, docs:build clean, listing-id advisory `13` unchanged (wikipedia/trending residual non-Twitter), `git diff --check` clean. GitHub: build×3 (ubuntu/macos/windows) SUCCESS, unit-test shards SUCCESS, bun-test SUCCESS, adapter-test SUCCESS, audit SUCCESS, doc-coverage SUCCESS, docs-build SUCCESS, smoke-test skipped, PR `CLEAN/MERGEABLE`. Reviewers: - Lead: @codex-mini1 (3 fix rounds, all caught proactively + amend P3 help consistency) - Aux: @First-principles-1 (better-solution triangulation on P2 queryId 三层 fallback + P5 invariant + P3 N=0 reference no-op + caught P2 limit silent normalize) - Author: @opencli-user (5-feature scope + 7-rule sediment co-author + corollary contributor)
150 lines
6.1 KiB
JavaScript
150 lines
6.1 KiB
JavaScript
import { ArgumentError } from '@jackwener/opencli/errors';
|
|
|
|
const QUERY_ID_PATTERN = /^[A-Za-z0-9_-]+$/;
|
|
const TWEET_PATH_PATTERN = /^\/(?:[^/]+|i)\/status\/(\d+)\/?$/;
|
|
const TWEET_HOSTS = new Set(['x.com', 'twitter.com']);
|
|
|
|
function isTwitterHost(hostname) {
|
|
return TWEET_HOSTS.has(hostname)
|
|
|| hostname.endsWith('.x.com')
|
|
|| hostname.endsWith('.twitter.com');
|
|
}
|
|
|
|
export function parseTweetUrl(rawUrl) {
|
|
const value = String(rawUrl ?? '').trim();
|
|
if (!value) {
|
|
throw new ArgumentError('twitter tweet URL cannot be empty', 'Example: opencli twitter retweet https://x.com/jack/status/20');
|
|
}
|
|
let parsed;
|
|
try {
|
|
parsed = new URL(value);
|
|
}
|
|
catch {
|
|
throw new ArgumentError(`Invalid tweet URL: ${value}`, 'Use a full https://x.com/<user>/status/<id> URL');
|
|
}
|
|
const hostname = parsed.hostname.toLowerCase();
|
|
if (parsed.protocol !== 'https:' || !isTwitterHost(hostname)) {
|
|
throw new ArgumentError(`Invalid tweet URL host: ${value}`, 'Use a full https://x.com/<user>/status/<id> URL');
|
|
}
|
|
const match = parsed.pathname.match(TWEET_PATH_PATTERN);
|
|
if (!match?.[1]) {
|
|
throw new ArgumentError(`Could not extract tweet ID from URL: ${value}`, 'Use a full https://x.com/<user>/status/<id> URL');
|
|
}
|
|
return {
|
|
id: match[1],
|
|
url: parsed.toString(),
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Build a JS source fragment that, when embedded inside a `page.evaluate(...)`
|
|
* IIFE, declares browser-side helpers for scoping operations to a specific
|
|
* tweet by status id. Sibling adapters historically inlined ad-hoc article
|
|
* lookups that either (a) skipped scoping entirely (silent: act on first
|
|
* matching button on a conversation page) or (b) used substring matches like
|
|
* `pathname.includes('/status/' + tweetId)` (silent: `/status/123` matches
|
|
* `/status/1234567`). This helper centralises the canonical pattern so all
|
|
* write-actions reuse the same exact-match guard.
|
|
*
|
|
* Declared bindings (available to the embedding IIFE):
|
|
* - `tweetId` : the requested status id (string)
|
|
* - `__twGetStatusIdFromHref(href)` : extract status id from a link href, or null
|
|
* - `__twHasLinkToTarget(root)` : true iff `root` contains any link to tweetId
|
|
* - `findTargetArticle()` : the <article> matching tweetId, or undefined
|
|
*/
|
|
export function buildTwitterArticleScopeSource(tweetId) {
|
|
return `
|
|
const tweetId = ${JSON.stringify(tweetId)};
|
|
const __twTweetPathRe = /^\\/(?:[^/]+|i)\\/status\\/(\\d+)\\/?$/;
|
|
const __twIsTwitterHost = (hostname) => hostname === 'x.com'
|
|
|| hostname === 'twitter.com'
|
|
|| hostname.endsWith('.x.com')
|
|
|| hostname.endsWith('.twitter.com');
|
|
const __twGetStatusIdFromHref = (href) => {
|
|
try {
|
|
const parsed = new URL(href, window.location.origin);
|
|
if (parsed.protocol !== 'https:' || !__twIsTwitterHost(parsed.hostname.toLowerCase())) {
|
|
return null;
|
|
}
|
|
return parsed.pathname.match(__twTweetPathRe)?.[1] || null;
|
|
} catch {
|
|
return null;
|
|
}
|
|
};
|
|
const __twHasLinkToTarget = (root) => Array.from(root.querySelectorAll('a[href*="/status/"]'))
|
|
.some((link) => __twGetStatusIdFromHref(link.href) === tweetId);
|
|
const findTargetArticle = () => Array.from(document.querySelectorAll('article'))
|
|
.find(__twHasLinkToTarget);
|
|
`;
|
|
}
|
|
|
|
export function sanitizeQueryId(resolved, fallbackId) {
|
|
return typeof resolved === 'string' && QUERY_ID_PATTERN.test(resolved) ? resolved : fallbackId;
|
|
}
|
|
export async function resolveTwitterQueryId(page, operationName, fallbackId) {
|
|
const resolved = await page.evaluate(`async () => {
|
|
const operationName = ${JSON.stringify(operationName)};
|
|
const controller = new AbortController();
|
|
const timeout = setTimeout(() => controller.abort(), 5000);
|
|
try {
|
|
const ghResp = await fetch('https://raw.githubusercontent.com/fa0311/twitter-openapi/refs/heads/main/src/config/placeholder.json', { signal: controller.signal });
|
|
clearTimeout(timeout);
|
|
if (ghResp.ok) {
|
|
const data = await ghResp.json();
|
|
const entry = data?.[operationName];
|
|
if (entry && entry.queryId) return entry.queryId;
|
|
}
|
|
} catch {
|
|
clearTimeout(timeout);
|
|
}
|
|
try {
|
|
const scripts = performance.getEntriesByType('resource')
|
|
.filter(r => r.name.includes('client-web') && r.name.endsWith('.js'))
|
|
.map(r => r.name);
|
|
for (const scriptUrl of scripts.slice(0, 15)) {
|
|
try {
|
|
const text = await (await fetch(scriptUrl)).text();
|
|
const re = new RegExp('queryId:"([A-Za-z0-9_-]+)"[^}]{0,200}operationName:"' + operationName + '"');
|
|
const match = text.match(re);
|
|
if (match) return match[1];
|
|
} catch {}
|
|
}
|
|
} catch {}
|
|
return null;
|
|
}`);
|
|
return sanitizeQueryId(resolved, fallbackId);
|
|
}
|
|
/**
|
|
* Extract media flags and URLs from a tweet's `legacy` object.
|
|
*
|
|
* Prefers `extended_entities.media` (superset with full video_info) and falls
|
|
* back to `entities.media` when the extended form is missing. For videos and
|
|
* animated GIFs, returns the mp4 variant URL; for photos, returns
|
|
* `media_url_https`.
|
|
*/
|
|
export function extractMedia(legacy) {
|
|
const media = legacy?.extended_entities?.media || legacy?.entities?.media;
|
|
if (!Array.isArray(media) || media.length === 0) {
|
|
return { has_media: false, media_urls: [] };
|
|
}
|
|
const urls = [];
|
|
for (const m of media) {
|
|
if (!m) continue;
|
|
if (m.type === 'video' || m.type === 'animated_gif') {
|
|
const variants = m.video_info?.variants || [];
|
|
const mp4 = variants.find((v) => v?.content_type === 'video/mp4');
|
|
const url = mp4?.url || m.media_url_https;
|
|
if (url) urls.push(url);
|
|
} else {
|
|
if (m.media_url_https) urls.push(m.media_url_https);
|
|
}
|
|
}
|
|
return { has_media: urls.length > 0, media_urls: urls };
|
|
}
|
|
export const __test__ = {
|
|
sanitizeQueryId,
|
|
extractMedia,
|
|
parseTweetUrl,
|
|
buildTwitterArticleScopeSource,
|
|
};
|