* feat(convention): listing↔detail id pairing rule + CI gate Adds a hard convention: when a site exposes both a listing-class command (search / hot / top / recent / ...) and a detail-class command (read / paper / article / view / ...), every listing row MUST surface an id-shaped column whose value round-trips into the detail command. Without that, an agent has no way to follow up on a listing row except re-searching by title or scraping URLs out of band — both of which break the agent-native contract. What's in this PR - docs/conventions/listing-detail-id-pairing.md — full rule, examples table, why-it-matters, what counts as id-shaped, exemption taxonomy, how to add an id column to a listing. - scripts/check-listing-id-pairing.mjs — validator that reads cli-manifest.json, classifies each entry as listing / detail / other, and fails when a listing on a site that also has a read-detail command is missing an id-shaped column. Exemption allowlist records WHY each pair is exempt so future maintainers know what to verify. - npm run check:listing-id-pairing — strict-mode wrapper. - CI: new step in build job runs the validator after the manifest freshness check on Linux. - docs/developer/ts-adapter.md — cross-link from the adapter authoring guide. - docs/.vitepress/config.mts — sidebar entries for the new conventions section. Fixes brought to zero violations - 1688/search: add offer_id (already extracted, just surfaced) - bluesky/user: add uri (AT URI round-trips into bluesky/thread) - tieba/search: add id + url (thread_id already extracted) - tieba/hot: add url (rows are topics, not threads — url is the best-effort round-trip handle, doc'd as such) Exemptions (intentional, doc'd in EXEMPT map with rationale) - nowcoder/hot, bluesky/trending, twitter/trending — listing rows are topic strings, not posts. - lesswrong/user, reddit/user — rows are profile-attribute key/value pairs, addressed by the username arg. - discord-app/search — desktop UI session, message ids not extractable. - notion/search — Strategy.UI Quick Find, page ids not exposed in DOM. Validator output after this PR: 32 sites scanned, 75 listings checked, 7 exempted, 0 violations. * fix(convention): tighten listing id gate * fix(convention): close url-derived id loophole
4.6 KiB
TypeScript Adapter Guide
Use TypeScript adapters when you need browser-side logic, multi-step flows, DOM manipulation, or complex data extraction that goes beyond simple API fetching.
Basic Structure
import { cli, Strategy } from '@jackwener/opencli/registry';
import { CommandExecutionError, EmptyResultError } from '@jackwener/opencli/errors';
cli({
site: 'mysite',
name: 'search',
description: 'Search MySite',
access: 'read', // 'read' | 'write'
domain: 'www.mysite.com',
strategy: Strategy.COOKIE, // PUBLIC | COOKIE | HEADER
args: [
{ name: 'query', required: true, help: 'Search query' },
{ name: 'limit', type: 'int', default: 10, help: 'Max results' },
],
columns: ['title', 'url', 'date'],
func: async (page, kwargs) => {
const { query, limit = 10 } = kwargs;
// Navigate and extract data
await page.goto('https://www.mysite.com');
const data = await page.evaluate(`
(async () => {
const res = await fetch('/api/search?q=${encodeURIComponent(String(query))}', {
credentials: 'include'
});
return (await res.json()).results;
})()
`);
if (!Array.isArray(data)) throw new CommandExecutionError('MySite returned an unexpected response');
if (!data.length) throw new EmptyResultError('mysite search', 'Try a different keyword');
return data.slice(0, Number(limit)).map((item: any) => ({
title: item.title,
url: item.url,
date: item.created_at,
}));
},
});
Access Metadata
Every adapter must declare access: 'read' | 'write'.
- Use
readwhen the command only retrieves data from the target product or account. - Use
writewhen the command changes remote product/account state, such as sending messages, publishing, liking, following, buying, deleting, creating remote assets, or starting paid/credit-consuming generation. downloadandexportcommands arereadwhen they only read remote data and write local files; local filesystem writes are a separate permission dimension.
Listing↔Detail ID Pairing
If your site exposes both a listing-class command (search / hot / top /
recent / ...) and a detail-class command (read / paper / article /
post / view / ...), every listing row MUST surface an id-shaped column
that round-trips into the detail command's positional arg. Without that, an
agent can't follow up on a row without re-searching by title or scraping a
URL out of band.
The CI gate npm run check:listing-id-pairing fails when a listing is
missing its id column. See Listing↔Detail ID Pairing
for the full rule, exemption rationale, and how to add an id to an existing
listing.
Strategy Types
| Strategy | Constant | Use Case |
|---|---|---|
| Public | Strategy.PUBLIC |
No auth needed |
| Cookie | Strategy.COOKIE |
Browser session cookies |
| Header | Strategy.HEADER |
Custom headers/tokens |
The page Object
The page parameter provides browser interaction methods:
page.goto(url)— Navigate to a URLpage.evaluate(script)— Execute JavaScript in the page contextpage.waitForSelector(selector)— Wait for an elementpage.click(selector)— Click an elementpage.type(selector, text)— Type text into an input
The kwargs Object
Contains parsed CLI arguments as key-value pairs. Always destructure with defaults:
const { query, limit = 10, format = 'json' } = kwargs;
For most search/read/detail commands, the main subject should be positional (opencli mysite search "rust", opencli mysite article 123) instead of a named flag such as --query or --id. Keep named flags for optional modifiers.
Error Handling
Prefer throwing CliError subclasses from src/errors.ts for expected adapter failures:
AuthRequiredErrorfor missing login / cookiesEmptyResultErrorfor empty but valid responsesCommandExecutionErrorfor unexpected API or browser failuresTimeoutErrorfor site timeoutsArgumentErrorfor invalid user input
Avoid raw Error for normal adapter control flow. This keeps top-level CLI output consistent and preserves hints for users.
AI-Assisted Development
Use the opencli-adapter-author skill plus the opencli browser * primitives to scaffold and verify adapters end-to-end:
# Recon on the target site
opencli browser open https://example.com
opencli browser network
opencli browser state
# Scaffold + verify
opencli browser init mysite/trending
opencli browser verify mysite/trending
See AI Workflow for the full loop and the adapter-author skill for the step-by-step runbook.