mirror of
https://github.com/jackwener/OpenCLI.git
synced 2026-09-14 18:25:42 +08:00
51a9456305
* feat(linkedin): add people-search command (#1621) Closes #1621. Adds opencli linkedin people-search <keywords> for finding people on standard LinkedIn (not Sales Navigator). Architecture note. Standard LinkedIn moved its people search results page to Server-Driven UI / React Server Components on the /flagship-web/rsc-action/... path stack. The legacy Voyager REST endpoint /voyager/api/search/dash/clusters returns HTTP 500 from a web context; its modern camelCase rename voyagerSearchDashClusters returns the same. The result list is rendered server-side and the page HTML IS the result payload; Voyager calls from the page are sidebar / notification concerns, not search results. Extraction strategy. LinkedIn SSR uses obfuscated CSS class hashes (e.g. _997b7c77) that rotate on every deploy AND display:contents wrappers that flatten the DOM tree. Class-based selectors, walk-up- to-card logic, and anchor-pair element ranges all fail because no element boundary matches a person's card. Working approach: extract main.innerText once, split by newline, slice between consecutive person names. The names come from the aria-hidden spans of /in/<handle> anchors. LinkedIn's SSR emits a card as a name line followed by degree badge / headline / location / action labels before the next card's name line - a layout that has been stable through several DOM refactors. Critical filter: /in/<handle> anchors over-count because LinkedIn renders each mutual connection as a /in/ anchor inside another card's result. The skip() predicate during name-line lookup drops mutual-connection lines ("X, Y and N other mutual connections"), so anchors that don't have a real name line are filtered out. CUL caveat. LinkedIn imposes a monthly Commercial Use Limit on people search against the standard site. Burst behaviour is irrelevant - the limit is a calendar-month counter. The adapter runs one navigation per invocation (no pagination) so a single call costs exactly one CUL query. --limit is capped at 10 to keep a single call's information density high without surfacing the "reached commercial use limit" yellow banner faster. Schema: rank, name, headline, location, profile_url Live verified against kyfw 12306-style throttled cadence (sleep 60s between dev iterations to keep CUL consumption visible): 5/5 rows populated with name + headline + location + profile_url for the keyword "reinforcement learning". Mutual-connection anchors correctly filtered out so the row order matches LinkedIn's own ranking. Tests: 10 unit tests covering URL construction, limit validation, extraction-script invariants (anchor enumeration, text-slice approach, mutual-connection filter, aria-hidden span as name source), limit slicing, AuthRequiredError on missing JSESSIONID, CUL- flavoured CommandExecutionError on redirect, EmptyResultError on zero rows, ArgumentError on empty keywords, and registry shape. * fix(linkedin): harden people search typed boundaries * fix(linkedin): fail people search candidate parser drift --------- Co-authored-by: jackwener <jakevingoo@gmail.com>