Commit Graph

207 Commits

Author SHA1 Message Date
木杉 6606594068 fix(suggest): surface both halves of a welded compound flag name (#2604)
`suggest.Closest` ranks flag/command suggestions by shared prefix then edit
distance. A hallucinated name welded from two real names -- e.g. `--sql-file`
from the real `--sql` and `--file` of `apps +db-execute` -- defeats both
signals: the leading half wins on prefix, and the trailing half (`file`, 4
edits from `sql-file`, budget 2) is dropped. The hint then names the flag the
caller did not want and omits the one that does exactly what they asked for.

The cost is not the rejected call. Steered to `--sql`, callers inline SQL
through the shell, where quoting mangles `DEFAULT ''` and
`current_setting(...)` into syntax errors that read as SQL-authoring bugs.
`--file` passes file contents verbatim and avoids that class entirely.

Treat a candidate that exactly equals one hyphen-delimited segment of the typed
name as plausible, however far the whole string drifted. Ranking is unchanged --
segment hits are admitted, not promoted -- so the leading segment still ranks
first, and an unrelated candidate list still yields no suggestions.
2026-09-03 11:33:52 +08:00
sang-neo03 59f6ad4900 fix(output): preserve non-data payloads in the api success envelope (#2601)
* fix: preserve non-data payload keys in SuccessEnvelopeData

When an API response uses a non-"data" key (e.g., /bot/v3/info returns
payload under "bot"), the previous implementation discarded the payload
and returned an empty object. Fall back to the envelope minus transport
fields (code, msg, data) so the business payload is preserved.

Fixes #2428

(cherry picked from commit a82585718c)

* fix(output): pass non-object bodies through and pin api envelope at command level

Follow-up to the cherry-picked fix for #2428: return nil bodies as {} and
non-object bodies untouched instead of collapsing them, align the new test
with its neighbours, and add cmd/api regression tests through the httpmock
path so the user-visible envelope is pinned where the bug was reported.

* test(output): pin null data with sibling payload, drop SDK-pinned array test

TestApiCmd_NonObjectBody_FailsLoudly asserted the SDK's pre-decode rejection
of non-object bodies, not the output-layer branch this PR added; reverting
that branch left it green. The unit test already covers the branch, so the
command-level copy is removed. Add a unit case for {"data": null, "bot": {..}},
which the previous code collapsed to {} and now returns as {"bot": {..}}.

* fix(output): normalize legacy bot payloads

---------

Co-authored-by: Wu Shuwen <108231307+dajiaohuang@users.noreply.github.com>
Co-authored-by: sang-neo03 <266690410+sang-neo03@users.noreply.github.com>
2026-09-03 01:14:22 +08:00
dc-bytedance d5148a88df fix: honor requiredScopes conjunction in CollectScopesForProjects (#1878) 2026-09-02 18:11:22 +08:00
sang-neo03 d12b39cf46 feat(vfs): allow absolute paths under a built-in path policy (#2580)
* feat(vfs): allow absolute paths under a built-in path allowlist

Path flags only accepted paths relative to the working directory, so an
agent passing a full path (typically under /tmp) failed on its first call
and had to retry with a relative one.

Absolute paths are now accepted when they resolve inside a built-in
allowlist: the working directory, /tmp, and ~/files. A built-in denylist
covers system and credential locations and wins over the allowlist,
including over the working directory. Both lists are compiled in and read
no environment variable, flag, or config file, so the effective policy is
fixed by the binary; upgrading is all it takes for the new behavior to
apply.

Containment is decided by file identity (device and inode) alongside the
resolved name, because a single directory has many spellings: APFS folds
U+017F onto "s", so ".sshh" spelled with it opens ~/.ssh, and NTFS and
APFS both compare case-insensitively.

Reads are hardened where the policy applies: O_NOFOLLOW pins the final
component, O_NONBLOCK keeps a FIFO from blocking before it can be
refused, and the opened descriptor is matched against the inspected
object, rejected when it is not a regular file, and rejected when it
carries extra hard links. The relaxed local-input tier used by apps
upload keeps its own contract (symlinks are legitimate arguments there)
and gains the denylist check instead.

Two behaviors are deliberate rather than incidental. Working inside a
denylisted directory now refuses even relative paths, since the denylist
is unconditional. Running as root leaves only the working directory and
/tmp, because the home directory is then /root, itself a deny root.

Existing tests asserted the old "every absolute path is refused"
baseline; they now assert the allowlist. Traversal fixtures escape to the
filesystem root, which stays outside every allowed root on Linux, where
the temp directory that hosts t.TempDir() is /tmp itself.

* fix(vfs): close two paths around the built-in denylist

A "~/..." argument had two readings: validation expanded it to the home
directory, while a caller that keeps the original string — SafeLocalFlagPath
returns it verbatim — opens whatever "~" names in the working directory. A
symlink there carried reads past the denylist, confirmed by reading
/etc/passwd through it. Every interpretation of an argument is now checked,
so the shorthand still reaches ~/files while the literal entry cannot
escape.

With no LARKSUITE_CLI_CONFIG_DIR and no reachable home directory,
core.GetBaseConfigDir keeps credentials in a bare ".lark-cli" resolved
against the working directory, which is an allow root. That fallback is now
mirrored as a deny root, so containers whose home lookup fails do not expose
their stored tokens.

* fix(vfs): enforce hard-link checks across readers

* fix(vfs): stop an output hard link from rewriting a file outside the allowlist

A hard link has no target for name resolution to follow, so a link inside an
allowed root looked like an allowed destination while sharing its inode with a
file outside every root. A caller that truncated the approved name in place
rewrote that outside file: `auth qrcode --output <link>` reported success and
replaced a 43-byte JSON file outside the allowlist with its PNG.

Output validation now refuses an existing target that carries more than one
name, which covers callers that write directly, and auth qrcode commits
through a temp file and a rename, which replaces the directory entry and
leaves the other names alone. Writers already going through FileIO.Save were
never affected, since that path has always committed by rename.

* fix(vfs): give the hard-link refusal a workable recovery hint

The message told the caller to copy the file into an allowed directory, which
answers a question they did not ask: the file that triggers this is normally
already inside one, with every one of its names there too. It now states what
the check actually cannot do — enumerate the other names a file is reachable
by — and offers the step that works, which is to copy the file and use the
copy.

* test(vfs): pick the denylist fixture for the platform under test

Two tests reached for "/etc/passwd" as a denylisted absolute path. That path
is not absolute on Windows, so one test met the foreign-path rejection instead
of the denylist it was asserting, and the other saw the path joined to the
working directory and no rejection at all. Both now ask for a deny root that
exists on the platform running them — the credential directories under the
account home qualify everywhere — which keeps the denylist covered on Windows
rather than skipping it there.

Verified on Windows 10.0.19045 by running the package's test binary from this
branch and from main: main passed, this branch failed these two, and both pass
after the change. The other packages this branch touches were compared the
same way and their Windows results are identical on both sides.

* fix(vfs): state the hard-link check as the condition it tests

The check read as "bail out unless the target can be inspected", which
nilerr reads as an error swallowed on the way out. It now names the case it
acts on — an existing regular file with more than one name — and the comment
carries what the early return used to imply: a target that cannot be
inspected has no link count to judge, and the write layer reports the real
failure with proper typing.

* docs(vfs): scope the policy's environment claim to what holds

The header promised that neither list accepts runtime input and that no
caller controlling the environment can widen them. Two inputs contradict
that: LARKSUITE_CLI_CONFIG_DIR contributes a deny root, and where the account
database cannot name the running uid, $HOME decides where ~/files points —
reproduced in a container running as an unregistered uid, which wrote into a
directory the environment chose.

The comments now state the preference and its boundary rather than a
guarantee, and record what the boundary costs: a directory named "files"
under the named path, with the home directory itself still outside the
allowlist and every candidate home still carrying the credential deny roots.
The trustedHome note also said the pure-Go lookup falls back to $HOME
silently; it does so only when $USER is set as well, and returns an error
otherwise, which drops the ~/files root instead of moving it.

No behavior change.

* fix(auth): keep the mode of a QR output file that already exists

Committing the QR write by rename fixed a hard link from rewriting a file
outside the allowlist, but it also changed what happens to the target's mode.
A rename installs the temp file's inode, mode included, where the previous
in-place write left the existing file's mode untouched. Overwriting a target
the caller had restricted to 0600 therefore published it as 0644.

The mode now comes from the file already at the path; only a path with
nothing at it takes the default. Verified against main, which preserved 0600
here, and covered by a test that fails when the fixed mode is restored.

* test(sheets): move the csv file-alias tests onto the new path baseline

Merging main brought #2559's tests for the --file → --csv alias, written
against the policy this branch replaces. Two of them fail on it, both because
the verdict they describe moved rather than disappeared.

The out-of-tree case used /tmp, which the allowlist now accepts, so the value
came back as a missing file instead of an out-of-tree one; it now names a path
no allow root can contain. The directory case is refused when the descriptor
is inspected, before a read is attempted, so the message reads "not a regular
file". What the caller sees of both — the flag named, the cause kept, stdin
offered — is unchanged.

That message listed the kinds it refuses and omitted directories, which is how
it reached a directory test reading as a mismatch. It now names them.

* fix(im): let the path policy judge a download target

`+messages-resources-download` refused an absolute --output before the shared
policy saw it, so the flag stayed relative-only after the policy learned to
accept full paths. It is the command behind 99% of a reported 1,189 download
path errors in one week, where 97.2% of first calls passed an absolute path
and every later success had switched to a relative one.

The shape checks are gone. Both call sites already hand the result to
ResolveSavePath, which applies the allowlist, the denylist and symlink
resolution, so refusing a shape here decided nothing the policy would not
decide better — an absolute path is now answered by where it points rather
than by how it is written.

The file-key checks stay, and they are what the batch caller relies on: it
embeds the key in the path, and a key carrying a separator is refused as a
malformed key, so a traversal cannot be built from one. Verified against a
real tenant: /tmp and ~/files now save, while ~/.ssh, /etc and a path outside
every root are still refused.

* test(im): pin the download output contract the policy now decides

The dry-run suite listed an absolute path among the values --output must
refuse. That held while the command rejected the shape itself; now that the
built-in policy decides, /tmp is an allowed root and the path is accepted, so
the case asserted a rule that no longer exists.

It is replaced by the two halves of the real contract: an absolute path
inside an allowed root reaches the request, and a path that resolves outside
every root — a parent escape from this working directory, or a denylisted
directory — is still turned down as a validation error naming --output.

* fix(vfs): hold a relative path to the working directory

Accepting /tmp as an allow root gave a relative path somewhere new to go.
A process whose working directory sits under /tmp — CI runners, containers
and agent sandboxes commonly arrange that — could climb out with "../" and
still satisfy the allowlist, because the sibling it landed in was also under
/tmp. /tmp is world-writable, so that sibling can belong to another user or
another session, and the write side commits by rename, which replaces an
existing target unconditionally. The previous policy refused this: it
required every resolved path to stay under the working directory.

Naming a full path and climbing out of the working directory are different
acts and no longer share one verdict. An absolute path is judged by the
allowlist, which is what this branch set out to allow; a relative one has to
resolve inside the working directory, whatever wider root contains it.

The home denylist grows at the same time and for the same reason: the working
directory is an allow root and running from the home directory is ordinary,
so a credential store there is reachable by a relative name unless the list
covers it. It now names the common ones — netrc, git and shell credentials,
kube, docker, azure, gh, gcloud, the language package registries — and the
shell histories, which carry pasted keys as reliably as a credential file.

---------
2026-09-01 22:18:29 +08:00
Zhibo Wu 20ef0af8a7 fix(event): preserve UTF-8 in truncated diagnostics (#2535)
* fix(event): preserve UTF-8 in truncated diagnostics

* test(event): preserve preflight decode cause coverage
2026-09-01 00:58:18 +08:00
xuzhigang 1181dafc76 feat(im): support rich-text message attachment zone in send/reply/mge… (#2515)
* feat(im): support rich-text message attachment zone in send/reply/mget/edit

Support the post message attachment zone (top-level files array) end to end:
- +messages-send / +messages-reply: repeatable --attachment file_key flags
  merged into the post content's files array (deduplicated).
- +messages-mget: render attachment-zone files/folders as <file>/<folder>
  tags in content, extract file keys for --download-resources.
- +messages-edit: new shortcut (PUT /open-apis/im/v1/messages/:id) with
  --set-attachments / --clear-attachments; body-only edits preserve the
  attachment zone by default.
- Attachment flags are mutually exclusive with --content carrying a files
  array (declare the zone via one or the other, not both).
- bot-only identity, matching server behavior (user token rejected).
- Fixes from review: attachments no longer bypass content mutual-exclusion
  validation (P1); merge dedups by key.
- Docs (SKILL.md, references, affordance) and unit tests updated.

* fix(im): address design-review findings (auto-infer post, dedup set, doc routing)

- --attachment/--set-attachments/--clear-attachments now infer msg_type=post
  automatically; only an explicit incompatible --msg-type conflicts.
  --text is rejected with attachments (text is a standalone message, not a
  post body) with a hint to use --markdown or --content.
- --set-attachments deduplicates repeated keys (docs promised this; the
  replace helper now enforces it).
- Shortcut Description no longer leaks the HTTP path or the raw server
  error phrase; it describes the command semantically.
- affordance/im.md +messages-edit now routes WHEN: interactive cards go to
  messages.patch, corrected messages go to +messages-send, and attachment
  tri-state tips are listed.
- mget doc no longer claims --format json exposes raw wire fields (the
  output is the rendered content); download eligibility clarified.
2026-08-28 16:55:28 +08:00
xiongyuanwen-byted decc9549b5 refactor(sheets): keep the success path off stderr (#2533)
* feat(sheets): reject local-office tokens in +workbook-export

A locally opened Office file (a local_office_ / fake_office_ token, or an
interleaved OFL0X one) names a file the Lark client is showing, not a cloud
document, so the drive export task can only fail on the backend -- and it
fails late, after the create and poll round trips, with an opaque message.

Refuse it up front with a typed failed_precondition that says the workbook is
already a file on disk, and points at +workbook-import for callers who want a
cloud spreadsheet they can export later. The check runs in Validate (so
--dry-run is covered too) and again after the wiki hop in Execute, where the
real spreadsheet token is first known.

* refactor(sheets): report success-path advisories in the result, not on stderr

Every sheets shortcut that had something to say on a successful run said it on
stderr: ignored sub-op locators, emulated dimension semantics, the deprecated
--dimension/--count and +cells-batch-set-style spellings, the dropdown
option-error steer, and the upload/export stage lines in the compatibility
layer. PowerShell's native-command handling and most agent harnesses read
non-empty stderr as failure, so a working call reported itself as an error --
and the facts a caller actually needed sat outside the JSON they parse.

Pure stage text ("Writing image", "Waiting for export task") is deleted: it
duplicates what the result already proves. Everything decision-relevant moves
into the payload:

  - data.warnings           ignored locators, colliding freezes, the dropdown
                            option-error steer (also shown in --dry-run now)
  - data.effective_operation +dim-insert's anchor shift under --inherit-style
                            before, and the whole (rows, cols) state a freeze
                            leaves behind
  - data.deprecation        +cells-batch-set-style and +dim-freeze's legacy
                            flag pair, under a key of its own rather than
                            mixed into warnings
  - data.upload             how +media-upload sent the file

Clean calls keep their exact previous payload shape: every field above is
added only when it has something to report.

Scope is shortcuts/sheets/** on purpose. The remaining success-path stderr in
this domain comes from shared code (the drive export/import core behind
+workbook-export / +workbook-import, the multipart media helper, the auto-grant
helper), which other domains share; cleaning those up belongs to their own
change. The one sheets-owned exception is +workbook-import's extension
correction, which has no slot in the import core's output envelope -- it is
documented at the call site and allowlisted in the guard test.

Tests pin the contract (a successful run leaves stderr empty) and each new
field, plus a source scan that stops new direct ErrOut writes from appearing.

* docs(sheets): point the dropdown option-error warning at data.warnings

The --source-range flag help still told callers the option-error steer arrives
on stderr; it now rides in the result. Mirrors the same edit in the upstream
spec (canonical-spec/spec-tables/flags.json), so the next sync is a no-op.

* fix(sheets): keep export identifiers in +export output, tighten the stderr guard

Review follow-ups on the success-path stderr change:

- +export --output-path lost file_token: on the download branch the token
  reached the caller only through the deleted "Export complete: file_token=…"
  stderr line, and the payload carried just saved_path and size_bytes. Both
  file_token and ticket now ride in the download result, so a caller can
  re-download or resume without re-running the export.

- The stderr guard allowlisted a whole file, hiding any future write in it.
  It now matches one exact statement in one file and asserts that write still
  exists, so both a new write and a stale exception fail the test.

- The contract comments claimed more than the tests prove. They now state
  that only sheets-OWNED code is silent, name the three commands whose noise
  comes from shared implementations (+workbook-export, +workbook-import,
  +media-upload over 20MB), and a new test pins that the shared export core
  does still write -- failing, by design, once that core is cleaned up.

* fix(drive): keep the export and import cores off stderr on success

+workbook-export and +workbook-import delegate to drive.RunExport /
drive.RunImport, so the sheets success-path contract could not hold while
those cores narrated every step: task creation, each poll attempt, completion,
"still in progress", and the import's media upload. Callers that read
non-empty stderr as failure saw a finished export report itself as an error.

The stage text is deleted -- ticket, ready, status, file_token, token and
next_command are all already in the payload. What the narration alone carried
moves into the result:

  - poll               attempts / transient_failures / last_error, added only
                       when a poll actually had to be retried, so a caller can
                       tell a clean run from one that limped to the finish
  - warnings           markdown export falling back to the token as file name
                       after a failed title lookup
  - input_corrections  a caller-supplied record of inputs the CLI rewrote
                       before the request ran; sheets +workbook-import uses it
                       for a mislabeled .xls that is really an .xlsx, which was
                       its last stderr write

drive +export / +import get the same treatment, since they share these cores.
Clean runs keep their exact previous payload shape.

With this, the sheets stderr guard needs no allowlist, and the contract test
covers both workbook commands end to end. Two shared paths a sheets caller can
still reach stay noisy and are named in the contract comment: multipart media
upload over 20MB, and the bot-identity auto-grant warning.

* test(sheets): cover the annotation shapes and both guard call sites

Review follow-ups, all test-side except one comment:

- +dim-insert's effective_operation had no test: a regression could drop the
  emulated-anchor block and still keep stderr empty. Now asserted field by
  field, plus the negative case (--inherit-style after rewrites nothing, so it
  must not gain the block).

- The local-office guard's second call site had no test. A /wiki/ URL only
  reveals its backing token after get_node runs in Execute, so that branch is
  now covered, asserting both the typed rejection and that no export task was
  created.

- annotateSheetsResult's three payload shapes are pinned: object annotated in
  place, array/scalar preserved under `result`, and an empty tool result left
  without an invented `result: null`. The doc comment now spells out that last
  case instead of lumping it in with non-object output.

- The export poll summary test asserted transient_failures but not attempts,
  so a wrong or missing count would have passed.

* fix: preserve recovery state on failure paths and TTY liveness during polls

Review round 2. Removing the success-path narration also removed information
from paths that fail after remote work has started, and removed the only
liveness signal an interactive user had:

- drive +import / sheets +workbook-import: once the import task exists, the
  ticket is the only handle back to it. A poll failure returned bare, so the
  ticket -- previously visible through the polling line -- was lost. It now
  rides on the typed error together with the +task_result command.

- sheets +export --output-path: a download or save failure happens after the
  artifact is ready, so the error now carries ticket, file_token and the
  +export-download command; re-running the whole export is not the recovery.
  A poll timeout carries the ticket for the same reason.

- sheets +batch-update / +batch-chart-*: batch_update is fail-fast without
  rollback, so the ignored-locator and colliding-freeze advisories matter most
  exactly when the call fails part-way -- they decide the safe retry set. They
  are now attached to the typed error's hint as well as the success payload.

- Bounded polls and the import upload are wrapped in RuntimeContext.StartSpinner,
  which is gated on StderrIsTerminal and is a strict no-op for pipes, CI and
  captured output. A human terminal gets liveness back; a machine caller's
  stderr stays empty (the contract tests, which capture stderr, still pass).

+workbook-export's rejection of Office tokens also stopped assuming the caller
holds the file: a local_office_ / fake_office_ prefix means the workbook is
already on their disk, but an interleaved OFL0X token is a file stored in Lark
that may never have been downloaded, so that class is now pointed at
drive +download (then +workbook-import if they want a Lark spreadsheet).

Each behaviour above has a regression test; httpmock's CapturedBodies doc
comment is corrected, since it is appended on every match, not only for
Reusable stubs.
2026-08-28 12:48:19 +08:00
liangshuo-1 8493800ebf feat(config): support keychain-backed tenant access tokens (#2488) 2026-08-25 21:39:50 +08:00
sang-neo03 6952d3aa7f feat(extension): add business command extension v1 (#2308)
* feat(shortcuts): import typed shortcut framework from c07b64621

* feat(extension): add public command contract

* feat(extension): compile and register business commands

* feat(cmd): assemble business command sets

* fix(auth): derive login domains from shortcuts

* test(extension): add business command test runtime

* ci: verify generated Go sources

* test(auth): match help and interactive domains

* fix(extension): validate command path segments

* fix(extension): satisfy generator and error guards

* fix(extension): follow generated domain naming contract

* fix(command): complete extension runtime contracts

* test(command): cover public extension surface

* fix(command): remove unused typed runtime APIs

* fix(command): address follow-up review findings

* fix(command): enforce extension v1 contracts

* fix(command): remove unreachable compatibility wrappers

* fix(auth): keep scope-less domains addressable via --domain, matching main

Move the declared-scope filter from allKnownDomains to the interactive
selector only. On main, a scope-less shortcut domain (event) passes
--domain validation and fails later with "no matching scopes found";
the previous unified filter changed that to "unknown domain" and
dropped it from the --help list. The interactive picker still hides
scope-less domains — selecting one can only fail.

* feat(command): add PathSegment for user-provided path values

Business code concatenates IDs into request paths but had no public
escape helper (internal/validate.EncodePathSegment is unreachable from
extension/command). Mirror its url.PathEscape semantics, use it in the
business command examples, and pin the traversal defense: the validator
decodes percent-encoding before the canonical check, so both raw and
escaped dot sequences fail same-origin validation.

* docs(command): add runnable chat-brief distribution example

Mirror the audit-observer precedent: a buildable wrapper main under
examples/ showing WithCommandSets against the real distribution shape
(plugins, strict mode, and service commands stay enabled). Covers the
single-read command (Validate, shared DryRun request, CallJSON,
PathSegment, Tips) and a Page[T] list command whose pagination flags
come from the compiler. The testdata/wrapper fixture stays test-only.

Verified offline: --help renders tag-driven parameters, +chat-brief-list
exposes --page-all/--page-limit/--page-delay, and --dry-run previews the
request with a fake env token and no network access.

* fix(command): normalize page envelopes and bound CollectAllPages

Two pagination contract fixes from the extension design (owner plan §8.3):

Page decoding accepted only a literal "items" array, so endpoints that
spell their list field differently (drive uses files, some responses use
records) walked every page while decoding nothing — CollectAllPages then
returned an empty set marked complete, and downstream writes ran against
it. Each page now normalizes its single top-level array field into
Page.Items; zero or multiple array fields fail closed with a typed
invalid-response error.

CollectAllPages previously reused the user-facing --page-limit maximum
(1000) as its walk bound. A complete-set collection holds every page in
memory before the workflow's writes run, so it now uses the design's
dedicated workflow bound of 100 pages.

* feat(commandhost): note bounded repetition in Page dry-run previews

A Page[T] command's dry-run can only show the first request; fabricating
response-dependent page tokens is forbidden. Append the bounded-repeat
explanation to the previewed request (preserving any business
description), matching the design's dry-run contract.

Also give the example's list command a --page-token resume flag seeded
into the request, documenting the resume convention: the framework owns
--page-all/--page-limit/--page-delay while the starting cursor is a
business-declared input, independent of --page-all.

* docs(command): generate domain constants with English comments

The generator emitted Chinese titles while the rest of the public
extension packages document in English. Switch the generator to the
"en" service title and regenerate. The generator already rejects a
domain missing either locale, so the switch keeps its own guard.

* docs(command): mark the host adapter read surface

The Host* types, InspectCommand, InspectDomain and CloneSets exist for
lark-cli's host adapter, not for business commands, but nothing said so
at the symbols themselves. They cannot move to a subpackage: a Command
holds its declaration unexported, so a sibling package has no way to
reach it, and moving the wire types to internal/ would cycle back
through CommandMetadata and CommandContext.

Also correct HostPagination, which is not adapter-only -- ContextOptions
and commandtest both carry it.

* refactor(command): type CommandMetadata.Service as DomainName

A set already declares its domain through ExtendDomain(DomainIm), yet
every command repeated the same domain as a bare string that only the
host compiler checked. Typing the field points authors at the generated
enum and forces an explicit conversion when the value comes from a
string variable.

This does not make a mistyped literal a compile error -- an untyped
constant still converts to DomainName -- so the mismatch check in
CompileSets stays the actual net. The example and the wrapper fixture
now declare command.DomainIm.

The chat-brief example was also not gofmt-clean, which the CI format
gate would have caught.

* refactor(command): extract the command-set assembly steps

buildInternalWithConfig is an orchestrator, and the business command
sets were compiled inline inside it. Move that step to
resolveShortcutSnapshot so the entry point reads as one call and the
built-in/external merge has a name.

newCommand carried six near-identical blocks that each nil-checked a
hook and wrapped it in the same type assertion. Split them into one
binder per hook shape; Normalize and Validate now share bindArgsHook
since their signatures match. The behaviour is unchanged: an undeclared
hook still erases to nil, and an empty renderer map still yields nil.

* perf(shortcuts): stop re-cloning an already-isolated snapshot

AllShortcuts deep-copies because a Shortcut carries slice fields whose
backing arrays a shallow copy would share: an external distribution
mutating registered[0].Flags[0] would corrupt the process-global list.
That copy is worth its ~165us over 500+ shortcuts.

Paying it four times per startup is not. auth, schema and the mount path
each cloned the snapshot again, but they receive it from
AllShortcutsWithExternal with no third-party code in between, and nothing
in this repository mutates a shortcut element -- mountDeclarative takes a
value receiver and only replaces slice headers. Drop those three copies
and document the boundary on AllShortcuts so the next reader does not
reintroduce them.

Startup drops from four full clones to one. Benchmarks pin the remaining
cost so a regression points at a new clone rather than at growth in the
shortcut set.

* fix(command): align the commandtest page bound and escape wrapper paths

Two review findings, both of which let a business command pass its tests
and then misbehave in production.

The commandtest recorder walked 1000 pages for a complete-set collection
while the host adapter stops at 100, so a command tested against 300
pages of fixtures would fail its first real --page-all run with
PaginationLimitError. The bound now lives in internal/pagination, which
both sides already import, and the hard-limit test scripts itself from
that constant instead of restating 1000 -- the literal was what let the
two drift apart.

The testdata wrapper concatenated args.ID straight into the request path,
contradicting PathSegment's own documented rule and the chat-brief
example. ValidateRequestView does not cover this: "abc/other-users-file"
cleans to itself, so an unescaped separator silently retargets the
request. Since testdata is what an integrator copies first, route both
call sites through one readRequest helper, mirroring chat-brief.

The e2e assertion could not have caught it either -- PathSegment("chat_1")
is "chat_1", so the check passed with or without the call. It now sends
"chat/1" and asserts %2F reaches the wire; removing PathSegment fails it.

* fix(command): deny network to the pre-confirmation hooks and four review findings

Normalize and Validate run before the high-risk confirmation gate, and
both received the full CommandContext, so a high-risk business command
could POST or DELETE from Validate and leave remote side effects behind
before the user was ever asked to confirm. Moving the gate earlier would
contradict the documented hook order and would also make --dry-run
require --yes. The design already forbids this from the other side --
Validate is specified as parameter checking that issues no request -- so
enforce that instead: Normalize and Validate get a context whose CallJSON
and CollectPages refuse, while PreflightScopes stays available. The guard
sits in CommandContext rather than in the wiring, so a future adapter
that wires the callbacks anyway still cannot reach the API. commandtest
mirrors it, otherwise a command would pass its tests and fail only in
production.

Page.Items now starts non-nil. It is declared required;nonnullable, but a
zero-item collection encoded as {"items":null}, which a caller generating
types from the published schema would reject.

NewCmdAuthWithRecovery and NewCmdSchemaWithVisibility are restored as
wrappers. Both were dropped for shortcut-aware variants, and both are
reachable from outside this module: CommandVisibility is an ordinary
exported func type, and *recovery.Projector cannot be named by an outside
caller but can be passed as nil. A signature test now pins them.

The path-traversal fixture said "../../secret", which the deterministic
gate rejects as a generic credential assignment -- the reason CI is
currently red. The filename carries no meaning; it is now "../../outside".

* test(cmd): exercise the retained constructors instead of naming them

The compatibility wrappers restored for outside callers are unreachable
from inside this repository by construction, so the incremental dead-code
gate rejected them. A signature-only assertion did not help: taking a
function value and discarding it leaves the body unreachable, and it
proved nothing about whether the wrapper still builds a working command.

Call each one and assert the command it returns. NewCmdAuthWithRecovery
is called with a nil projector, which is the exact call an outside module
can make and the reason the wrapper has to keep compiling.

Verified with the same deadcode version CI runs: neither function is
reported, and no other function in this branch's files is either.

* docs(command): document the hook contract and pin dry-run note idempotency

Hooks is the first type a business author reads and carried no field
documentation, so the rules lived only in the design doc: which of DryRun
and DryRunE to set, that setting both fails to compile, that Execute owns
the API call and must not write stdout, and that Normalize and Validate
run before the confirmation gate and therefore get no network.

The choice between DryRun and DryRunE is not old-versus-new -- neither is
legacy. It follows from whether building the preview can fail, which is
now what the field docs say.

Also pin the dry-run note as idempotent. convertDryRun writes the
bounded-repeat note into the projection it builds, never back into the
hook's *DryRun, and DryRunAPI.Desc assigns rather than appends, so a hook
that caches and returns the same preview cannot accumulate the note. Both
properties were true and neither was tested.

* refactor(command): settle the dry-run constructor, tips, and domain enum

Three narrowings of the V1 business-command contract, none of which has a
published compatibility surface: extension/command does not exist on main.

Preview and NewDryRun were the same constructor twice -- one empty, one
seeded with requests. Fold them into a variadic NewDryRun. Every existing
NewDryRun() call keeps compiling, and the domain word in the contract is
now spelled one way. The type DryRun already owns that identifier in this
package, so naming the constructor DryRun outright cannot compile.

Drop Metadata.Tips. It was pure passthrough into common.Shortcut.Tips and
nothing in the execution path read it, so business commands lose only the
ability to declare help tips; the repository's own typed shortcuts keep
theirs. The mount test asserted a tip reached the rendered help as proof
that metadata survives the extension -> commandhost -> common.Shortcut ->
help conversion; it now asserts the risk line, which travels the same path.

Hand-write the domain enumeration and delete the generator. Generating
from shortcuts.AllShortcuts silently omitted approval, attendance and
mindnotes: all three are published under `lark-cli --help` and served by
typed and raw API commands, they just own no shortcut. The enum is now
the 23 domains the CLI actually exposes.

Those three would otherwise have been constants that compile and always
fail, because CompileSets derived its mountable domains from the same
shortcut list. It now reads the service registry, and shortcuts/register.go
already creates a domain command group on demand when no built-in occupies
it, so a business command can mount under a shortcut-less domain.

* ci: stop generating extension/command

The domain enumeration is hand-written now and extension/command holds no
go:generate directive, so the path was a no-op that still read as if the
package carried generated files.

* refactor(command): drop DryRunE and let the preview render like a built-in

DryRunE has no counterpart in the shipped CLI: `git show
main:shortcuts/common/types.go` has no such field, it arrived with the
typed-shortcut framework this branch imported, and no shortcut in the
repository sets one. Business commands get the single DryRun hook that
built-in shortcuts have. Validate already runs before it and owns the
error channel, so a preview that cannot be built still fails there with a
typed error -- which is what the repointed tests now assert, end to end
through --dry-run and through commandtest.Preview.

Also stop appending the bounded-repeat note. convertDryRun added "with
--page-all, repeats with the returned page_token until exhaustion or
--page-limit" to every Page[T] preview, so an external command's dry-run
carried a sentence its author never wrote. Built-in paginated shortcuts
say this themselves when they want it (im_chat_members_list.go calls
dry.Desc), and external commands now do the same: the framework renders
the description it was given and nothing else.

The dry-run context keeps refusing requests -- runner.go does the same for
built-ins, so removing that would be the divergence, not the alignment.
Only the word changes: "offline" was our own vocabulary for what the rest
of the CLI calls dry-run.

* refactor(command): drop the partial-failure outcome from the contract

Business commands now return Success only. Partial, OutcomeDefinition,
PartialFailureDefinition and FailedItemDefinition leave the public surface
along with Execution.Partial and the host adapter's receipt conversion.

Result keeps its outcome field. It is no longer a choice -- Success is the
only value -- but it is also how the host tells a returned Result apart
from the zero value that accompanies an error, which is the check
commandtest.Execute makes before reporting "returned both Result and
error". Collapsing it to nothing would delete that signal.

The exemplar commands that returned Partial keep their scenarios: a
best-effort scope failure still marks every item failed and appends the
snapshot, and the multi-call audit still records the owner it could not
resolve. That information lives in the command's own Data (Items[].State
plus Failures), not in the outcome, so the tests assert the same facts and
only the outcome assertion is gone. The deep-copy test moved its nested
JSON exemplar from FailedValues to InputDefault.Value, keeping
cloneJSONValue covered.

shortcuts/common still defines PartialFailure for built-in typed
shortcuts. That is the imported framework, untouched here.

* refactor(shortcuts): walk pages once for built-in and external commands

Commit 4d0c6ea61 added internal/pagination, moved PaginateInto onto it,
and then wrote a second caller for externally declared commands. Both
assembled the same Walk options, cloned the same params, read the same
cursor and mapped the same walk error; only the policy source, the call
path and the accumulator ever differed.

Those three now parameterize one pageWalk. PaginateInto keeps calling
through the RuntimeContext and keeps its per-page progress line; external
commands keep CallTypedAPI, the walker's context and their undecoded
pages, which the public contract needs because it decodes them into its
own Page[T]. Behavior is unchanged on both sides -- the external walk
still leaves Wait nil, which internal/pagination fills with WaitContext,
so --page-delay works exactly as before.

pageWalk is deliberately generic-free so one struct serves both callers;
the typed half of the built-in path moved to addDecodedPage.

* refactor(shortcuts): make the external page walk PaginateInto's twin

CollectCommandPages now differs from PaginateInto only where the context
type forces it. It takes the same PageAccumulator, decodes each page into
T through the same addDecodedPage, returns the same *output.PaginationMeta
and reads the same state out of the same walk, in the same order.

What is left is what the interface cannot supply. An externally declared
command compiles in the business module, so it holds a CommandContext
rather than a *RuntimeContext: the context arrives as a parameter because
the interface carries none, the call goes through CallTypedAPI, and there
is no progress line because deciding to print one needs StderrIsTerminal,
JqExpr and Format, none of which the interface exposes.

The all parameter stays. It is the complete-set policy CollectAllPages
depends on -- collect to exhaustion under the hard page bound instead of
obeying --page-all and --page-limit -- and PaginateInto has no way to
express it, since resolvePaginationPolicy only ever reads flags. Dropping
it would quietly turn a command that must see the whole set into one a
user can truncate with --page-limit 1.

CommandPageCollection is gone with it: pages accumulate in commandhost's
own accumulator, the way every built-in shortcut already accumulates its
own. One consequence of sharing the decode: a page whose response carries
no data object is now an error on this path too, as it always was for
built-ins.

* fix(shortcuts): reject recursive Data and Args types during compilation

shapeForType and compileStructShape called each other without recording the
Go types already being walked, so a self-referential type recursed forever.
The walk ran during command registration and ended in a stack overflow --
a fatal runtime error rather than a panic, so no recover boundary could
contain it and one extension command took the whole CLI down before --help,
schema, or any unrelated command could run. That also broke
CompileErasedDefinition's documented promise to compile without panic.

Thread the struct types open on the current recursion path through both
functions and return a compile error on a repeat visit, pointing at the
explicit Shape escape hatch. Membership is scoped to the path, not the whole
walk, so a type reused as a sibling or at another depth stays legal.

Covers self-reference through a slice and through a pointer, mutual
recursion through two types, the JSON-encoded Args path, and the public
CompileErasedDefinition contract.

* fix(command): share one result protocol between the host and commandtest

commandtest.Execute only checked that Data carried the expected type. Generic
erasure leaves a correctly typed zero Data behind, so a business command that
returned Result{} instead of Success(data) passed the type assertion and the
test reported success. Production rejects the same result, which left
extension authors with green tests and a command that failed on every real
invocation -- exactly the guarantee commandtest exists to provide.

Add ValidateHostResult in extension/command and call it from both
commandtest.Execute and the host adapter's execute hook, so the two surfaces
cannot drift. It rejects an empty or unsupported outcome and mirrors the
pagination receipt checks the host already applies: declared Page output,
pages of at least one, non-negative items, and next-token state consistent
with completeness.

RunWithFlags is covered because it delegates to Execute.

* feat(command): expose reusable download capabilities

* test(command): make the chat-brief example testable and cover its hooks

The example declared both commands as inline Definition literals handed
straight to Define. Define erases the type parameters and returns an opaque
Command that cannot hand its Definition back, while commandtest.Execute takes
the Definition -- so the shape the example demonstrated could not be unit
tested at all. Neither shipped example had a test file, so nobody had walked
the copy-the-example-then-add-a-test path.

Lift both declarations into Definition-returning functions, the shape the
repository's own commandtest suites already use, and keep the compiled
Commands as package vars so main is unchanged. The configuration bodies are
untouched.

Add the tests that shape exists for: the single-read projection and its
Validate rejection, plus the Page[T] contract for default single-page reads
and a --page-all walk through RunWithFlags. All four run offline through the
commandtest recorder.

* test(commandtest): cover the ordered URL download script

Recorder.ReplyURL shipped without a caller, so the incremental dead code
gate flagged it as new unreachable code and blocked the branch. The existing
URL download test uses the unordered RespondFile constructor, which never
exercises the URL assertion ReplyURL exists for.

Mirror the ReplyJSON pair: one test walks two scripted URLs in order and
checks the recorded source URLs, content types, and artifacts; the other
points DownloadURL at an unscripted URL and expects the mismatch error.

* feat(command): accept @file input for external commands

V1 rejected the file value source, leaving external commands with inline
flags and stdin only. A process has one stdin, so a command whose body is
too large or too quoted for the shell -- an XML document update is the case
that surfaced this -- had no second way to receive it, and the caller had to
fall back to shell escaping.

Nothing downstream was missing: resolveInputFlags already resolves @path and
the @@ escape through the invocation's FileIO, help renders the "@file"
affordance, and legacyInputSources maps the source onto the compiled flag.
The gap was the public constant and the host allow-list.

Export SourceFile and let it through compilation. Unknown sources still fail
the same way, which the rewritten host test now pins alongside the compiled
flag actually carrying both extra sources -- silently dropping a declared
source is the regression worth catching.

Give the wrapper fixture a command whose content flag declares all three
sources, and drive it end to end: @file and stdin both reach the request
body, and the help text advertises them.

* revert(content): drop the default content exports from the extension surface

Exporting the repository's embedded skills and affordance trees turned two
content directories into Go packages, because go:embed cannot reach up out of
a package directory and the repository root is package main. That cost two
things: a .go file living inside an authored-content tree, where a future
content type is silently omitted until someone edits the glob, and an embed
directive per tree where the root previously covered both in one.

The need it served does not exist. The distribution driving this work ships
its own skills and does not consume the official set, and a wrapper that does
want them can supply a tree through cmd.SetEmbeddedSkillContent, which is what
extension/platform already documents.

Restore content_embed.go, the SkillsOverlay comments, and the platform README
to their main state, and delete the two exporting packages. The example and
fixture wrappers now ship no embedded content, which is what a wrapper that
does not compile the repository root actually gets; the e2e assertion that
depended on inheriting lark-doc goes with it.

* fix(command): close the review findings on API surface and storage commit

Five findings from the extension-v1 review, each verified by a test that fails
against the previous implementation.

Source compatibility of auth.LoginOptions. The shortcut snapshot was an
unexported field on a struct that appears in the exported runF signature, which
ends positional literals for every caller outside this module. It becomes a
closure capture plus an explicit authLoginRun parameter, and the six domain
helpers collapse into domainResolver methods, removing five xxxWithShortcuts
twins that only tests reached. common.Shortcut.DryRunE goes too: it was a new
exported field with no production caller. The unexported typed field stays, so
Shortcut itself is still not positional-literal compatible -- that is a
deliberate remaining gap, since relocating it would need a global mutable map or
a wider signature change.

One canonical wire projection. queryValues stringified and dropped nils for the
live call while the dry-run preview and the pagination walk forwarded raw values,
so a preview could describe a request the runtime would never send. canonicalQuery
is now the only projection and all three consumers derive from it. Dry-run output
for numeric parameters therefore reads "20" instead of 20, matching what the
query string actually carries.

No-clobber as a storage guarantee. IfExistsFail checked existence, downloaded,
then committed with a rename that replaces unconditionally, so a target created
during the transfer was silently overwritten. The commit step is now an optional
ExclusiveFileIO capability: content lands in a temp file and is published with
Link, which refuses an existing target and never exposes a partial file. A
provider without the capability is refused rather than served a guarantee it
cannot keep.

V1 public surface. Removes NewDomain and its options (host compilation rejected
them), HostDomain.IsNew, reservedRootNames, and the unproducible
ResultMetaDefinition.Count; narrows Hooks.Renderers to a single PrettyRenderer,
since pretty was the only key the compiler accepted; and demotes the generic
authoring layer in shortcuts/common to unexported, as no production code outside
that package used it.

commandtest runs the production compiler. Execute, RunWithFlags and Preview now
share compileForTest, so a wrong tag, Shape or relation fails in the unit test
instead of at CLI startup. Applying this surfaced two long-standing contract
violations in the package's own fixture.

* refactor(command): keep a single authoring contract

* revert(schema): keep the schema command blind to shortcuts

The schema command serves the generated API catalog only, matching main.
Drop the shortcut contract lookup and completion from cmd/schema and
restore the constructor surface. ShortcutSchema stays on the sealed
commandbridge surface, where the host compiler tests assert the contract;
the CLI itself no longer consumes it. The surface scenario and wrapper
e2e pin the boundary from the other side: schema resolution and
completion must not see mounted shortcuts.

---------

Co-authored-by: sang-neo03 <266690410+sang-neo03@users.noreply.github.com>
Co-authored-by: liangshuo-1 <266696938+liangshuo-1@users.noreply.github.com>
2026-08-25 20:22:13 +08:00
zhangjun-bytedance aec659c461 feat: words replace and minutes fix (#2490) 2026-08-25 17:18:07 +08:00
zhicong666-bytedance 35bd5ecfcd feat(vc): add agent meeting control shortcuts (#2466)
* feat(vc): add agent calendar meeting actions

* feat(vc): add meeting screenshot shortcut for visual context

* feat(vc): add meeting countdown commands and events

* fix(skills): avoid screenshot recall from meeting description

* fix(vc): omit screenshot log ID on success

---------

Co-authored-by: renaocheng <renaocheng@bytedance.com>
Co-authored-by: shike.11 <shike.11@bytedance.com>
2026-08-24 21:26:44 +08:00
chenxingyang1019 e0e90a4e1b fix(apps): classify the online DDL ban and the file storage quota failure (#2460)
* fix(apps): classify the online DDL/DCL ban on +db-execute

Running DDL against the online branch of a multi-env app came back as
api/server_error with exit 1, hinted "fix the SQL and re-run", and carried a
statement position that did not exist. All three point the caller the wrong way:

  - server_error means "upstream 5xx, retryable"; this is a product rule and no
    number of retries changes it;
  - the SQL is fine — the target environment is what has to change;
  - "(at statement 1 of 1)" is fabricated. The server pre-validates the whole
    batch and returns a single ERROR sentinel, so a 5-statement request with the
    DDL in position 4 still rendered as "1 of 1". The CLI does not split the SQL,
    so it cannot know the real count and cannot detect the mismatch generally —
    only codes known to be batch-level rejections can drop the suffix.

Give code 4000001 its own arm: validation/failed_precondition (exit 2, "change
the environment, do not retry"), a hint pointing at dev plus +db-env-migrate, and
no statement position. Every other code keeps its current classification, wording
and position suffix; a test pins that.

4000001 is a dedicated server-side code (ErrOnlineEnvForbidDDLDCL, client-error
band), raised only by the pre-validation pass when env==online on a multi-env
workspace. Syntax errors and PG errors use different codes, so keying on it is
safe. Matching on the numeric value also covers the "k_dl_4000001" wire form,
since codeString already strips that prefix — both forms are tested.

"No statements were applied" is stated rather than inferred here: the validator
walks every statement and rejects the batch on the first DDL, so nothing lands.
The default arm would have inferred the opposite for a DDL in a later position
("Earlier statements were committed"), which is wrong for this code.

* fix(apps): classify tenant file storage quota exceeded

+file-upload against a tenant whose file storage is full returned api/unknown with
no hint at all, so a caller could not tell "the quota is full, stop" from "the
upstream had a bad minute, retry" — and had no next step either.

Register 400000055 as api/quota_exceeded. That is a dedicated server-side code
(ErrTenantStorageQuotaExceeded, client-error band) raised only on the upload path,
and the subtype already carries "retrying will not help", so the framework's
existing quota wording is enough and no domain-specific hint is added.

Left in CategoryAPI (exit 1) rather than Validation (exit 2): a full quota is not
something a different argument fixes, and exit 2 would imply it is.

The test asserts the hint is non-empty on purpose. The wording comes from the
shared APIHint table, so if quota_exceeded is ever dropped from there this fails
and says the code now needs its own wording, instead of silently shipping an
empty hint.
2026-08-24 16:49:33 +08:00
Vee fbd1aa49cd feat: add IM read status shortcuts (#2318) 2026-08-21 16:47:49 +08:00
zhaoleibd e525beb8d6 feat(skills): unify meeting related skills (#2387)
* feat(skills): unify meeting guidance

* fix(meeting): restore domain boundary guidance

* docs(meeting): remove agent rollout qualification guidance

* docs(meeting): front-load skill routing description

* docs(meeting): refine identity and command guidance

* docs(meeting): clarify identity and pagination guidance

* docs(meeting): fix minutes todo detail command

* fix(meeting): clarify artifact query routing

* fix(meeting): improve live meeting skill recall

* fix(qualitygate): generate valid minute token placeholders

* fix(meeting): address unified skill review findings

* docs(meeting): add minutes permission guidance

* docs(lark-meeting): 更新SKILL.md并新增会议问答引导脚本

1. 优化SKILL.md表格排版与快速行动章节内容,新增批量获取当日会议脚本的使用说明
2. 新增meeting_qa_bootstrap.py脚本,实现一站式采集当日进行中、已结束会议及未来日程,生成可直接执行的命令引导

* docs(calendar): clarify today's meeting lookup

* revert(meeting): remove meeting Q&A bootstrap guidance

* fix(skills): register lark-meeting suite keywords

---------

Co-authored-by: maozhixiang <maozhixiang@bytedance.com>
2026-08-21 10:49:46 +08:00
chenxingyang1019 ca35f60616 fix(apps): make cache-clear ask first, and make apps failures classifiable (#2415)
* docs(skills): require explicit confirmation before apps +cache-clear

Asked to clear an app's online cache, an agent read `Risk: high-risk-write` from
--help and then supplied `--yes` itself on the first call, wiping production
cache without ever hitting the confirmation gate.

The CLI gate is fine: no --yes -> exit 10 confirmation_required, and --dry-run ->
exit 0 without triggering it. The wording was not. It only forbade appending
`--yes` *after* an exit-10, and said "已明确授权可直接带 --yes" without defining
authorization — so "clear my cache" read as authorization.

- `+cache-clear` gets a CAUTION block: never self-supply `--yes` on the first
  call; without confirmation, either --dry-run or ask, then stop and wait. exit
  10 is not a signal to retry with --yes.
- Add a zero-ambiguity table separating a *request* to clear ("clear the online
  cache") from a *confirmation* ("我确认清 dev"), so blocking the accidental wipe
  does not also kill the cases that were already correct: an explicit
  confirmation still goes straight to `--yes`, and a request with no environment
  named still has to ask instead of picking one.
- Note that online needs a confirmation phrase even when named explicitly.

`+cache-delete` gains the response field an agent has to read
(`deleted_key_count`): 0 means the key never existed, not "deleted
successfully", plus the get -> delete -> get chain needed to prove a delete took
effect — a single miss afterwards cannot tell the two apart.

SKILL.md: add +cache-clear to 禁止预授权判定底线, the one list a pre-authorized
run cannot skip; a reference-level rule alone would be bypassed there. The
routing table is left alone — no other row annotates risk, including
+file-delete, +role-delete and +member-remove.

* fix(apps): stop attaching request-shaped hints to precondition failures

`+db-execute` against a tenant that never activated Miaoda returns code 221800
"miaoda UAT not activated" with the hint "verify table/column names with
`+db-table-get` ... target the dev database with --environment dev". Neither step
can help: the failure is tenant-level, so a caller following the hint loops over
table lookups and env retries that fail identically.

Two causes. 221800 was unregistered, so it degraded to api/unknown — nothing in
the envelope distinguished "your tenant is not activated, stop" from "your SQL
was wrong, fix it and retry". And withAppsHint filled the caller's hint whenever
the server sent none, without looking at what failed: the hints are
command-scoped ("verify --app-id", "verify table/column names", "list releases"),
so every one of them describes the request, and the request is exactly what
failed_precondition says was fine.

Register 221800 as validation/failed_precondition (same shape as 400002465 "app
has no database yet") and gate the hint fallback on the subtype.

Blast radius is two codes, since that is all the subtype covers here:
  - 221800 — now withheld; message and code still carry the meaning.
  - 400002655 "no running container" — only when it reaches a non-observability
    command; the observability pair rewrites it first, and "verify --app-id" was
    never the fix for an undeployed app.
400002465 / 500002759 are intercepted by the isAppNoDatabaseError branch above
the gate, and 400002479 is served by withDBSyncHint, which does not delegate
here. The other 78 call sites take the original path for every input.

Gate on the one subtype, not on Category: this package asserts on purpose that an
authentication failure on +role-list (99991663) keeps the app-access hint and a
503 on credential issuance keeps the developer-access hint. Those hints are broad
enough to survive a caller-standing failure; only the precondition class is
misdescribed by construction. A test pins that, so widening the gate to Category
fails loudly instead of silently dropping those hints.

The gate is asserted on the real classification path (BuildAPIError -> the code
table -> withAppsHint), not only on a hand-built Problem. Constructing
SubtypeFailedPrecondition directly feeds the gate the input it wants and passes
whether or not 221800 is registered, so the registration itself has to be part of
what the test covers.

No recovery hint for 221800 — the activation path is a product procedure, and
guessing one is what made this failure misleading in the first place.

* fix(apps): classify file-storage and app-level failures

Five Spark business codes reached the CLI unregistered, so every one of them came
out as api/unknown with exit 1: a caller could not tell "your app id is wrong"
from "you lack permission" from "the upstream is having a bad minute", and the
exit code offered no way to branch either.

  400002484  app not found            -> validation/invalid_argument     exit 2
  400002467  no admin/developer perm  -> authorization/permission_denied exit 3
  500002761  ditto, pre-4xx renumber  -> same
  400000034  file not found/no access -> api/not_found                   exit 1
  500000034  ditto, pre-4xx renumber  -> same

400002467 is not file-specific: db commands (+db-table-list, +db-table-get,
+db-quota-get, +db-changelog-list) return it for an app the caller cannot access,
so registering it fixes both domains at once.

400002484 covers a well-formed id that does not exist AND a malformed one
("notanappid", "app_1" return it too), so the argument itself is the failure ->
invalid_argument, whose exit 2 separates "you passed the wrong id" from an
upstream fault. Environments that have not picked it up answer with 400002465
instead, conflating it with "app has no database yet"; the CLI cannot tell those
apart on the old code, so nothing here keys on that.

Both the current and the pre-4xx number are registered for each file failure.
The domain is moving its client-class errors from the 5xxxxxxxx band into
4xxxxxxxx, rolled out per environment, so both are live at once and dropping the
old one would silently return the un-migrated half to api/unknown — the same trap
that made the no-database recovery flow disappear when the server renumbered it
(see appNoDatabaseCode). The new number is not derivable from the old either:
500002761 became 400002467, tail digits included.

No hints added: permission_denied already has framework recovery wording, and a
domain-specific one would have to invent a remedy.
2026-08-20 17:36:08 +08:00
guokexin.02 e0867e6ebb feat: support separate and suite skill layouts (#2211) 2026-08-18 20:59:31 +08:00
陈家名 679ebd5289 fix(api): reject query strings and fragments in paths (#2375) 2026-08-18 16:49:00 +08:00
kiraWangRuilong ac6d7ce4f8 feat: propagate credential metadata (#2305) 2026-08-13 14:42:44 +08:00
木杉 1e87f67244 fix(apps): friendly-ize "Container not exists" for observability commands (#2302)
* fix(apps): friendly-ize "Container not exists" for observability commands

+metric-list and +analytics-list passed the upstream business code 400002655
("Container not exists") through verbatim. The message reads like an
infrastructure fault and misleads callers (including AI agents) into retrying a
non-retryable, expected business state: an app with no running container simply
has no metrics/analytics to query yet.

Rewrite it at a scoped observability helper (withObservabilityHint) into a
user-facing explanation plus a deploy-then-retry next step, mirroring the
existing isAppNoDatabaseError override. Detection is code-OR-message so a server
renumber alone does not silently drop the rewrite. Classification, code, and the
wrapped cause are preserved; unrelated failures still fall through to the shared
app-id recovery hint (and its own no-database override).

* test(apps): add execution-path regression tests for observability container hint

common_test.go proves withObservabilityHint in isolation but stays green if a
call site reverts to withAppsHint. Drive +metric-list and +analytics-list
Execute with a mocked 400002655 "Container not exists" envelope and assert the
container-specific message/hint/code, so a revert fails the build. Also closes
the two uncovered call-site lines flagged by coverage.

* fix(errclass): classify no-container code as validation/failed_precondition

Register 400002655 in sparkCodeMeta mirroring its no-database twin
(400002465) so both "expected precondition not met" business states expose
the same validation/failed_precondition classification to machine consumers,
instead of falling back to api/unknown. The shortcut-layer message rewrite
already keyed off the raw code, so this only aligns the typed envelope's
category/subtype; update the execution-path tests to pin the new
classification.

* fix(apps): gate the no-container hint's release behind user authorization

The no-container hint told a harness to deploy via +release-create, a "write"
that takes the whole app live and can affect existing production traffic —
without the user-confirmation gate its no-database twin deliberately carries.
Since the hint's audience is an AI agent that acts on it, a failed metrics read
could trigger an unconfirmed go-live. Lead with a read-only +release-list
status check and gate +release-create behind an explicit user confirmation,
mirroring appNoDatabaseHint.
2026-08-12 15:01:10 +08:00
ethan-zhx 7436a53afc feat(slides): inline slides reference docs into help output (#2181) 2026-08-11 19:29:57 +08:00
木杉 4488da0b14 feat(apps): add +user-id-convert shortcut for Miaoda↔Feishu ID conversion (#2270)
* feat(apps): add +user-id-convert shortcut for Miaoda↔Feishu ID conversion

Wrap the platform id_convert OpenAPI as a read-only shortcut that maps
Miaoda user_id ↔ Feishu open platform IDs (open_id / union_id / Feishu
user_id). It does one thing — conversion — with no local mapping table,
caching, permission pre-check, or direction guessing.

- --convert-type enum → server id_convert_type (10/11/20/21/40)
- --ids: csv / @file / stdin, 1-100 per call, not de-duped, input order
- reconstructs data.missed by diffing input positions against returned
  source_ids (server silently drops unresolved IDs), keyed by 0-based index
- meta counters (total/hit_count/missed_count) via pointer fields on
  output.Meta so an explicit missed_count: 0 survives omitempty

* test(apps): address review feedback on +user-id-convert

- reject empty --ids CSV entries (e.g. "a,,b") with a typed validation
  error instead of silently dropping them, since a dropped entry shifts
  every later result's 0-based index and breaks the position-keyed
  items/missed contract; add an interior-empty-element test
- reuse common.GetSlice / common.GetString for response projection
  (house convention) instead of local asSlice/asString helpers
- requireConvertValidation now asserts CategoryValidation +
  SubtypeInvalidArgument via errs.ProblemOf, keeping ValidationError.Param
- table-drive TestResolveConvertType over all five directions so every
  --convert-type → id_convert_type mapping (10/11/20/21/40) is protected

* fix(apps): split newline-delimited --ids for +user-id-convert @file/stdin

@file and - (stdin) input arrives verbatim from the framework as
one-ID-per-line text, but parseConvertIDs only split on commas, so such a
block was sent as a single malformed request ID. Treat a newline as
equivalent to a comma, tolerating a file's trailing newline while still
rejecting interior empty entries so position-keyed result indices stay
aligned. Add @file and stdin tests asserting the request body's ids are
split into discrete IDs.

* fix(apps): stringify numeric JSON IDs in +user-id-convert results

Responses decode with json.Number (client.ParseJSONResponse uses
dec.UseNumber()), so a server that emits source_id/target_id as bare
numbers — plausible for the numeric Miaoda user_id form — was silently
coerced to "" by buildConvertResult's strict string assertion: the
source_id got dropped (false not_found) and the target_id blanked
(false success).

Add common.GetStringLoose, which stringifies string/json.Number/int64/
float64 via literal text (large integer IDs keep full precision, never
routed through a lossy float64), and use it for both id reads. Cover it
with a package-level table test plus an end-to-end regression asserting a
numeric-JSON response yields intact, non-blank ids and no false miss.

Also exercise resolveConvertType's non-empty "not a valid direction"
branch directly, since the runner's enum gate preempts it in normal flow.

* test(common): tighten GetStringLoose numeric coverage

Add an int-branch case (was only covering int64) and swap the float64
fixture from 42 — which no formatter would render in exponent form — to
1e-7, whose fixed-point rendering "0.0000001" fails under the 'g' verb.
This turns the "no scientific notation" case into a real guard for the
'f' verb choice, per CodeRabbit review on c64cca39.
2026-08-11 19:14:52 +08:00
jinjiuzhe 158d15b3fd feat: add apps database sync shortcuts for Base-to-database import (#2251)
* feat: add apps database sync shortcuts

Add Base-to-database sync shortcuts for preview, create, list, get, enable, disable, update, and delete flows.

Cover OpenAPI request contracts, typed sync error classification, dry-run E2E coverage, and lark-apps skill guidance.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(apps): send db-sync task_id and config in request body

The enable/disable/delete/update sync commands placed task_id (and
update's config) in query params, but the OpenAPI contract binds these
fields via api.json (request body). BOE testing returned
"field validation failed" (99992402) because the body was empty.

Move task_id to the request body for enable/disable/delete, and move
both task_id and config to the body for update. Dry-run previews now
render these under body, and unit tests pin the body binding so a
regression to query params fails.

* fix(apps): use POST for db-sync-delete action endpoint

The delete command issued an HTTP DELETE to db/sync_del, but the
action-style endpoint is registered as POST (like sync_create and
sync_disable). The method mismatch made the gateway return a plaintext
404, surfacing as "API returned a non-object JSON response".

Switch the request and dry-run preview to POST, and pin the method in
the delete unit tests so a regression to DELETE fails.

* test: pin db-sync update base_url as optional contract

* test: pin db-sync update omits base_url without silent default

* docs(skills): clarify db-sync source.base_url create-required update-optional contract

* fix(apps): send db-sync env in request body not query params

The +db-sync-create and +db-sync-update endpoints read env from the
request body (peer of config/preview/task_id), not the query string.
Placing env in query params left the body env empty, so the server
treated every request as online and rejected DDL operations
(code 500002776: forbid ddl/dcl operation in online env), making it
impossible to create/update sync tasks against a dev environment.

Move env into the request body via a new dbEnvBody helper that mirrors
dbEnvParams' omit-empty contract, so unset env still lets the server
auto-select the branch. Pin the contract in unit and e2e dry-run tests
by asserting body.env and that env is absent from query params.

* test: align db-sync operate/delete e2e with request-body contract

The enable/disable/delete dry-run e2e still asserted the pre-migration
wire shape: delete on DELETE and task_id in query params. The shortcuts
now POST these actions with task_id in the request body (commits moving
task_id and the delete verb), so the stale assertions failed against a
current binary.

Assert POST + body.task_id and that task_id is absent from query params,
pinning the same body-over-query contract the env fix established.

* fix(apps): improve db-sync create ergonomics and error guidance

Refine +db-sync-create/update validation, error hints, and docs so AI
agents recover from common Base-to-database sync failures without guessing:

- source.table.name: document that a user-named table must be set, name
  takes precedence over the base_url ?table= token; fix test fixtures that
  used a fictional source.table.url instead of source.base_url.
- Preflight source table locate: reject create locally when base_url has no
  ?table= and source.table.name is empty, pointing at base +table-list.
- Online DDL ban: attach a precise hint for code 500002776 + subcode
  k_dl_4000001 telling multi-env apps to create tables on --environment dev.
- Missing record-id column: extend the 500002783 hint to add a unique text
  column via +db-execute before retrying.
- Optional field_maps on create: allow omitted or empty field_maps so the
  server auto-matches and creates the task; keep update requiring an enabled
  mapping and still reject an all-disabled array.
- Environment default: db-sync commands use online when --environment is
  omitted; align help text, comments, and skill docs.

* fix(apps): migrate db-sync error codes to the 4xx client-error range

The backend moved the seven db-sync error codes from the 5000027xx
server-error range to the 4000024xx client-input range to reflect that
they are client-input errors. Mirror the new codes in the CLI so error
classification and recovery hints keep matching:

- 500002783 -> 400002477 (mapping invalid)
- 500002784 -> 400002478 (target schema mismatch)
- 500002785 -> 400002479 (operation not allowed)
- 500002786 -> 400002480 (task not found)
- 500002787 -> 400002481 (invalid task id)
- 500002788 -> 400002482 (source table not found)
- 500002789 -> 400002483 (target table not found)

Category, subtype, hint text, and behavior are unchanged; 500002776
(online DDL ban) is untouched.

* fix(apps): tighten db-sync preview validation and pretty output

Address review follow-ups on the db-sync shortcuts:

- +db-sync-get pretty output no longer prints <nil> for a missing
  schema_only nor Go map syntax for statistics; render a bare bool and
  deterministic key=value pairs instead.
- Reject a non-array field_maps in +db-sync-create --preview as well as
  commit, so the malformed shape is caught locally rather than forwarded
  to the backend.
- Clarify in lark-apps-db.md that +db-sync-create --preview needs no
  confirmation and only a real create requires --yes.
- Harden the db-sync dry-run validation tests to assert exit code 2 and
  the structured stderr envelope (type/subtype/param), and add coverage
  for the preview non-array field_maps rejection and batch pretty output.

* fix(apps): guard db-sync preview output and neutralize update hint

Address the next db-sync review round:

- +db-sync-create --preview --output no longer writes a "null" file and
  exits success when the response omits data.config; project config into
  a typed object and return internal/invalid_response without writing.
- Make the 400002482 code hint command-neutral so +db-sync-update is not
  steered into a create-only recovery path that risks duplicate tasks.
- lark-apps-db.md: carry --environment on the update lifecycle examples
  and split failure recovery by streaming (can update) vs batch (cannot
  update; recreate instead), removing the batch/update contradiction.

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-11 18:21:32 +08:00
liangshuo-1 79eb16c50b feat(affordance): support domain skill lists (#2291)
Co-authored-by: liangshuo-1 <266696938+liangshuo-1@users.noreply.github.com>
2026-08-11 16:27:37 +08:00
fangshuyu-768 82e628bf79 fix(drive): harden export and push failure recovery (#2279) 2026-08-11 14:20:42 +08:00
zhoujunteng-max 5b734238d7 feat: report upload file events (#2093)
* feat: report upload file events

* test(drive): skip import workflow without tenant token

* docs: document upload report helpers

* docs: improve function documentation coverage

* docs: complete incremental function documentation

* docs: complete function documentation coverage

* fix: report every upload file event

* refactor: move file event reporting to internal package

* fix: harden upload file event reporting

* fix: omit upload mode from file event reports

* test: fix workbook import dry-run token assertion
2026-08-10 19:12:38 +08:00
SunPeiYang996 a6f3e635d4 feat(docs): add local authoring and resource workflows (#1921)
* feat(docs): add local authoring and resource workflows

Add docs +script workflows for isolated draft initialization and tolerant XML/Markdown profiling.

Support local and remote document resources across create and update flows with safe, bounded-concurrency uploads, binding verification, and cleanup.

Synchronize shared credential-source selection during concurrent uploads, refresh lark-doc guidance, and expand unit, dry-run, and live E2E coverage.

* docs(lark-doc): clarify genre reference paths

* fix(docs): address PR validation feedback

* docs(lark-doc): clarify remote image handling

* feat: streamline docs draft workflow

* fix(docs): clarify script input and resource cleanup

* fix(docs): align script dry-run test with auth flow

* fix(docs): authenticate local script e2e test

* fix(docs): refine script diagnostics and image preflight

* fix(docs): align draft workspace cleanup with VFS

* fix(docs): route workspace cleanup through FileIO

* docs(lark-doc): simplify profile check guidance
2026-08-07 18:15:28 +08:00
liangshuo-1 0d6f2c65d7 fix(im): harden resource downloads with validated ranged streams (#2223) 2026-08-07 17:45:40 +08:00
huarenmin13 7c4f6c023f fix(base): improve field creation and query guidance (#2114)
* fix(base): improve field creation and query guidance

1. Document batch field creation in shortcut help and the delivered Base skill.
2. Choose field types from stored values instead of business-purpose names.
3. Add generic common filter values to the data-query quick guide with contract tests.

说明:
- Combines the accepted fixes for base_table_096, base_table_028, and base_table_087;
    MR 1275 contributes round 4 only.

```ai-signature
改动范围: Base field-create 帮助与 Skill 指南、data-query 快速指南,以及对应的 shortcuts/base 契约测试
思考过程: 保留三个实验的最终通用规则,合并 public main 上新增的字段读回提示,并排除没有 benchmark 支撑的 MR 1275 round6
改动原因: 让代理发现批量字段接口、按存储值选择字段类型,并用常见通用形状构造 data-query 过滤条件
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: a53fdb6b6b82d74e4deff5e6d32591ec897cd2dc6ad662885c961a27e67d4666

* test(base): strengthen guidance contracts

1. Assert every common filter fragment introduced by the data-query quick guide.
2. Keep the field-create argument table compliant with markdown table spacing.

```ai-signature
改动范围: data-query 指南契约测试与 field-create Markdown 表格后的空行
思考过程: 逐条核对 CodeRabbit 建议,只补会防止新增指南片段回退的断言和确定性的 MD058 格式问题,不改生产提示语义
改动原因: 关闭 PR 2114 的两条有效自动审查意见并保持变更可回归
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 81096f474001f2c9c59af48ee64150f205101710ebcf3b909ab9d300a019f46f

* fix(base): generalize data-query filter guidance

1. Replace the date-and-status scenario template with reusable Condition.value shape rules.
2. Update the contract test to require relative-date guidance and reject evaluation-shaped placehold
    ers.

```ai-signature
改动范围: lark-base data-query quick guide 与对应 shortcuts/base 契约测试
思考过程: 保留 select、datetime、empty 的通用 value shape,删除日期字段和状态字段组合模板,避免将 base_table_087 的解题路径固化到公共指南
改动原因: benchmark 显示当前文案能引导目标题,但组合示例与测试过度贴合单题,需要收敛为跨场景可复用的不变量
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 7b6916f05dccf6489dce65610fe757f7b346bcd0f8f5500cfbfb1f30c60471f3

* fix(base): generalize creation guidance

* fix(base): report partial field-create results

* fix(base): preserve partial field-create recovery metadata

* fix(base): separate partial field recovery

* perf(base): 缩短字段批量创建的空等

1. 将固定 1 秒批次等待改为 500ms 最小请求起点间隔,并让请求耗时抵扣等待
2. 新增节流计算契约测试,覆盖首次请求、快速响应和慢响应
3. 同输入 150 字段 A/B 从 269.64s 降至 118.80s,且两侧均创建 150/150

说明:
- 保持同表写入串行和 partial failure 输出不变

```ai-signature
改动范围: shortcuts/base/field_ops.go 与 shortcuts/base/base_execute_test.go,仅调整 field-create 数组批次的串行节流计算和回归测试
思考过程: 保留同表串行写入,以 500ms 作为请求起点最小间隔,并把请求耗时计入间隔,避免固定空等同时降低写冲突风险
改动原因: PR 引导 Agent 使用数组批量创建后触发既有每项固定 1 秒等待,导致 150 字段用例产生约 149 秒可归因耗时回退
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 69b888d362d2f484c4ce6ad050bdbfe2de7368948eb79ba516bcaa6806ec9107

* fix(base): 收紧不支持字段行为的终止边界

1. 派生、自动、同步或回填行为只使用已记录能力,无法实现时禁止探测、占位或虚假完成
2. 删除 data-query 契约测试对旧题模板占位符的反向黑名单,只保留原子规则和真实 case 污染检查
3. field-create 指引替换前后均为 36 个英文词,不扩大该帮助项的词数

```ai-signature
改动范围: shortcuts/base/field_create.go、base_shortcuts_test.go 与 data_query_guide_contract_test.go,仅收口通用终止规则和测试泛化
思考过程: 采纳 review 中可独立闭环的两点,不增加翻译专用规则,不修改运行时能力;用等词数替换避免帮助上下文继续增长
改动原因: 当前规则能阻止按业务名猜字段类型,却仍允许退化成普通文本占位;同时测试记住旧题模板会阻碍未来合理示例
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 759b7f4044c6381b963c8ea299b70967e634ab9d43b6ebfbe94ef38bb69ba46a

* perf(base): 压缩字段批量创建的评测开销

1. 批量成功与部分失败仅返回字段 id/name/type,保留恢复所需身份并减少大响应上下文
2. 引导数组在调用方超时范围内一次提交,并为生成的大数组推荐 @file 或 argv-safe 调用
3. 字段列表默认页大小提升到 API 上限 200,避免百字段以上场景的帮助查询与重试

说明:
- 同口径 case032:raw token 432292→428304,weighted token 126408→117761,耗时 272675ms→235022ms,两侧均读回 154 字段

```ai-signature
改动范围: Base field-create 批量输出、帮助提示、field-list 默认分页及对应契约测试与 Skill 返回说明
思考过程: 从同口径 trace 定位固定分块、大字段对象回传、100 条分页和 shell 双重转义四个确定性开销,保持单字段与部分失败恢复语义不变并逐项用测试锁定
改动原因: PR 引导数组批量创建后虽降低耗时,但多回合大输出会推高 raw token;需要在不牺牲正确率和恢复信息的前提下同时压缩 token 与耗时
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 8a2bd2d79c9abdb1e36ae45822dd04b2a3384498a1dc268cffcd47cff00d94ed

* perf(base): 收敛字段批量创建的上下文

1. 大数组成功路径推荐保留摘要的 --jq 投影,失败路径仍原样保留部分失败明细
2. 为一个或多个简单 text 字段提供 help fast path,并让 next_step:done 终止默认回读
3. 补充 Skill、帮助与执行结果契约测试,锁定有界输出和可恢复失败语义

说明:
- 最新 main 同题 A/B:两侧均回读 154 字段,raw token 下降 5.9%,耗时下降 19.3%,峰值上下文下降 13.3%,工具调用下降 20%

```ai-signature
改动范围: Base field-create 帮助、简单字段成功提示、lark-base Skill 路由与对应契约测试
思考过程: 从最终 A/B trace 分别定位批量成功展开 150 项、简单 text 读取冗余指南和成功后整表回读三类可控上下文开销,用成功摘要与失败全量明细分流来保留恢复能力
改动原因: 继续优化 PR 2114 的 token 和耗时,同时要求任何回退不能归因到 PR;需要让大批量成功路径有界且不削弱正确性或部分失败恢复
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 626b288a7f5444dfecb872e6cd29d4fe0788404f5001dcb47aac64e2d55d9b5a

* fix(base): 恢复批量字段输出与分页默认契约

1. 批量创建完整成功时保留服务端字段元数据,部分失败仍返回精简 identity
2. 将 +field-list 默认页大小恢复为 100,继续支持显式 --limit 200
3. 补充回归测试与字段创建文档,保留 --jq 有界输出指导

说明:
- 定向、Base 全量、race、仓库单测、构建、vet 与 lint 均通过

```ai-signature
改动范围: Base 批量字段创建成功输出、字段列表默认分页、对应测试与文档
思考过程: 先用契约测试复现完整字段元数据丢失和默认分页翻倍,再只恢复主干既有成功输出与默认值,同时锁定部分失败精简输出不变
改动原因: 移除可归因到 PR 的兼容性和 token 回退,并保留显式 jq 投影、节流、快速路径及部分失败恢复带来的通用收益
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 6351ed9d228611e3b6f5bc30faf278835b463e07f4eb020f225abd41e264399c

* fix(base): present batch field errors before partial output

1. Project the typed field-create error through Runtime.PresentError before copying result fields
2. Read Error, ProblemOf, and permission extensions from the presented clone
3. Cover visible scoped authorization and concealed recovery without fabricating missing_scopes

说明:
- Targeted, Base, full race, build, vet, format, lint, and module checks pass

```ai-signature
改动范围: Base 批量字段创建部分失败的错误呈现,以及 visible 和 concealed 恢复契约测试
思考过程: 先在最新 main 合并树上复现无 scope 授权提示和隐藏命令泄露,再复用兄弟批量命令的 PresentError 边界,仅替换 payload 复制时的错误来源
改动原因: OutPartialFailure 不会再次呈现根错误,必须在复制 typed error 字段前应用命令 scope 与发行隐藏策略
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: deee600dc8d4bb889629895178c1f217ba7762fbd5b3163da3b9d621c8585ac0

* fix(base): allow recovery before retrying failed field writes

1. Clarify that retryable gates unchanged automatic retries, not corrected resubmissions
2. Classify Base error 1254291 as a retryable conflict with canonical wait guidance
3. Cover authorization recovery, write conflicts, and reference contract consistency

```ai-signature
改动范围: internal/errclass/codemeta_base.go、shortcuts/base/field_ops.go、对应 Base 回归测试与 field-create 参考文档
思考过程: 将 retryable 限定为同一请求原样自动重试资格,保留授权或输入修正后重新提交,并复用现有 conflict 恢复提示
改动原因: 部分失败顶层提示会与权限恢复 hint 冲突,且 1254291 未分类导致等待重试规则无法由结构化错误驱动
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 65cd95f24968c5e0b0c43cf858ed787eabb6d9d63b3d5199a294bd284816c35e

* fix(base): preserve partial field recovery contracts

1. Preserve presented typed-error extensions without allowing them to overwrite batch ledger fields
2. Align field creation guidance with command-specific name semantics and caller-timeout recovery
3. Cover security challenges, extension collisions, and storage-type selection with regression tests

```ai-signature
改动范围: shortcuts/base 的 field-create 部分失败输出、命令提示、Base Skill 写入规则、field-create reference 与对应回归测试
思考过程: 复用 Runtime.PresentError 后 concrete typed error 的 JSON wire shape 作为扩展字段单源;对批次账本自有键统一生成无冲突 error_ 别名,并只收敛已证实的同名、fast path 与 timeout 契约矛盾
改动原因: 部分失败会丢失 challenge_url 等恢复字段,自定义 typed error 还可覆盖 status/index/error 导致错误账本;过度绝对的字段类型、同名和超时文案也会形成可归因正确率回退
Break Change: 否
```

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 02377560f14202ea652f530389ac102647189923e45548337814ec3871f5a4de

---------

Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
2026-08-07 16:19:19 +08:00
chenxingyang1019 6402080328 fix(apps): detect the no-database failure by code or message (#2217)
* fix(apps): detect the no-database failure by code or message

The recovery flow for "db command against an app that has no database"
keyed on one business code (500002759). The server has since renumbered
that case to 400002465, which silently disabled the flow: users now see
the raw internal message about workspace / app-id mapping and lose the
cloud-development recovery steps entirely.

Nothing catches the regression. There is no compile error, the unit tests
compare against the same constant they set, and the dry-run E2E does not
exercise a real response — the failure only shows up against a server that
has already renumbered.

Detect on code OR message instead. Both known codes are kept, plus narrow
lowercase markers of the server's internal wording. The two channels have
opposite failure modes: a code is precise but gets renumbered, a message
survives renumbering but breaks on rewording or localization. Requiring
either to match means one channel changing degrades nothing, and only a
simultaneous change of both regresses.

Markers stay deliberately narrow. "no db branch" in particular must not
also swallow env-pull's "invalid db branch" case, which needs its own
hint; a comment records that widening them requires a test proving the
neighbours still pass through.

Classification and the cause chain are untouched: the helper still mutates
the problem in place and returns the same error value.

* test(apps): assert the full typed-error contract in no-database cases

Review feedback: the new subtests checked only Message and Hint, so a
change that reclassified the failure — or replaced the error value and
dropped the cause chain — would still have passed.

Each case now asserts Category, Subtype and Code are untouched by the
rewrite, and that the helper returns the same error value. Inputs use a
concrete subtype rather than Unknown, so a clobbered classification is
actually observable. One new case wraps a cause and asserts errors.Is
still finds it through the rewrite.

Also covers the predicate's defensive nil guard, which withAppsHint cannot
reach on its own (ProblemOf returns ok=false for untyped errors), closing
the two uncovered lines the coverage report flagged. Both withAppsHint and
isAppNoDatabaseError are now at 100%.

* fix(errclass): classify the db-domain business codes

Three codes reaching the Apps db commands were absent from the Spark
table, so BuildAPIError fell through to the CategoryAPI + SubtypeUnknown
catch-all and the envelope carried no usable classification.

"App has no database yet" registers as Validation / FailedPrecondition:
the app resolves fine and the request is well-formed, but a prerequisite
the caller must create first is missing, so retrying unchanged can never
succeed. This moves its exit code from 1 to 2 — "fix the state" rather
than "the call failed" — and a test pins that so a future
reclassification has to be deliberate. Two codes cover it because the
server renumbered the case into the 4xx band; the legacy one stays for
older servers.

"Table does not exist" registers as API / NotFound, an ordinary
missing-resource lookup with no exit-code change. SubtypeNotFound has no
APIHint default, which matters here: the Apps layer fills its
command-scoped hint only when the classifier left Hint empty, so a
context-free default would displace the more actionable one. A test
guards that too.
2026-08-07 15:13:31 +08:00
xiongyuanwen-byted be2a96f490 feat(sheets): harden error prescriptions, batch updates, and read workflows
Aggregate the sheets work from feat/lark-sheets-develop:

- Improve validation errors with schema hints, aggregated issues, enum guidance, and prescriptive flag/style-field messages.
- Harden +batch-update input contracts, key normalization, style vocabulary handling, and resource-budget checks.
- Add read offload and truncation handling for cells, csv, and table-get, with typed output-path errors and safer jq/output-path semantics.
- Correct freeze semantics by emitting full-state freeze/unfreeze operations and adding --rows/--cols for +dim-freeze.
- Improve +styles-put and shared --styles parsing for styles, merges, row/column sizing, freeze, and sheet-prefixed range validation.
- Fix dim-insert inherit-style mapping, table-get date/time handling, table-put style anchors, and CSV path-shaped input guards.
- Update lark-sheets skill docs, scripts, tests, and generated flag data.

Tested with:
- go test ./shortcuts/common ./shortcuts/sheets/...
- go test ./shortcuts/... ./internal/...
- python3 -m py_compile skills/lark-sheets/scripts/*.py
2026-08-07 11:21:12 +08:00
sang-neo03 164d3ccd27 fix(registry): describe the attendance and mindnotes domains (#2210)
Both are registered, reachable business domains — `lark-cli attendance
user_tasks query` and `lark-cli mindnotes nodes create` run, and both appear in
`auth login --domain` — but neither had an entry in
service_descriptions.json. GetServiceTitle/GetServiceDescription returned "",
so buildDomainMeta fell back to the typed service spec, which carries one
string for both languages. The result was visible in `--help`: mindnotes
rendered its Chinese description in the English domain list, and attendance
rendered "attendance record query", a lowercase fragment shown in both locales.

Both keep their own scope namespaces (attendance:task:readonly,
mindnote:node:create/read) and stay independent auth domains, so no
auth_domain is set. whiteboard remains the only domain folded into docs.

TestGetDomainMetadata_HasTitleAndDescription could not catch this: it asserts
on buildDomainMeta's output, which the fallback had already made non-empty.
The new reconciliation test walks EmbeddedServicesTyped plus AllShortcuts and
asserts both languages resolve from the config itself, before any fallback.
EmbeddedServicesTyped is the overlay-free parse, so the test does not depend on
what remote_meta.json happens to hold on the machine — a domain that only ever
arrives via remote overlay stays out of its reach.
2026-08-06 19:02:21 +08:00
kiraWangRuilong 9f713e60f7 fix(auth): harden token refresh and concurrency handling (#2135)
1. Improve retry behavior for refresh failures.
2. Improve authentication token refresh reliability during concurrent activity.
 - Add a token-storage writability probe before refresh.
 - Add generation-safe token updates.
 - Add lock for all set/update/delete token operation.
2026-08-06 17:50:45 +08:00
liangshuo-1 9759167542 fix(im): unpin search identity after bot support landed (#2208)
#2194 extended `im +messages-search` to `AuthTypes: {user, bot}` but left the
affordance example and the skill reference asserting user-only, so the
dual-identity guard added by #2199 fails on main.
2026-08-06 15:36:59 +08:00
leave330 fcdef499bb feat(event): compile the catalog and harden the consume pipeline (#2142) 2026-08-06 02:44:23 +08:00
liangshuo-1 43fe6c8787 fix: surface retry metadata for TAT rate limits (#2200)
Co-authored-by: liangshuo-1 <266696938+liangshuo-1@users.noreply.github.com>
2026-08-06 00:31:57 +08:00
liangshuo-1 960bdf6d71 feat: migrate application and IM guidance to affordance (#2199) 2026-08-05 23:20:39 +08:00
liangshuo-1 47c9741ce0 feat: profile selection from environment with source-aware errors (#2198) 2026-08-05 23:09:22 +08:00
evandance 09feefe96b fix: make agent recovery and concealment reliable (#2189) 2026-08-05 19:44:11 +08:00
fongwave 8e884928f6 feat: add Base table copy shortcuts (#2019)
* feat: add Base table copy shortcuts

* test: cover Base table copy edge cases

* fix: preserve table copy task state

* fix: align table copy recovery with API errors

* fix: preserve table copy auth recovery

* test: verify copied table schema

---------

Co-authored-by: fongwave <272393974+fongwave@users.noreply.github.com>
2026-08-05 15:43:10 +08:00
evandance ebdeda854d feat(extension): present restricted commands as absent and trim skills (#1837) 2026-08-05 01:44:50 +08:00
liangshuo-1 2a1613484a feat: add framework flag aliases and unified IM pagination (#2146)
* feat: add framework flag aliases and unified IM pagination

Introduce declarative exact-name flag aliases at the shortcut framework boundary while keeping semantic compatibility domain-owned. Add a shared, format-aware IM pagination pipeline with consistent flags, metadata, safety bounds, resumable cursors, and request throttling.

* fix: align alias attribution and pagination contracts

* fix: align alias contracts and documentation

* test: remove environment-dependent contact bot e2e

* test: restore contact bot e2e

* docs: reduce IM pagination guidance noise

---------

Co-authored-by: liangshuo-1 <266696938+liangshuo-1@users.noreply.github.com>
2026-08-03 19:20:40 +08:00
evandance 40a0a9de66 feat(enhancement): centralize HTTP transport policies (#2021) 2026-08-02 14:55:05 +08:00
HanShaoshuai-k cfe76ad56a ci: add protected public domain allowlists (#2111)
Co-authored-by: HanShaoshuai-k <268785735+HanShaoshuai-k@users.noreply.github.com>
2026-07-31 11:02:04 +08:00
zhouyue-bytedance 4a16139348 fix(base): resolve Base URL block types accurately (#2099)
* fix: resolve Base URL block types accurately

* fix: resolve Base block selection from Wiki URLs

* fix(base): guide resolved folder and docx blocks

* fix(base): avoid field fallback for untyped URL blocks

* docs(base): specify URL example fence language

* test(base): cover unmatched URL block resolution
2026-07-30 20:29:47 +08:00
liangshuo-1 c167163d70 feat: propagate invocation metadata (#2097) 2026-07-29 19:39:53 +08:00
kiraWangRuilong e7d5ecdd01 feat: add risk-control protection (#1910)
1. Add baseline safe protection for Feishu/Lark API endpoints.
2. Add lark-cli config risk-control on|off|default command for workspace-level safety protection control.
2026-07-24 17:12:10 +08:00
liangshuo-1 b8f56dbc0b feat(apps): support absolute and relative upload paths (#2005) 2026-07-23 17:52:49 +08:00
sang-neo03 4c1a92caa6 refactor: converge success output through a single Emitter that owns the write (#1899)
* refactor: add output emitter contract and differential harness

Introduce a leaf Emitter in internal/output that composes the existing
output primitives (content-safety scan, envelope, jq, format rendering,
notice) behind a single command-scoped port. The emitter is unwired: no
production caller is migrated, so CLI output stays byte-for-byte unchanged.

A differential test harness drives the real legacy entry points
(RuntimeContext.Out/OutRaw/OutFormat/..., WriteSuccessEnvelope and the
pagination formatter) and asserts byte-identical stdout/stderr plus typed
errors, locking behavior before later slices migrate callers.

* refactor: tighten emitter API and cover pagination with real tests

- split Emitter.Success/PartialFailure and drop EmitOptions.OK so a
  missing ok flag can no longer silently emit ok:false
- give StreamPage its own StreamOptions (format + pretty) instead of
  reusing EmitOptions, making "jq needs aggregation" a compile-time fact
- pin the Emitter jq-error contract (returns error, writes no stderr);
  the caller adapter re-emits the legacy stderr line on migration
- add in-package tests driving the real apiPaginate/servicePaginate over
  a mock transport: multi-page aggregation, empty-result fallback,
  MarkRaw handling, and the business-error raw-response red line

* test: use standard TestFactory harness for pagination tests

Replace the hand-rolled RoundTripper + APIClient construction in the
apiPaginate/servicePaginate tests with cmdutil.TestFactory and its
httpmock.Registry, and isolate LARKSUITE_CLI_CONFIG_DIR to t.TempDir(),
matching the repo's standard HTTP-mocked test convention. Assertions and
coverage (multi-page aggregation, empty-result fallback, MarkRaw, and the
business-error raw-response red line) are unchanged.

* refactor: route success output through the single Emitter port

Migrate the success-output surfaces onto internal/output's Emitter,
byte-for-byte identical (proven by frozen golden diffs and the real
paginate/HandleResponse tests):

- RuntimeContext.Out/OutRaw/OutFormat/OutFormatRaw/OutPartialFailure now
  build an Emitter and call Success/PartialFailure; emit and outFormat are
  removed. An adapter maps the returned error back to the legacy
  outputErrOnce / jq-error stderr / exit-code behavior.
- WriteSuccessEnvelope degrades to a thin Emitter.Success delegate; its 8
  callers are unchanged.
- apiPaginate/servicePaginate stream pages via Emitter.StreamPage; the
  aggregate and business-error raw-response branches are untouched.
- HandleResponse routes its non-JSON structured-response branch through
  Emitter.Success.

Frozen golden fixtures replace the runtime legacy oracles so the
differential harness cannot go self-referential after migration.

* fix: keep _notice on struct payloads in Emitter's unknown-format fallback

printLegacyDataJSON now normalizes via toGeneric first (matching FormatValue), so a struct / named-map payload retains its injected _notice on the unknown-format -> JSON fallback rather than dropping it silently. Add a regression test that fails against the pre-fix path.

* refactor: make the Emitter own write failures and stop mutating inputs

Route every Emitter stdout path through a render-to-buffer-then-copy helper so a marshal/render failure leaves stdout empty and surfaces a typed internal error (with cause), and a stdout write failure is propagated instead of silently swallowed. Leaf writers gain error-returning Write* cores; the legacy Print*/FormatValue wrappers keep their exact behavior for unmigrated callers.

- handleEmitterError now captures every error, not only the jq/safety branches; flip OutRaw's write-error test to assert propagation.
- Clone the map before injecting _notice so a caller's payload is never mutated and an existing _notice is never overwritten.
- Preserve jq's own typed error (validation/api) on a bad expression or runtime failure; only wrap genuine stdout write failures.
- Split tests: normative emitter_contract_test.go vs frozen emitter_legacy_compat_test.go (base SHA recorded, self-update env vars removed).

* fix: satisfy license-header and forbidigo lint on the emitter changes

- Move the base-SHA note below the copyright header in the renamed legacy-compat test so the license-header check sees a valid header at the top.
- Route the leaf wrappers' marshal/format stderr messages through a single legacyStderrf helper (one //nolint:forbidigo) instead of bare os.Stderr, preserving exact legacy behavior for unmigrated direct callers while passing forbidigo; drop the now-unused os imports.

* fix: stop legacy CSV wrappers reporting write failures to stderr

Align FormatAsCSV/FormatAsCSVPaginated and FormatValue/FormatPage's CSV branch with the other leaf wrappers: report only marshal failures, swallow write failures. Previously they emitted a 'csv write error' for the (empty) line and the JSON-fallback write failures that the pre-refactor code ignored, and mislabeled a JSON write failure as a CSV one. Failure-path only; success output is unchanged (golden double-diff still byte-for-byte).
2026-07-21 14:32:47 +08:00
HanShaoshuai-k 577ff035c3 fix: allow jq examples in quality gate dry-runs 2026-07-21 14:07:57 +08:00
luozhixiong01 d8fb368ce4 test: isolate unit tests from user state (#1883) 2026-07-20 22:22:39 +08:00