* fix(apps): friendly-ize "Container not exists" for observability commands
+metric-list and +analytics-list passed the upstream business code 400002655
("Container not exists") through verbatim. The message reads like an
infrastructure fault and misleads callers (including AI agents) into retrying a
non-retryable, expected business state: an app with no running container simply
has no metrics/analytics to query yet.
Rewrite it at a scoped observability helper (withObservabilityHint) into a
user-facing explanation plus a deploy-then-retry next step, mirroring the
existing isAppNoDatabaseError override. Detection is code-OR-message so a server
renumber alone does not silently drop the rewrite. Classification, code, and the
wrapped cause are preserved; unrelated failures still fall through to the shared
app-id recovery hint (and its own no-database override).
* test(apps): add execution-path regression tests for observability container hint
common_test.go proves withObservabilityHint in isolation but stays green if a
call site reverts to withAppsHint. Drive +metric-list and +analytics-list
Execute with a mocked 400002655 "Container not exists" envelope and assert the
container-specific message/hint/code, so a revert fails the build. Also closes
the two uncovered call-site lines flagged by coverage.
* fix(errclass): classify no-container code as validation/failed_precondition
Register 400002655 in sparkCodeMeta mirroring its no-database twin
(400002465) so both "expected precondition not met" business states expose
the same validation/failed_precondition classification to machine consumers,
instead of falling back to api/unknown. The shortcut-layer message rewrite
already keyed off the raw code, so this only aligns the typed envelope's
category/subtype; update the execution-path tests to pin the new
classification.
* fix(apps): gate the no-container hint's release behind user authorization
The no-container hint told a harness to deploy via +release-create, a "write"
that takes the whole app live and can affect existing production traffic —
without the user-confirmation gate its no-database twin deliberately carries.
Since the hint's audience is an AI agent that acts on it, a failed metrics read
could trigger an unconfirmed go-live. Lead with a read-only +release-list
status check and gate +release-create behind an explicit user confirmation,
mirroring appNoDatabaseHint.
* feat: add apps database sync shortcuts
Add Base-to-database sync shortcuts for preview, create, list, get, enable, disable, update, and delete flows.
Cover OpenAPI request contracts, typed sync error classification, dry-run E2E coverage, and lark-apps skill guidance.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(apps): send db-sync task_id and config in request body
The enable/disable/delete/update sync commands placed task_id (and
update's config) in query params, but the OpenAPI contract binds these
fields via api.json (request body). BOE testing returned
"field validation failed" (99992402) because the body was empty.
Move task_id to the request body for enable/disable/delete, and move
both task_id and config to the body for update. Dry-run previews now
render these under body, and unit tests pin the body binding so a
regression to query params fails.
* fix(apps): use POST for db-sync-delete action endpoint
The delete command issued an HTTP DELETE to db/sync_del, but the
action-style endpoint is registered as POST (like sync_create and
sync_disable). The method mismatch made the gateway return a plaintext
404, surfacing as "API returned a non-object JSON response".
Switch the request and dry-run preview to POST, and pin the method in
the delete unit tests so a regression to DELETE fails.
* test: pin db-sync update base_url as optional contract
* test: pin db-sync update omits base_url without silent default
* docs(skills): clarify db-sync source.base_url create-required update-optional contract
* fix(apps): send db-sync env in request body not query params
The +db-sync-create and +db-sync-update endpoints read env from the
request body (peer of config/preview/task_id), not the query string.
Placing env in query params left the body env empty, so the server
treated every request as online and rejected DDL operations
(code 500002776: forbid ddl/dcl operation in online env), making it
impossible to create/update sync tasks against a dev environment.
Move env into the request body via a new dbEnvBody helper that mirrors
dbEnvParams' omit-empty contract, so unset env still lets the server
auto-select the branch. Pin the contract in unit and e2e dry-run tests
by asserting body.env and that env is absent from query params.
* test: align db-sync operate/delete e2e with request-body contract
The enable/disable/delete dry-run e2e still asserted the pre-migration
wire shape: delete on DELETE and task_id in query params. The shortcuts
now POST these actions with task_id in the request body (commits moving
task_id and the delete verb), so the stale assertions failed against a
current binary.
Assert POST + body.task_id and that task_id is absent from query params,
pinning the same body-over-query contract the env fix established.
* fix(apps): improve db-sync create ergonomics and error guidance
Refine +db-sync-create/update validation, error hints, and docs so AI
agents recover from common Base-to-database sync failures without guessing:
- source.table.name: document that a user-named table must be set, name
takes precedence over the base_url ?table= token; fix test fixtures that
used a fictional source.table.url instead of source.base_url.
- Preflight source table locate: reject create locally when base_url has no
?table= and source.table.name is empty, pointing at base +table-list.
- Online DDL ban: attach a precise hint for code 500002776 + subcode
k_dl_4000001 telling multi-env apps to create tables on --environment dev.
- Missing record-id column: extend the 500002783 hint to add a unique text
column via +db-execute before retrying.
- Optional field_maps on create: allow omitted or empty field_maps so the
server auto-matches and creates the task; keep update requiring an enabled
mapping and still reject an all-disabled array.
- Environment default: db-sync commands use online when --environment is
omitted; align help text, comments, and skill docs.
* fix(apps): migrate db-sync error codes to the 4xx client-error range
The backend moved the seven db-sync error codes from the 5000027xx
server-error range to the 4000024xx client-input range to reflect that
they are client-input errors. Mirror the new codes in the CLI so error
classification and recovery hints keep matching:
- 500002783 -> 400002477 (mapping invalid)
- 500002784 -> 400002478 (target schema mismatch)
- 500002785 -> 400002479 (operation not allowed)
- 500002786 -> 400002480 (task not found)
- 500002787 -> 400002481 (invalid task id)
- 500002788 -> 400002482 (source table not found)
- 500002789 -> 400002483 (target table not found)
Category, subtype, hint text, and behavior are unchanged; 500002776
(online DDL ban) is untouched.
* fix(apps): tighten db-sync preview validation and pretty output
Address review follow-ups on the db-sync shortcuts:
- +db-sync-get pretty output no longer prints <nil> for a missing
schema_only nor Go map syntax for statistics; render a bare bool and
deterministic key=value pairs instead.
- Reject a non-array field_maps in +db-sync-create --preview as well as
commit, so the malformed shape is caught locally rather than forwarded
to the backend.
- Clarify in lark-apps-db.md that +db-sync-create --preview needs no
confirmation and only a real create requires --yes.
- Harden the db-sync dry-run validation tests to assert exit code 2 and
the structured stderr envelope (type/subtype/param), and add coverage
for the preview non-array field_maps rejection and batch pretty output.
* fix(apps): guard db-sync preview output and neutralize update hint
Address the next db-sync review round:
- +db-sync-create --preview --output no longer writes a "null" file and
exits success when the response omits data.config; project config into
a typed object and return internal/invalid_response without writing.
- Make the 400002482 code hint command-neutral so +db-sync-update is not
steered into a create-only recovery path that risks duplicate tasks.
- lark-apps-db.md: carry --environment on the update lifecycle examples
and split failure recovery by streaming (can update) vs batch (cannot
update; recreate instead), removing the batch/update contradiction.
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(apps): detect the no-database failure by code or message
The recovery flow for "db command against an app that has no database"
keyed on one business code (500002759). The server has since renumbered
that case to 400002465, which silently disabled the flow: users now see
the raw internal message about workspace / app-id mapping and lose the
cloud-development recovery steps entirely.
Nothing catches the regression. There is no compile error, the unit tests
compare against the same constant they set, and the dry-run E2E does not
exercise a real response — the failure only shows up against a server that
has already renumbered.
Detect on code OR message instead. Both known codes are kept, plus narrow
lowercase markers of the server's internal wording. The two channels have
opposite failure modes: a code is precise but gets renumbered, a message
survives renumbering but breaks on rewording or localization. Requiring
either to match means one channel changing degrades nothing, and only a
simultaneous change of both regresses.
Markers stay deliberately narrow. "no db branch" in particular must not
also swallow env-pull's "invalid db branch" case, which needs its own
hint; a comment records that widening them requires a test proving the
neighbours still pass through.
Classification and the cause chain are untouched: the helper still mutates
the problem in place and returns the same error value.
* test(apps): assert the full typed-error contract in no-database cases
Review feedback: the new subtests checked only Message and Hint, so a
change that reclassified the failure — or replaced the error value and
dropped the cause chain — would still have passed.
Each case now asserts Category, Subtype and Code are untouched by the
rewrite, and that the helper returns the same error value. Inputs use a
concrete subtype rather than Unknown, so a clobbered classification is
actually observable. One new case wraps a cause and asserts errors.Is
still finds it through the rewrite.
Also covers the predicate's defensive nil guard, which withAppsHint cannot
reach on its own (ProblemOf returns ok=false for untyped errors), closing
the two uncovered lines the coverage report flagged. Both withAppsHint and
isAppNoDatabaseError are now at 100%.
* fix(errclass): classify the db-domain business codes
Three codes reaching the Apps db commands were absent from the Spark
table, so BuildAPIError fell through to the CategoryAPI + SubtypeUnknown
catch-all and the envelope carried no usable classification.
"App has no database yet" registers as Validation / FailedPrecondition:
the app resolves fine and the request is well-formed, but a prerequisite
the caller must create first is missing, so retrying unchanged can never
succeed. This moves its exit code from 1 to 2 — "fix the state" rather
than "the call failed" — and a test pins that so a future
reclassification has to be deliberate. Two codes cover it because the
server renumbered the case into the 4xx band; the legacy one stays for
older servers.
"Table does not exist" registers as API / NotFound, an ordinary
missing-resource lookup with no exit-code change. SubtypeNotFound has no
APIHint default, which matters here: the Apps layer fills its
command-scoped hint only when the classifier left Hint empty, so a
context-free default would displace the more actionable one. A test
guards that too.
Failures from the minutes and video-conference commands now surface as
structured, typed errors carrying a stable category and subtype — spanning
input validation, missing permissions, network and file-I/O failures, and
remote API errors — so callers can branch on the error kind instead of
parsing free-form text. Batch commands report partial failures explicitly,
emitting per-item results with a non-zero exit instead of masking them.
Calendar commands now return structured, typed error envelopes for every
failure mode — input validation, internal faults, and API responses —
instead of legacy generic errors. Callers and AI agents get consistent
exit codes and a machine-readable shape (type / subtype / code / hint),
and can tell bad input, an internal fault, and an API rejection apart.
Validation errors are attributed to the offending flag.
Server-supplied error details (e.g. why an event time was rejected) are
surfaced on the typed error's hint via a shared classifier improvement
that benefits every domain. Multi-step operations (create-with-attendees
rollback, multi-field update) preserve the real failure's classification
and report which steps completed.
The whole calendar domain is now lint-locked against reintroducing legacy
error constructors.
* feat(mail): return typed error envelopes across the mail domain
Replace every produced error path in shortcuts/mail with typed errs.* envelopes, so consumers get stable category, subtype, param/params, hint, retryable, and log_id metadata for classification and recovery instead of free-form message text.
- Locally constructed mail errors move from output.Err* / output.Errorf / final fmt.Errorf / common legacy helpers to errs.* builders, with structured params on multi-flag validation and failed-precondition states kept non-retryable.
- API-call failures move from runtime.CallAPI / DoAPIJSON legacy boundaries to runtime.CallAPITyped or runtime.ClassifyAPIResponse, and mail-specific enrichers read errs.ProblemOf so typed code, subtype, hint, and log_id metadata are preserved.
- Batch draft-send partial failures now use runtime.OutPartialFailure so successful and failed draft sends stay in stdout while the command exits through a typed multi-status signal.
- Add mail-domain typed helpers, mail API code metadata, and guard wiring to keep shortcuts/mail from reintroducing legacy envelopes or legacy API calls.
- Keep genuine intermediate fmt.Errorf wraps in parser/builder layers annotated with nolint comments; command-facing paths wrap them into typed validation, API, network, or internal errors.
* fix(mail): report aborted draft-send batches as a single failure result
When an account-level failure interrupts a batch send after some drafts
already went out, the command previously produced two machine-readable
failure results: the partial-failure ledger on stdout and a second error
envelope on stderr. Consumers could not tell which one to recover from.
The batch ledger is now the only failure result for that case: it gains
aborted and abort_error fields carrying the typed cause, so callers can
see which drafts were sent, which failed, why the batch stopped, and how
to recover — all from stdout. A --stop-on-error stop keeps these fields
unset because stopping early there is the caller's own choice.
Drive-domain errors now leave the CLI as typed, machine-branchable
envelopes — a stable `type` plus `subtype` and named fields (param,
params, retryable, log_id, hint) — so scripts and AI agents can branch on
structure and act on a recovery hint instead of parsing prose.
Changes:
- Every error produced in the drive domain — validation, file I/O, and the
failures returned from its Lark API calls — is emitted as a typed errs.*
error; the exit code is derived from the error category. Drive's API calls
now go through a shared typed classifier, so failures carry subtype,
troubleshooter, a recovery hint, and the request's log_id whether the
server returns it in the response body or the x-tt-logid header; an
already-typed network/auth error is never downgraded into a generic API
error.
- Known API conditions (resource conflict, cross-tenant, cross-brand, ...)
carry a recovery hint keyed by their error class; a command can refine
that hint with command-specific guidance.
- Batch partial failures (+push / +pull / +sync, where some items succeed
and some fail) now report an honest ok:false multi-status result on
stdout — the summary and every per-item outcome stay machine-readable —
and exit non-zero, instead of a misleading ok:true success envelope.
- Duplicate rel_path conflicts report each colliding path as a structured
params entry (RFC 7807 invalid-params style).
- Static guards lock the drive path so legacy error construction — direct
envelopes or the auto-classifying API helpers — cannot be reintroduced,
making drive the template for the remaining domains.
Output changes worth noting for consumers:
- Error envelopes now carry typed type/subtype and named fields; exit
codes follow the error category (malformed or incomplete API responses
are reported as internal errors rather than generic API errors).
- Batch partial failures (+push / +pull / +sync) emit an ok:false result
envelope on stdout (summary + per-item items[]) and exit non-zero; the
per-item results stay on stdout rather than in a stderr error envelope.
Errors surfaced through shared cross-domain helpers (scope precheck, media
import upload, metadata lookup, save-path resolution) are not yet typed;
they migrate with the shared layer in a follow-up change.
Every failure on the authentication, authorization, and configuration
path now surfaces as a typed structured error instead of an ad-hoc
envelope. Users and scripts that consume CLI output get:
- a fixed nine-category taxonomy on the wire, each mapped to a
stable shell exit code (authentication/authorization/config = 3,
network = 4, internal = 5, policy = 6, confirmation = 10)
- identity-aware detail fields (missing_scopes, requested_scopes,
granted_scopes, console_url, log_id, retryable, hint) carried
uniformly on the envelope
- a single canonical policy envelope at exit 6; the legacy
auth_error carve-out is retired
- per-subtype canonical message + hint that preserves Lark's
diagnostic phrasing and routes recovery to the right actor:
app developer (app_scope_not_applied), user (missing_scope,
token_scope_insufficient, user_unauthorized), or tenant admin
(app_unavailable, app_disabled)
- wrong app credentials classify as config/invalid_client whether
surfaced by the Open API endpoint (99991543) or the tenant
access-token mint endpoint (10003 / 10014), instead of
collapsing to a transport error or api/unknown
- local shortcut scope preflight emits the same
authorization/missing_scope envelope (identity + deterministic
missing-scope set) used by the post-call permission path, so AI
consumers read the same structured shape from precheck and from
server-returned permission denial
- streaming download/upload failures keep the same network subtype
split (timeout / TLS / DNS / transport) as the non-stream path
instead of collapsing every cause to a generic transport failure
- console_url is carried only on the bot-perspective
app_scope_not_applied envelope (where the recovery action is
"developer applies the scope at the developer console"); the
user-perspective missing_scope envelope drops the field, since
the only actionable user recovery is `lark-cli auth login --scope`
and pointing an end user at a console they cannot modify is
misleading
- bind workflows (Hermes / OpenClaw / lark-channel) flatten dynamic
Type tags to wire 'config' with the original module name kept
as a metric label
All 10 typed errors are cause-bearing, nil-safe on .Error() and
.Unwrap(), and defensively clone slice setter inputs. Four lint
rules (CheckNilSafeError / CheckBuilderImmutable / CheckUnwrapSymmetry
/ CheckBuildAPIErrorArms) lock these invariants on migrated paths.
Introduce a typed error contract framework for lark-cli so in-process
Go callers can branch via errors.As(&errs.XxxError{}) and shell scripts,
AI agents, and protocol adapters can branch on stable JSON type/subtype
fields instead of regex-parsing free-form messages.
Adds:
- Canonical taxonomy under errs/ (9 categories + typed Error structs
embedding a shared Problem, RFC 7807-aligned)
- Centralized Lark code metadata + identity-aware BuildAPIError dispatch
- Typed JSON envelope writer alongside the legacy envelope writer
- MCP / OAuth (RFC 6750 Bearer) projection adapters
- Five CI lint guards preventing ad-hoc taxonomy drift
Backward compatibility: legacy *output.ExitError producers (ErrAPI,
ErrWithHint, Errorf, ErrBare) and business shortcuts that use them
continue to render the legacy envelope unchanged. SecurityPolicyError
wire format and exit code are preserved via a carve-out; taxonomy
migration is deferred to PR 2. Domain-specific business migration is
staged across PR 3+.
Framework-direct paths now return typed *errs.*Error: ErrAuth /
ErrValidation / ErrNetwork emit category literals on the wire
(authentication / validation / network), *core.ConfigError is promoted
at the cmd/root boundary with exit code aligned from 2 to 3, and Lark
API permission denials classified by BuildAPIError exit 3.
At the SDK boundary, WrapDoAPIError preserves any already-classified
error (legacy *output.ExitError or typed *errs.*) so output.ErrAuth
from missing credentials surfaces with the auth category and exit 3
intact instead of being downgraded to a network error. Policy responses
classified by BuildAPIError (codes 21000 / 21001) extract challenge_url
and the canonical hint from the response body, matching what the
auth transport already surfaces at the HTTP layer; non-https
challenge URLs are dropped.
First PR in the feat/error-contract-* series.