Commit Graph

6 Commits

Author SHA1 Message Date
chenxingyang1019 e0e90a4e1b fix(apps): classify the online DDL ban and the file storage quota failure (#2460)
* fix(apps): classify the online DDL/DCL ban on +db-execute

Running DDL against the online branch of a multi-env app came back as
api/server_error with exit 1, hinted "fix the SQL and re-run", and carried a
statement position that did not exist. All three point the caller the wrong way:

  - server_error means "upstream 5xx, retryable"; this is a product rule and no
    number of retries changes it;
  - the SQL is fine — the target environment is what has to change;
  - "(at statement 1 of 1)" is fabricated. The server pre-validates the whole
    batch and returns a single ERROR sentinel, so a 5-statement request with the
    DDL in position 4 still rendered as "1 of 1". The CLI does not split the SQL,
    so it cannot know the real count and cannot detect the mismatch generally —
    only codes known to be batch-level rejections can drop the suffix.

Give code 4000001 its own arm: validation/failed_precondition (exit 2, "change
the environment, do not retry"), a hint pointing at dev plus +db-env-migrate, and
no statement position. Every other code keeps its current classification, wording
and position suffix; a test pins that.

4000001 is a dedicated server-side code (ErrOnlineEnvForbidDDLDCL, client-error
band), raised only by the pre-validation pass when env==online on a multi-env
workspace. Syntax errors and PG errors use different codes, so keying on it is
safe. Matching on the numeric value also covers the "k_dl_4000001" wire form,
since codeString already strips that prefix — both forms are tested.

"No statements were applied" is stated rather than inferred here: the validator
walks every statement and rejects the batch on the first DDL, so nothing lands.
The default arm would have inferred the opposite for a DDL in a later position
("Earlier statements were committed"), which is wrong for this code.

* fix(apps): classify tenant file storage quota exceeded

+file-upload against a tenant whose file storage is full returned api/unknown with
no hint at all, so a caller could not tell "the quota is full, stop" from "the
upstream had a bad minute, retry" — and had no next step either.

Register 400000055 as api/quota_exceeded. That is a dedicated server-side code
(ErrTenantStorageQuotaExceeded, client-error band) raised only on the upload path,
and the subtype already carries "retrying will not help", so the framework's
existing quota wording is enough and no domain-specific hint is added.

Left in CategoryAPI (exit 1) rather than Validation (exit 2): a full quota is not
something a different argument fixes, and exit 2 would imply it is.

The test asserts the hint is non-empty on purpose. The wording comes from the
shared APIHint table, so if quota_exceeded is ever dropped from there this fails
and says the code now needs its own wording, instead of silently shipping an
empty hint.
2026-08-24 16:49:33 +08:00
chenxingyang1019 ca35f60616 fix(apps): make cache-clear ask first, and make apps failures classifiable (#2415)
* docs(skills): require explicit confirmation before apps +cache-clear

Asked to clear an app's online cache, an agent read `Risk: high-risk-write` from
--help and then supplied `--yes` itself on the first call, wiping production
cache without ever hitting the confirmation gate.

The CLI gate is fine: no --yes -> exit 10 confirmation_required, and --dry-run ->
exit 0 without triggering it. The wording was not. It only forbade appending
`--yes` *after* an exit-10, and said "已明确授权可直接带 --yes" without defining
authorization — so "clear my cache" read as authorization.

- `+cache-clear` gets a CAUTION block: never self-supply `--yes` on the first
  call; without confirmation, either --dry-run or ask, then stop and wait. exit
  10 is not a signal to retry with --yes.
- Add a zero-ambiguity table separating a *request* to clear ("clear the online
  cache") from a *confirmation* ("我确认清 dev"), so blocking the accidental wipe
  does not also kill the cases that were already correct: an explicit
  confirmation still goes straight to `--yes`, and a request with no environment
  named still has to ask instead of picking one.
- Note that online needs a confirmation phrase even when named explicitly.

`+cache-delete` gains the response field an agent has to read
(`deleted_key_count`): 0 means the key never existed, not "deleted
successfully", plus the get -> delete -> get chain needed to prove a delete took
effect — a single miss afterwards cannot tell the two apart.

SKILL.md: add +cache-clear to 禁止预授权判定底线, the one list a pre-authorized
run cannot skip; a reference-level rule alone would be bypassed there. The
routing table is left alone — no other row annotates risk, including
+file-delete, +role-delete and +member-remove.

* fix(apps): stop attaching request-shaped hints to precondition failures

`+db-execute` against a tenant that never activated Miaoda returns code 221800
"miaoda UAT not activated" with the hint "verify table/column names with
`+db-table-get` ... target the dev database with --environment dev". Neither step
can help: the failure is tenant-level, so a caller following the hint loops over
table lookups and env retries that fail identically.

Two causes. 221800 was unregistered, so it degraded to api/unknown — nothing in
the envelope distinguished "your tenant is not activated, stop" from "your SQL
was wrong, fix it and retry". And withAppsHint filled the caller's hint whenever
the server sent none, without looking at what failed: the hints are
command-scoped ("verify --app-id", "verify table/column names", "list releases"),
so every one of them describes the request, and the request is exactly what
failed_precondition says was fine.

Register 221800 as validation/failed_precondition (same shape as 400002465 "app
has no database yet") and gate the hint fallback on the subtype.

Blast radius is two codes, since that is all the subtype covers here:
  - 221800 — now withheld; message and code still carry the meaning.
  - 400002655 "no running container" — only when it reaches a non-observability
    command; the observability pair rewrites it first, and "verify --app-id" was
    never the fix for an undeployed app.
400002465 / 500002759 are intercepted by the isAppNoDatabaseError branch above
the gate, and 400002479 is served by withDBSyncHint, which does not delegate
here. The other 78 call sites take the original path for every input.

Gate on the one subtype, not on Category: this package asserts on purpose that an
authentication failure on +role-list (99991663) keeps the app-access hint and a
503 on credential issuance keeps the developer-access hint. Those hints are broad
enough to survive a caller-standing failure; only the precondition class is
misdescribed by construction. A test pins that, so widening the gate to Category
fails loudly instead of silently dropping those hints.

The gate is asserted on the real classification path (BuildAPIError -> the code
table -> withAppsHint), not only on a hand-built Problem. Constructing
SubtypeFailedPrecondition directly feeds the gate the input it wants and passes
whether or not 221800 is registered, so the registration itself has to be part of
what the test covers.

No recovery hint for 221800 — the activation path is a product procedure, and
guessing one is what made this failure misleading in the first place.

* fix(apps): classify file-storage and app-level failures

Five Spark business codes reached the CLI unregistered, so every one of them came
out as api/unknown with exit 1: a caller could not tell "your app id is wrong"
from "you lack permission" from "the upstream is having a bad minute", and the
exit code offered no way to branch either.

  400002484  app not found            -> validation/invalid_argument     exit 2
  400002467  no admin/developer perm  -> authorization/permission_denied exit 3
  500002761  ditto, pre-4xx renumber  -> same
  400000034  file not found/no access -> api/not_found                   exit 1
  500000034  ditto, pre-4xx renumber  -> same

400002467 is not file-specific: db commands (+db-table-list, +db-table-get,
+db-quota-get, +db-changelog-list) return it for an app the caller cannot access,
so registering it fixes both domains at once.

400002484 covers a well-formed id that does not exist AND a malformed one
("notanappid", "app_1" return it too), so the argument itself is the failure ->
invalid_argument, whose exit 2 separates "you passed the wrong id" from an
upstream fault. Environments that have not picked it up answer with 400002465
instead, conflating it with "app has no database yet"; the CLI cannot tell those
apart on the old code, so nothing here keys on that.

Both the current and the pre-4xx number are registered for each file failure.
The domain is moving its client-class errors from the 5xxxxxxxx band into
4xxxxxxxx, rolled out per environment, so both are live at once and dropping the
old one would silently return the un-migrated half to api/unknown — the same trap
that made the no-database recovery flow disappear when the server renumbered it
(see appNoDatabaseCode). The new number is not derivable from the old either:
500002761 became 400002467, tail digits included.

No hints added: permission_denied already has framework recovery wording, and a
domain-specific one would have to invent a remedy.
2026-08-20 17:36:08 +08:00
木杉 1e87f67244 fix(apps): friendly-ize "Container not exists" for observability commands (#2302)
* fix(apps): friendly-ize "Container not exists" for observability commands

+metric-list and +analytics-list passed the upstream business code 400002655
("Container not exists") through verbatim. The message reads like an
infrastructure fault and misleads callers (including AI agents) into retrying a
non-retryable, expected business state: an app with no running container simply
has no metrics/analytics to query yet.

Rewrite it at a scoped observability helper (withObservabilityHint) into a
user-facing explanation plus a deploy-then-retry next step, mirroring the
existing isAppNoDatabaseError override. Detection is code-OR-message so a server
renumber alone does not silently drop the rewrite. Classification, code, and the
wrapped cause are preserved; unrelated failures still fall through to the shared
app-id recovery hint (and its own no-database override).

* test(apps): add execution-path regression tests for observability container hint

common_test.go proves withObservabilityHint in isolation but stays green if a
call site reverts to withAppsHint. Drive +metric-list and +analytics-list
Execute with a mocked 400002655 "Container not exists" envelope and assert the
container-specific message/hint/code, so a revert fails the build. Also closes
the two uncovered call-site lines flagged by coverage.

* fix(errclass): classify no-container code as validation/failed_precondition

Register 400002655 in sparkCodeMeta mirroring its no-database twin
(400002465) so both "expected precondition not met" business states expose
the same validation/failed_precondition classification to machine consumers,
instead of falling back to api/unknown. The shortcut-layer message rewrite
already keyed off the raw code, so this only aligns the typed envelope's
category/subtype; update the execution-path tests to pin the new
classification.

* fix(apps): gate the no-container hint's release behind user authorization

The no-container hint told a harness to deploy via +release-create, a "write"
that takes the whole app live and can affect existing production traffic —
without the user-confirmation gate its no-database twin deliberately carries.
Since the hint's audience is an AI agent that acts on it, a failed metrics read
could trigger an unconfirmed go-live. Lead with a read-only +release-list
status check and gate +release-create behind an explicit user confirmation,
mirroring appNoDatabaseHint.
2026-08-12 15:01:10 +08:00
jinjiuzhe 158d15b3fd feat: add apps database sync shortcuts for Base-to-database import (#2251)
* feat: add apps database sync shortcuts

Add Base-to-database sync shortcuts for preview, create, list, get, enable, disable, update, and delete flows.

Cover OpenAPI request contracts, typed sync error classification, dry-run E2E coverage, and lark-apps skill guidance.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(apps): send db-sync task_id and config in request body

The enable/disable/delete/update sync commands placed task_id (and
update's config) in query params, but the OpenAPI contract binds these
fields via api.json (request body). BOE testing returned
"field validation failed" (99992402) because the body was empty.

Move task_id to the request body for enable/disable/delete, and move
both task_id and config to the body for update. Dry-run previews now
render these under body, and unit tests pin the body binding so a
regression to query params fails.

* fix(apps): use POST for db-sync-delete action endpoint

The delete command issued an HTTP DELETE to db/sync_del, but the
action-style endpoint is registered as POST (like sync_create and
sync_disable). The method mismatch made the gateway return a plaintext
404, surfacing as "API returned a non-object JSON response".

Switch the request and dry-run preview to POST, and pin the method in
the delete unit tests so a regression to DELETE fails.

* test: pin db-sync update base_url as optional contract

* test: pin db-sync update omits base_url without silent default

* docs(skills): clarify db-sync source.base_url create-required update-optional contract

* fix(apps): send db-sync env in request body not query params

The +db-sync-create and +db-sync-update endpoints read env from the
request body (peer of config/preview/task_id), not the query string.
Placing env in query params left the body env empty, so the server
treated every request as online and rejected DDL operations
(code 500002776: forbid ddl/dcl operation in online env), making it
impossible to create/update sync tasks against a dev environment.

Move env into the request body via a new dbEnvBody helper that mirrors
dbEnvParams' omit-empty contract, so unset env still lets the server
auto-select the branch. Pin the contract in unit and e2e dry-run tests
by asserting body.env and that env is absent from query params.

* test: align db-sync operate/delete e2e with request-body contract

The enable/disable/delete dry-run e2e still asserted the pre-migration
wire shape: delete on DELETE and task_id in query params. The shortcuts
now POST these actions with task_id in the request body (commits moving
task_id and the delete verb), so the stale assertions failed against a
current binary.

Assert POST + body.task_id and that task_id is absent from query params,
pinning the same body-over-query contract the env fix established.

* fix(apps): improve db-sync create ergonomics and error guidance

Refine +db-sync-create/update validation, error hints, and docs so AI
agents recover from common Base-to-database sync failures without guessing:

- source.table.name: document that a user-named table must be set, name
  takes precedence over the base_url ?table= token; fix test fixtures that
  used a fictional source.table.url instead of source.base_url.
- Preflight source table locate: reject create locally when base_url has no
  ?table= and source.table.name is empty, pointing at base +table-list.
- Online DDL ban: attach a precise hint for code 500002776 + subcode
  k_dl_4000001 telling multi-env apps to create tables on --environment dev.
- Missing record-id column: extend the 500002783 hint to add a unique text
  column via +db-execute before retrying.
- Optional field_maps on create: allow omitted or empty field_maps so the
  server auto-matches and creates the task; keep update requiring an enabled
  mapping and still reject an all-disabled array.
- Environment default: db-sync commands use online when --environment is
  omitted; align help text, comments, and skill docs.

* fix(apps): migrate db-sync error codes to the 4xx client-error range

The backend moved the seven db-sync error codes from the 5000027xx
server-error range to the 4000024xx client-input range to reflect that
they are client-input errors. Mirror the new codes in the CLI so error
classification and recovery hints keep matching:

- 500002783 -> 400002477 (mapping invalid)
- 500002784 -> 400002478 (target schema mismatch)
- 500002785 -> 400002479 (operation not allowed)
- 500002786 -> 400002480 (task not found)
- 500002787 -> 400002481 (invalid task id)
- 500002788 -> 400002482 (source table not found)
- 500002789 -> 400002483 (target table not found)

Category, subtype, hint text, and behavior are unchanged; 500002776
(online DDL ban) is untouched.

* fix(apps): tighten db-sync preview validation and pretty output

Address review follow-ups on the db-sync shortcuts:

- +db-sync-get pretty output no longer prints <nil> for a missing
  schema_only nor Go map syntax for statistics; render a bare bool and
  deterministic key=value pairs instead.
- Reject a non-array field_maps in +db-sync-create --preview as well as
  commit, so the malformed shape is caught locally rather than forwarded
  to the backend.
- Clarify in lark-apps-db.md that +db-sync-create --preview needs no
  confirmation and only a real create requires --yes.
- Harden the db-sync dry-run validation tests to assert exit code 2 and
  the structured stderr envelope (type/subtype/param), and add coverage
  for the preview non-array field_maps rejection and batch pretty output.

* fix(apps): guard db-sync preview output and neutralize update hint

Address the next db-sync review round:

- +db-sync-create --preview --output no longer writes a "null" file and
  exits success when the response omits data.config; project config into
  a typed object and return internal/invalid_response without writing.
- Make the 400002482 code hint command-neutral so +db-sync-update is not
  steered into a create-only recovery path that risks duplicate tasks.
- lark-apps-db.md: carry --environment on the update lifecycle examples
  and split failure recovery by streaming (can update) vs batch (cannot
  update; recreate instead), removing the batch/update contradiction.

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-11 18:21:32 +08:00
chenxingyang1019 6402080328 fix(apps): detect the no-database failure by code or message (#2217)
* fix(apps): detect the no-database failure by code or message

The recovery flow for "db command against an app that has no database"
keyed on one business code (500002759). The server has since renumbered
that case to 400002465, which silently disabled the flow: users now see
the raw internal message about workspace / app-id mapping and lose the
cloud-development recovery steps entirely.

Nothing catches the regression. There is no compile error, the unit tests
compare against the same constant they set, and the dry-run E2E does not
exercise a real response — the failure only shows up against a server that
has already renumbered.

Detect on code OR message instead. Both known codes are kept, plus narrow
lowercase markers of the server's internal wording. The two channels have
opposite failure modes: a code is precise but gets renumbered, a message
survives renumbering but breaks on rewording or localization. Requiring
either to match means one channel changing degrades nothing, and only a
simultaneous change of both regresses.

Markers stay deliberately narrow. "no db branch" in particular must not
also swallow env-pull's "invalid db branch" case, which needs its own
hint; a comment records that widening them requires a test proving the
neighbours still pass through.

Classification and the cause chain are untouched: the helper still mutates
the problem in place and returns the same error value.

* test(apps): assert the full typed-error contract in no-database cases

Review feedback: the new subtests checked only Message and Hint, so a
change that reclassified the failure — or replaced the error value and
dropped the cause chain — would still have passed.

Each case now asserts Category, Subtype and Code are untouched by the
rewrite, and that the helper returns the same error value. Inputs use a
concrete subtype rather than Unknown, so a clobbered classification is
actually observable. One new case wraps a cause and asserts errors.Is
still finds it through the rewrite.

Also covers the predicate's defensive nil guard, which withAppsHint cannot
reach on its own (ProblemOf returns ok=false for untyped errors), closing
the two uncovered lines the coverage report flagged. Both withAppsHint and
isAppNoDatabaseError are now at 100%.

* fix(errclass): classify the db-domain business codes

Three codes reaching the Apps db commands were absent from the Spark
table, so BuildAPIError fell through to the CategoryAPI + SubtypeUnknown
catch-all and the envelope carried no usable classification.

"App has no database yet" registers as Validation / FailedPrecondition:
the app resolves fine and the request is well-formed, but a prerequisite
the caller must create first is missing, so retrying unchanged can never
succeed. This moves its exit code from 1 to 2 — "fix the state" rather
than "the call failed" — and a test pins that so a future
reclassification has to be deliberate. Two codes cover it because the
server renumbered the case into the 4xx band; the legacy one stays for
older servers.

"Table does not exist" registers as API / NotFound, an ordinary
missing-resource lookup with no exit-code change. SubtypeNotFound has no
APIHint default, which matters here: the Apps layer fills its
command-scoped hint only when the classifier left Hint empty, so a
context-free default would displace the more actionable one. A test
guards that too.
2026-08-07 15:13:31 +08:00
linchao5102 baf6050f8e feat(apps): add role management shortcuts (#1881) 2026-07-16 14:09:41 +08:00