mirror of
https://github.com/larksuite/cli.git
synced 2026-09-14 18:42:53 +08:00
main
23 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e0e90a4e1b |
fix(apps): classify the online DDL ban and the file storage quota failure (#2460)
* fix(apps): classify the online DDL/DCL ban on +db-execute
Running DDL against the online branch of a multi-env app came back as
api/server_error with exit 1, hinted "fix the SQL and re-run", and carried a
statement position that did not exist. All three point the caller the wrong way:
- server_error means "upstream 5xx, retryable"; this is a product rule and no
number of retries changes it;
- the SQL is fine — the target environment is what has to change;
- "(at statement 1 of 1)" is fabricated. The server pre-validates the whole
batch and returns a single ERROR sentinel, so a 5-statement request with the
DDL in position 4 still rendered as "1 of 1". The CLI does not split the SQL,
so it cannot know the real count and cannot detect the mismatch generally —
only codes known to be batch-level rejections can drop the suffix.
Give code 4000001 its own arm: validation/failed_precondition (exit 2, "change
the environment, do not retry"), a hint pointing at dev plus +db-env-migrate, and
no statement position. Every other code keeps its current classification, wording
and position suffix; a test pins that.
4000001 is a dedicated server-side code (ErrOnlineEnvForbidDDLDCL, client-error
band), raised only by the pre-validation pass when env==online on a multi-env
workspace. Syntax errors and PG errors use different codes, so keying on it is
safe. Matching on the numeric value also covers the "k_dl_4000001" wire form,
since codeString already strips that prefix — both forms are tested.
"No statements were applied" is stated rather than inferred here: the validator
walks every statement and rejects the batch on the first DDL, so nothing lands.
The default arm would have inferred the opposite for a DDL in a later position
("Earlier statements were committed"), which is wrong for this code.
* fix(apps): classify tenant file storage quota exceeded
+file-upload against a tenant whose file storage is full returned api/unknown with
no hint at all, so a caller could not tell "the quota is full, stop" from "the
upstream had a bad minute, retry" — and had no next step either.
Register 400000055 as api/quota_exceeded. That is a dedicated server-side code
(ErrTenantStorageQuotaExceeded, client-error band) raised only on the upload path,
and the subtype already carries "retrying will not help", so the framework's
existing quota wording is enough and no domain-specific hint is added.
Left in CategoryAPI (exit 1) rather than Validation (exit 2): a full quota is not
something a different argument fixes, and exit 2 would imply it is.
The test asserts the hint is non-empty on purpose. The wording comes from the
shared APIHint table, so if quota_exceeded is ever dropped from there this fails
and says the code now needs its own wording, instead of silently shipping an
empty hint.
|
||
|
|
ca35f60616 |
fix(apps): make cache-clear ask first, and make apps failures classifiable (#2415)
* docs(skills): require explicit confirmation before apps +cache-clear
Asked to clear an app's online cache, an agent read `Risk: high-risk-write` from
--help and then supplied `--yes` itself on the first call, wiping production
cache without ever hitting the confirmation gate.
The CLI gate is fine: no --yes -> exit 10 confirmation_required, and --dry-run ->
exit 0 without triggering it. The wording was not. It only forbade appending
`--yes` *after* an exit-10, and said "已明确授权可直接带 --yes" without defining
authorization — so "clear my cache" read as authorization.
- `+cache-clear` gets a CAUTION block: never self-supply `--yes` on the first
call; without confirmation, either --dry-run or ask, then stop and wait. exit
10 is not a signal to retry with --yes.
- Add a zero-ambiguity table separating a *request* to clear ("clear the online
cache") from a *confirmation* ("我确认清 dev"), so blocking the accidental wipe
does not also kill the cases that were already correct: an explicit
confirmation still goes straight to `--yes`, and a request with no environment
named still has to ask instead of picking one.
- Note that online needs a confirmation phrase even when named explicitly.
`+cache-delete` gains the response field an agent has to read
(`deleted_key_count`): 0 means the key never existed, not "deleted
successfully", plus the get -> delete -> get chain needed to prove a delete took
effect — a single miss afterwards cannot tell the two apart.
SKILL.md: add +cache-clear to 禁止预授权判定底线, the one list a pre-authorized
run cannot skip; a reference-level rule alone would be bypassed there. The
routing table is left alone — no other row annotates risk, including
+file-delete, +role-delete and +member-remove.
* fix(apps): stop attaching request-shaped hints to precondition failures
`+db-execute` against a tenant that never activated Miaoda returns code 221800
"miaoda UAT not activated" with the hint "verify table/column names with
`+db-table-get` ... target the dev database with --environment dev". Neither step
can help: the failure is tenant-level, so a caller following the hint loops over
table lookups and env retries that fail identically.
Two causes. 221800 was unregistered, so it degraded to api/unknown — nothing in
the envelope distinguished "your tenant is not activated, stop" from "your SQL
was wrong, fix it and retry". And withAppsHint filled the caller's hint whenever
the server sent none, without looking at what failed: the hints are
command-scoped ("verify --app-id", "verify table/column names", "list releases"),
so every one of them describes the request, and the request is exactly what
failed_precondition says was fine.
Register 221800 as validation/failed_precondition (same shape as 400002465 "app
has no database yet") and gate the hint fallback on the subtype.
Blast radius is two codes, since that is all the subtype covers here:
- 221800 — now withheld; message and code still carry the meaning.
- 400002655 "no running container" — only when it reaches a non-observability
command; the observability pair rewrites it first, and "verify --app-id" was
never the fix for an undeployed app.
400002465 / 500002759 are intercepted by the isAppNoDatabaseError branch above
the gate, and 400002479 is served by withDBSyncHint, which does not delegate
here. The other 78 call sites take the original path for every input.
Gate on the one subtype, not on Category: this package asserts on purpose that an
authentication failure on +role-list (99991663) keeps the app-access hint and a
503 on credential issuance keeps the developer-access hint. Those hints are broad
enough to survive a caller-standing failure; only the precondition class is
misdescribed by construction. A test pins that, so widening the gate to Category
fails loudly instead of silently dropping those hints.
The gate is asserted on the real classification path (BuildAPIError -> the code
table -> withAppsHint), not only on a hand-built Problem. Constructing
SubtypeFailedPrecondition directly feeds the gate the input it wants and passes
whether or not 221800 is registered, so the registration itself has to be part of
what the test covers.
No recovery hint for 221800 — the activation path is a product procedure, and
guessing one is what made this failure misleading in the first place.
* fix(apps): classify file-storage and app-level failures
Five Spark business codes reached the CLI unregistered, so every one of them came
out as api/unknown with exit 1: a caller could not tell "your app id is wrong"
from "you lack permission" from "the upstream is having a bad minute", and the
exit code offered no way to branch either.
400002484 app not found -> validation/invalid_argument exit 2
400002467 no admin/developer perm -> authorization/permission_denied exit 3
500002761 ditto, pre-4xx renumber -> same
400000034 file not found/no access -> api/not_found exit 1
500000034 ditto, pre-4xx renumber -> same
400002467 is not file-specific: db commands (+db-table-list, +db-table-get,
+db-quota-get, +db-changelog-list) return it for an app the caller cannot access,
so registering it fixes both domains at once.
400002484 covers a well-formed id that does not exist AND a malformed one
("notanappid", "app_1" return it too), so the argument itself is the failure ->
invalid_argument, whose exit 2 separates "you passed the wrong id" from an
upstream fault. Environments that have not picked it up answer with 400002465
instead, conflating it with "app has no database yet"; the CLI cannot tell those
apart on the old code, so nothing here keys on that.
Both the current and the pre-4xx number are registered for each file failure.
The domain is moving its client-class errors from the 5xxxxxxxx band into
4xxxxxxxx, rolled out per environment, so both are live at once and dropping the
old one would silently return the un-migrated half to api/unknown — the same trap
that made the no-database recovery flow disappear when the server renumbered it
(see appNoDatabaseCode). The new number is not derivable from the old either:
500002761 became 400002467, tail digits included.
No hints added: permission_denied already has framework recovery wording, and a
domain-specific one would have to invent a remedy.
|
||
|
|
1e87f67244 |
fix(apps): friendly-ize "Container not exists" for observability commands (#2302)
* fix(apps): friendly-ize "Container not exists" for observability commands
+metric-list and +analytics-list passed the upstream business code 400002655
("Container not exists") through verbatim. The message reads like an
infrastructure fault and misleads callers (including AI agents) into retrying a
non-retryable, expected business state: an app with no running container simply
has no metrics/analytics to query yet.
Rewrite it at a scoped observability helper (withObservabilityHint) into a
user-facing explanation plus a deploy-then-retry next step, mirroring the
existing isAppNoDatabaseError override. Detection is code-OR-message so a server
renumber alone does not silently drop the rewrite. Classification, code, and the
wrapped cause are preserved; unrelated failures still fall through to the shared
app-id recovery hint (and its own no-database override).
* test(apps): add execution-path regression tests for observability container hint
common_test.go proves withObservabilityHint in isolation but stays green if a
call site reverts to withAppsHint. Drive +metric-list and +analytics-list
Execute with a mocked 400002655 "Container not exists" envelope and assert the
container-specific message/hint/code, so a revert fails the build. Also closes
the two uncovered call-site lines flagged by coverage.
* fix(errclass): classify no-container code as validation/failed_precondition
Register 400002655 in sparkCodeMeta mirroring its no-database twin
(400002465) so both "expected precondition not met" business states expose
the same validation/failed_precondition classification to machine consumers,
instead of falling back to api/unknown. The shortcut-layer message rewrite
already keyed off the raw code, so this only aligns the typed envelope's
category/subtype; update the execution-path tests to pin the new
classification.
* fix(apps): gate the no-container hint's release behind user authorization
The no-container hint told a harness to deploy via +release-create, a "write"
that takes the whole app live and can affect existing production traffic —
without the user-confirmation gate its no-database twin deliberately carries.
Since the hint's audience is an AI agent that acts on it, a failed metrics read
could trigger an unconfirmed go-live. Lead with a read-only +release-list
status check and gate +release-create behind an explicit user confirmation,
mirroring appNoDatabaseHint.
|
||
|
|
158d15b3fd |
feat: add apps database sync shortcuts for Base-to-database import (#2251)
* feat: add apps database sync shortcuts Add Base-to-database sync shortcuts for preview, create, list, get, enable, disable, update, and delete flows. Cover OpenAPI request contracts, typed sync error classification, dry-run E2E coverage, and lark-apps skill guidance. Co-authored-by: TRAE CLI <noreply@bytedance.com> * fix(apps): send db-sync task_id and config in request body The enable/disable/delete/update sync commands placed task_id (and update's config) in query params, but the OpenAPI contract binds these fields via api.json (request body). BOE testing returned "field validation failed" (99992402) because the body was empty. Move task_id to the request body for enable/disable/delete, and move both task_id and config to the body for update. Dry-run previews now render these under body, and unit tests pin the body binding so a regression to query params fails. * fix(apps): use POST for db-sync-delete action endpoint The delete command issued an HTTP DELETE to db/sync_del, but the action-style endpoint is registered as POST (like sync_create and sync_disable). The method mismatch made the gateway return a plaintext 404, surfacing as "API returned a non-object JSON response". Switch the request and dry-run preview to POST, and pin the method in the delete unit tests so a regression to DELETE fails. * test: pin db-sync update base_url as optional contract * test: pin db-sync update omits base_url without silent default * docs(skills): clarify db-sync source.base_url create-required update-optional contract * fix(apps): send db-sync env in request body not query params The +db-sync-create and +db-sync-update endpoints read env from the request body (peer of config/preview/task_id), not the query string. Placing env in query params left the body env empty, so the server treated every request as online and rejected DDL operations (code 500002776: forbid ddl/dcl operation in online env), making it impossible to create/update sync tasks against a dev environment. Move env into the request body via a new dbEnvBody helper that mirrors dbEnvParams' omit-empty contract, so unset env still lets the server auto-select the branch. Pin the contract in unit and e2e dry-run tests by asserting body.env and that env is absent from query params. * test: align db-sync operate/delete e2e with request-body contract The enable/disable/delete dry-run e2e still asserted the pre-migration wire shape: delete on DELETE and task_id in query params. The shortcuts now POST these actions with task_id in the request body (commits moving task_id and the delete verb), so the stale assertions failed against a current binary. Assert POST + body.task_id and that task_id is absent from query params, pinning the same body-over-query contract the env fix established. * fix(apps): improve db-sync create ergonomics and error guidance Refine +db-sync-create/update validation, error hints, and docs so AI agents recover from common Base-to-database sync failures without guessing: - source.table.name: document that a user-named table must be set, name takes precedence over the base_url ?table= token; fix test fixtures that used a fictional source.table.url instead of source.base_url. - Preflight source table locate: reject create locally when base_url has no ?table= and source.table.name is empty, pointing at base +table-list. - Online DDL ban: attach a precise hint for code 500002776 + subcode k_dl_4000001 telling multi-env apps to create tables on --environment dev. - Missing record-id column: extend the 500002783 hint to add a unique text column via +db-execute before retrying. - Optional field_maps on create: allow omitted or empty field_maps so the server auto-matches and creates the task; keep update requiring an enabled mapping and still reject an all-disabled array. - Environment default: db-sync commands use online when --environment is omitted; align help text, comments, and skill docs. * fix(apps): migrate db-sync error codes to the 4xx client-error range The backend moved the seven db-sync error codes from the 5000027xx server-error range to the 4000024xx client-input range to reflect that they are client-input errors. Mirror the new codes in the CLI so error classification and recovery hints keep matching: - 500002783 -> 400002477 (mapping invalid) - 500002784 -> 400002478 (target schema mismatch) - 500002785 -> 400002479 (operation not allowed) - 500002786 -> 400002480 (task not found) - 500002787 -> 400002481 (invalid task id) - 500002788 -> 400002482 (source table not found) - 500002789 -> 400002483 (target table not found) Category, subtype, hint text, and behavior are unchanged; 500002776 (online DDL ban) is untouched. * fix(apps): tighten db-sync preview validation and pretty output Address review follow-ups on the db-sync shortcuts: - +db-sync-get pretty output no longer prints <nil> for a missing schema_only nor Go map syntax for statistics; render a bare bool and deterministic key=value pairs instead. - Reject a non-array field_maps in +db-sync-create --preview as well as commit, so the malformed shape is caught locally rather than forwarded to the backend. - Clarify in lark-apps-db.md that +db-sync-create --preview needs no confirmation and only a real create requires --yes. - Harden the db-sync dry-run validation tests to assert exit code 2 and the structured stderr envelope (type/subtype/param), and add coverage for the preview non-array field_maps rejection and batch pretty output. * fix(apps): guard db-sync preview output and neutralize update hint Address the next db-sync review round: - +db-sync-create --preview --output no longer writes a "null" file and exits success when the response omits data.config; project config into a typed object and return internal/invalid_response without writing. - Make the 400002482 code hint command-neutral so +db-sync-update is not steered into a create-only recovery path that risks duplicate tasks. - lark-apps-db.md: carry --environment on the update lifecycle examples and split failure recovery by streaming (can update) vs batch (cannot update; recreate instead), removing the batch/update contradiction. --------- Co-authored-by: TRAE CLI <noreply@bytedance.com> |
||
|
|
82e628bf79 | fix(drive): harden export and push failure recovery (#2279) | ||
|
|
7c4f6c023f |
fix(base): improve field creation and query guidance (#2114)
* fix(base): improve field creation and query guidance
1. Document batch field creation in shortcut help and the delivered Base skill.
2. Choose field types from stored values instead of business-purpose names.
3. Add generic common filter values to the data-query quick guide with contract tests.
说明:
- Combines the accepted fixes for base_table_096, base_table_028, and base_table_087;
MR 1275 contributes round 4 only.
```ai-signature
改动范围: Base field-create 帮助与 Skill 指南、data-query 快速指南,以及对应的 shortcuts/base 契约测试
思考过程: 保留三个实验的最终通用规则,合并 public main 上新增的字段读回提示,并排除没有 benchmark 支撑的 MR 1275 round6
改动原因: 让代理发现批量字段接口、按存储值选择字段类型,并用常见通用形状构造 data-query 过滤条件
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: a53fdb6b6b82d74e4deff5e6d32591ec897cd2dc6ad662885c961a27e67d4666
* test(base): strengthen guidance contracts
1. Assert every common filter fragment introduced by the data-query quick guide.
2. Keep the field-create argument table compliant with markdown table spacing.
```ai-signature
改动范围: data-query 指南契约测试与 field-create Markdown 表格后的空行
思考过程: 逐条核对 CodeRabbit 建议,只补会防止新增指南片段回退的断言和确定性的 MD058 格式问题,不改生产提示语义
改动原因: 关闭 PR 2114 的两条有效自动审查意见并保持变更可回归
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 81096f474001f2c9c59af48ee64150f205101710ebcf3b909ab9d300a019f46f
* fix(base): generalize data-query filter guidance
1. Replace the date-and-status scenario template with reusable Condition.value shape rules.
2. Update the contract test to require relative-date guidance and reject evaluation-shaped placehold
ers.
```ai-signature
改动范围: lark-base data-query quick guide 与对应 shortcuts/base 契约测试
思考过程: 保留 select、datetime、empty 的通用 value shape,删除日期字段和状态字段组合模板,避免将 base_table_087 的解题路径固化到公共指南
改动原因: benchmark 显示当前文案能引导目标题,但组合示例与测试过度贴合单题,需要收敛为跨场景可复用的不变量
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 7b6916f05dccf6489dce65610fe757f7b346bcd0f8f5500cfbfb1f30c60471f3
* fix(base): generalize creation guidance
* fix(base): report partial field-create results
* fix(base): preserve partial field-create recovery metadata
* fix(base): separate partial field recovery
* perf(base): 缩短字段批量创建的空等
1. 将固定 1 秒批次等待改为 500ms 最小请求起点间隔,并让请求耗时抵扣等待
2. 新增节流计算契约测试,覆盖首次请求、快速响应和慢响应
3. 同输入 150 字段 A/B 从 269.64s 降至 118.80s,且两侧均创建 150/150
说明:
- 保持同表写入串行和 partial failure 输出不变
```ai-signature
改动范围: shortcuts/base/field_ops.go 与 shortcuts/base/base_execute_test.go,仅调整 field-create 数组批次的串行节流计算和回归测试
思考过程: 保留同表串行写入,以 500ms 作为请求起点最小间隔,并把请求耗时计入间隔,避免固定空等同时降低写冲突风险
改动原因: PR 引导 Agent 使用数组批量创建后触发既有每项固定 1 秒等待,导致 150 字段用例产生约 149 秒可归因耗时回退
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 69b888d362d2f484c4ce6ad050bdbfe2de7368948eb79ba516bcaa6806ec9107
* fix(base): 收紧不支持字段行为的终止边界
1. 派生、自动、同步或回填行为只使用已记录能力,无法实现时禁止探测、占位或虚假完成
2. 删除 data-query 契约测试对旧题模板占位符的反向黑名单,只保留原子规则和真实 case 污染检查
3. field-create 指引替换前后均为 36 个英文词,不扩大该帮助项的词数
```ai-signature
改动范围: shortcuts/base/field_create.go、base_shortcuts_test.go 与 data_query_guide_contract_test.go,仅收口通用终止规则和测试泛化
思考过程: 采纳 review 中可独立闭环的两点,不增加翻译专用规则,不修改运行时能力;用等词数替换避免帮助上下文继续增长
改动原因: 当前规则能阻止按业务名猜字段类型,却仍允许退化成普通文本占位;同时测试记住旧题模板会阻碍未来合理示例
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 759b7f4044c6381b963c8ea299b70967e634ab9d43b6ebfbe94ef38bb69ba46a
* perf(base): 压缩字段批量创建的评测开销
1. 批量成功与部分失败仅返回字段 id/name/type,保留恢复所需身份并减少大响应上下文
2. 引导数组在调用方超时范围内一次提交,并为生成的大数组推荐 @file 或 argv-safe 调用
3. 字段列表默认页大小提升到 API 上限 200,避免百字段以上场景的帮助查询与重试
说明:
- 同口径 case032:raw token 432292→428304,weighted token 126408→117761,耗时 272675ms→235022ms,两侧均读回 154 字段
```ai-signature
改动范围: Base field-create 批量输出、帮助提示、field-list 默认分页及对应契约测试与 Skill 返回说明
思考过程: 从同口径 trace 定位固定分块、大字段对象回传、100 条分页和 shell 双重转义四个确定性开销,保持单字段与部分失败恢复语义不变并逐项用测试锁定
改动原因: PR 引导数组批量创建后虽降低耗时,但多回合大输出会推高 raw token;需要在不牺牲正确率和恢复信息的前提下同时压缩 token 与耗时
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 8a2bd2d79c9abdb1e36ae45822dd04b2a3384498a1dc268cffcd47cff00d94ed
* perf(base): 收敛字段批量创建的上下文
1. 大数组成功路径推荐保留摘要的 --jq 投影,失败路径仍原样保留部分失败明细
2. 为一个或多个简单 text 字段提供 help fast path,并让 next_step:done 终止默认回读
3. 补充 Skill、帮助与执行结果契约测试,锁定有界输出和可恢复失败语义
说明:
- 最新 main 同题 A/B:两侧均回读 154 字段,raw token 下降 5.9%,耗时下降 19.3%,峰值上下文下降 13.3%,工具调用下降 20%
```ai-signature
改动范围: Base field-create 帮助、简单字段成功提示、lark-base Skill 路由与对应契约测试
思考过程: 从最终 A/B trace 分别定位批量成功展开 150 项、简单 text 读取冗余指南和成功后整表回读三类可控上下文开销,用成功摘要与失败全量明细分流来保留恢复能力
改动原因: 继续优化 PR 2114 的 token 和耗时,同时要求任何回退不能归因到 PR;需要让大批量成功路径有界且不削弱正确性或部分失败恢复
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 626b288a7f5444dfecb872e6cd29d4fe0788404f5001dcb47aac64e2d55d9b5a
* fix(base): 恢复批量字段输出与分页默认契约
1. 批量创建完整成功时保留服务端字段元数据,部分失败仍返回精简 identity
2. 将 +field-list 默认页大小恢复为 100,继续支持显式 --limit 200
3. 补充回归测试与字段创建文档,保留 --jq 有界输出指导
说明:
- 定向、Base 全量、race、仓库单测、构建、vet 与 lint 均通过
```ai-signature
改动范围: Base 批量字段创建成功输出、字段列表默认分页、对应测试与文档
思考过程: 先用契约测试复现完整字段元数据丢失和默认分页翻倍,再只恢复主干既有成功输出与默认值,同时锁定部分失败精简输出不变
改动原因: 移除可归因到 PR 的兼容性和 token 回退,并保留显式 jq 投影、节流、快速路径及部分失败恢复带来的通用收益
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 6351ed9d228611e3b6f5bc30faf278835b463e07f4eb020f225abd41e264399c
* fix(base): present batch field errors before partial output
1. Project the typed field-create error through Runtime.PresentError before copying result fields
2. Read Error, ProblemOf, and permission extensions from the presented clone
3. Cover visible scoped authorization and concealed recovery without fabricating missing_scopes
说明:
- Targeted, Base, full race, build, vet, format, lint, and module checks pass
```ai-signature
改动范围: Base 批量字段创建部分失败的错误呈现,以及 visible 和 concealed 恢复契约测试
思考过程: 先在最新 main 合并树上复现无 scope 授权提示和隐藏命令泄露,再复用兄弟批量命令的 PresentError 边界,仅替换 payload 复制时的错误来源
改动原因: OutPartialFailure 不会再次呈现根错误,必须在复制 typed error 字段前应用命令 scope 与发行隐藏策略
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: deee600dc8d4bb889629895178c1f217ba7762fbd5b3163da3b9d621c8585ac0
* fix(base): allow recovery before retrying failed field writes
1. Clarify that retryable gates unchanged automatic retries, not corrected resubmissions
2. Classify Base error 1254291 as a retryable conflict with canonical wait guidance
3. Cover authorization recovery, write conflicts, and reference contract consistency
```ai-signature
改动范围: internal/errclass/codemeta_base.go、shortcuts/base/field_ops.go、对应 Base 回归测试与 field-create 参考文档
思考过程: 将 retryable 限定为同一请求原样自动重试资格,保留授权或输入修正后重新提交,并复用现有 conflict 恢复提示
改动原因: 部分失败顶层提示会与权限恢复 hint 冲突,且 1254291 未分类导致等待重试规则无法由结构化错误驱动
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 65cd95f24968c5e0b0c43cf858ed787eabb6d9d63b3d5199a294bd284816c35e
* fix(base): preserve partial field recovery contracts
1. Preserve presented typed-error extensions without allowing them to overwrite batch ledger fields
2. Align field creation guidance with command-specific name semantics and caller-timeout recovery
3. Cover security challenges, extension collisions, and storage-type selection with regression tests
```ai-signature
改动范围: shortcuts/base 的 field-create 部分失败输出、命令提示、Base Skill 写入规则、field-create reference 与对应回归测试
思考过程: 复用 Runtime.PresentError 后 concrete typed error 的 JSON wire shape 作为扩展字段单源;对批次账本自有键统一生成无冲突 error_ 别名,并只收敛已证实的同名、fast path 与 timeout 契约矛盾
改动原因: 部分失败会丢失 challenge_url 等恢复字段,自定义 typed error 还可覆盖 status/index/error 导致错误账本;过度绝对的字段类型、同名和超时文案也会形成可归因正确率回退
Break Change: 否
```
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
AI-SHA256: 02377560f14202ea652f530389ac102647189923e45548337814ec3871f5a4de
---------
Co-authored-by: BASE Infra Harness <ai@base-infra-harness.noreply.local>
|
||
|
|
6402080328 |
fix(apps): detect the no-database failure by code or message (#2217)
* fix(apps): detect the no-database failure by code or message The recovery flow for "db command against an app that has no database" keyed on one business code (500002759). The server has since renumbered that case to 400002465, which silently disabled the flow: users now see the raw internal message about workspace / app-id mapping and lose the cloud-development recovery steps entirely. Nothing catches the regression. There is no compile error, the unit tests compare against the same constant they set, and the dry-run E2E does not exercise a real response — the failure only shows up against a server that has already renumbered. Detect on code OR message instead. Both known codes are kept, plus narrow lowercase markers of the server's internal wording. The two channels have opposite failure modes: a code is precise but gets renumbered, a message survives renumbering but breaks on rewording or localization. Requiring either to match means one channel changing degrades nothing, and only a simultaneous change of both regresses. Markers stay deliberately narrow. "no db branch" in particular must not also swallow env-pull's "invalid db branch" case, which needs its own hint; a comment records that widening them requires a test proving the neighbours still pass through. Classification and the cause chain are untouched: the helper still mutates the problem in place and returns the same error value. * test(apps): assert the full typed-error contract in no-database cases Review feedback: the new subtests checked only Message and Hint, so a change that reclassified the failure — or replaced the error value and dropped the cause chain — would still have passed. Each case now asserts Category, Subtype and Code are untouched by the rewrite, and that the helper returns the same error value. Inputs use a concrete subtype rather than Unknown, so a clobbered classification is actually observable. One new case wraps a cause and asserts errors.Is still finds it through the rewrite. Also covers the predicate's defensive nil guard, which withAppsHint cannot reach on its own (ProblemOf returns ok=false for untyped errors), closing the two uncovered lines the coverage report flagged. Both withAppsHint and isAppNoDatabaseError are now at 100%. * fix(errclass): classify the db-domain business codes Three codes reaching the Apps db commands were absent from the Spark table, so BuildAPIError fell through to the CategoryAPI + SubtypeUnknown catch-all and the envelope carried no usable classification. "App has no database yet" registers as Validation / FailedPrecondition: the app resolves fine and the request is well-formed, but a prerequisite the caller must create first is missing, so retrying unchanged can never succeed. This moves its exit code from 1 to 2 — "fix the state" rather than "the call failed" — and a test pins that so a future reclassification has to be deliberate. Two codes cover it because the server renumbered the case into the 4xx band; the legacy one stays for older servers. "Table does not exist" registers as API / NotFound, an ordinary missing-resource lookup with no exit-code change. SubtypeNotFound has no APIHint default, which matters here: the Apps layer fills its command-scoped hint only when the classifier left Hint empty, so a context-free default would displace the more actionable one. A test guards that too. |
||
|
|
09feefe96b | fix: make agent recovery and concealment reliable (#2189) | ||
|
|
8e884928f6 |
feat: add Base table copy shortcuts (#2019)
* feat: add Base table copy shortcuts * test: cover Base table copy edge cases * fix: preserve table copy task state * fix: align table copy recovery with API errors * fix: preserve table copy auth recovery * test: verify copied table schema --------- Co-authored-by: fongwave <272393974+fongwave@users.noreply.github.com> |
||
|
|
ebdeda854d | feat(extension): present restricted commands as absent and trim skills (#1837) | ||
|
|
baf6050f8e | feat(apps): add role management shortcuts (#1881) | ||
|
|
80fadf1801 | fix(drive): abort push on parent sibling limit (#1813) | ||
|
|
f98dbfe247 | Improve agent-facing error guidance for drive, markdown, and wiki (#1779) | ||
|
|
578e2db4e0 | fix: point permission-apply link at official /page/scope-apply entry (#1722) | ||
|
|
d0cde9a414 |
Improve secure label error handling (#1707)
* Improve secure label error handling * Address secure label review feedback |
||
|
|
a6797ac2e4 | Improve drive batch failure handling (#1703) | ||
|
|
842be3fdc5 | feat(token): mint TAT via unified OAuth v3 Token Endpoint (#1408) | ||
|
|
076f4d579f |
feat(minutes,vc): emit typed error envelopes across both domains (#1234)
Failures from the minutes and video-conference commands now surface as structured, typed errors carrying a stable category and subtype — spanning input validation, missing permissions, network and file-I/O failures, and remote API errors — so callers can branch on the error kind instead of parsing free-form text. Batch commands report partial failures explicitly, emitting per-item results with a non-zero exit instead of masking them. |
||
|
|
f3949f04c4 |
feat(calendar): emit typed error envelopes across the calendar domain (#1232)
Calendar commands now return structured, typed error envelopes for every failure mode — input validation, internal faults, and API responses — instead of legacy generic errors. Callers and AI agents get consistent exit codes and a machine-readable shape (type / subtype / code / hint), and can tell bad input, an internal fault, and an API rejection apart. Validation errors are attributed to the offending flag. Server-supplied error details (e.g. why an event time was rejected) are surfaced on the typed error's hint via a shared classifier improvement that benefits every domain. Multi-step operations (create-with-attendees rollback, multi-field update) preserve the real failure's classification and report which steps completed. The whole calendar domain is now lint-locked against reintroducing legacy error constructors. |
||
|
|
5e6a3eb857 |
feat(mail): return typed error envelopes across the mail domain (#1250)
* feat(mail): return typed error envelopes across the mail domain Replace every produced error path in shortcuts/mail with typed errs.* envelopes, so consumers get stable category, subtype, param/params, hint, retryable, and log_id metadata for classification and recovery instead of free-form message text. - Locally constructed mail errors move from output.Err* / output.Errorf / final fmt.Errorf / common legacy helpers to errs.* builders, with structured params on multi-flag validation and failed-precondition states kept non-retryable. - API-call failures move from runtime.CallAPI / DoAPIJSON legacy boundaries to runtime.CallAPITyped or runtime.ClassifyAPIResponse, and mail-specific enrichers read errs.ProblemOf so typed code, subtype, hint, and log_id metadata are preserved. - Batch draft-send partial failures now use runtime.OutPartialFailure so successful and failed draft sends stay in stdout while the command exits through a typed multi-status signal. - Add mail-domain typed helpers, mail API code metadata, and guard wiring to keep shortcuts/mail from reintroducing legacy envelopes or legacy API calls. - Keep genuine intermediate fmt.Errorf wraps in parser/builder layers annotated with nolint comments; command-facing paths wrap them into typed validation, API, network, or internal errors. * fix(mail): report aborted draft-send batches as a single failure result When an account-level failure interrupts a batch send after some drafts already went out, the command previously produced two machine-readable failure results: the partial-failure ledger on stdout and a second error envelope on stderr. Consumers could not tell which one to recover from. The batch ledger is now the only failure result for that case: it gains aborted and abort_error fields carrying the typed cause, so callers can see which drafts were sent, which failed, why the batch stopped, and how to recover — all from stdout. A --stop-on-error stop keeps these fields unset because stopping early there is the caller's own choice. |
||
|
|
98173ae5a9 |
feat(drive): emit typed error envelopes across the drive domain (#1205)
Drive-domain errors now leave the CLI as typed, machine-branchable envelopes — a stable `type` plus `subtype` and named fields (param, params, retryable, log_id, hint) — so scripts and AI agents can branch on structure and act on a recovery hint instead of parsing prose. Changes: - Every error produced in the drive domain — validation, file I/O, and the failures returned from its Lark API calls — is emitted as a typed errs.* error; the exit code is derived from the error category. Drive's API calls now go through a shared typed classifier, so failures carry subtype, troubleshooter, a recovery hint, and the request's log_id whether the server returns it in the response body or the x-tt-logid header; an already-typed network/auth error is never downgraded into a generic API error. - Known API conditions (resource conflict, cross-tenant, cross-brand, ...) carry a recovery hint keyed by their error class; a command can refine that hint with command-specific guidance. - Batch partial failures (+push / +pull / +sync, where some items succeed and some fail) now report an honest ok:false multi-status result on stdout — the summary and every per-item outcome stay machine-readable — and exit non-zero, instead of a misleading ok:true success envelope. - Duplicate rel_path conflicts report each colliding path as a structured params entry (RFC 7807 invalid-params style). - Static guards lock the drive path so legacy error construction — direct envelopes or the auto-classifying API helpers — cannot be reintroduced, making drive the template for the remaining domains. Output changes worth noting for consumers: - Error envelopes now carry typed type/subtype and named fields; exit codes follow the error category (malformed or incomplete API responses are reported as internal errors rather than generic API errors). - Batch partial failures (+push / +pull / +sync) emit an ok:false result envelope on stdout (summary + per-item items[]) and exit non-zero; the per-item results stay on stdout rather than in a stderr error envelope. Errors surfaced through shared cross-domain helpers (scope precheck, media import upload, metadata lookup, save-path resolution) are not yet typed; they migrate with the shared layer in a follow-up change. |
||
|
|
99e314fe0b |
feat(errs): typed envelope contract for auth-domain errors (#1135)
Every failure on the authentication, authorization, and configuration
path now surfaces as a typed structured error instead of an ad-hoc
envelope. Users and scripts that consume CLI output get:
- a fixed nine-category taxonomy on the wire, each mapped to a
stable shell exit code (authentication/authorization/config = 3,
network = 4, internal = 5, policy = 6, confirmation = 10)
- identity-aware detail fields (missing_scopes, requested_scopes,
granted_scopes, console_url, log_id, retryable, hint) carried
uniformly on the envelope
- a single canonical policy envelope at exit 6; the legacy
auth_error carve-out is retired
- per-subtype canonical message + hint that preserves Lark's
diagnostic phrasing and routes recovery to the right actor:
app developer (app_scope_not_applied), user (missing_scope,
token_scope_insufficient, user_unauthorized), or tenant admin
(app_unavailable, app_disabled)
- wrong app credentials classify as config/invalid_client whether
surfaced by the Open API endpoint (99991543) or the tenant
access-token mint endpoint (10003 / 10014), instead of
collapsing to a transport error or api/unknown
- local shortcut scope preflight emits the same
authorization/missing_scope envelope (identity + deterministic
missing-scope set) used by the post-call permission path, so AI
consumers read the same structured shape from precheck and from
server-returned permission denial
- streaming download/upload failures keep the same network subtype
split (timeout / TLS / DNS / transport) as the non-stream path
instead of collapsing every cause to a generic transport failure
- console_url is carried only on the bot-perspective
app_scope_not_applied envelope (where the recovery action is
"developer applies the scope at the developer console"); the
user-perspective missing_scope envelope drops the field, since
the only actionable user recovery is `lark-cli auth login --scope`
and pointing an end user at a console they cannot modify is
misleading
- bind workflows (Hermes / OpenClaw / lark-channel) flatten dynamic
Type tags to wire 'config' with the original module name kept
as a metric label
All 10 typed errors are cause-bearing, nil-safe on .Error() and
.Unwrap(), and defensively clone slice setter inputs. Four lint
rules (CheckNilSafeError / CheckBuilderImmutable / CheckUnwrapSymmetry
/ CheckBuildAPIErrorArms) lock these invariants on migrated paths.
|
||
|
|
fe72e41fb2 |
feat(errs): add structured CLI error contract (#984)
Introduce a typed error contract framework for lark-cli so in-process
Go callers can branch via errors.As(&errs.XxxError{}) and shell scripts,
AI agents, and protocol adapters can branch on stable JSON type/subtype
fields instead of regex-parsing free-form messages.
Adds:
- Canonical taxonomy under errs/ (9 categories + typed Error structs
embedding a shared Problem, RFC 7807-aligned)
- Centralized Lark code metadata + identity-aware BuildAPIError dispatch
- Typed JSON envelope writer alongside the legacy envelope writer
- MCP / OAuth (RFC 6750 Bearer) projection adapters
- Five CI lint guards preventing ad-hoc taxonomy drift
Backward compatibility: legacy *output.ExitError producers (ErrAPI,
ErrWithHint, Errorf, ErrBare) and business shortcuts that use them
continue to render the legacy envelope unchanged. SecurityPolicyError
wire format and exit code are preserved via a carve-out; taxonomy
migration is deferred to PR 2. Domain-specific business migration is
staged across PR 3+.
Framework-direct paths now return typed *errs.*Error: ErrAuth /
ErrValidation / ErrNetwork emit category literals on the wire
(authentication / validation / network), *core.ConfigError is promoted
at the cmd/root boundary with exit code aligned from 2 to 3, and Lark
API permission denials classified by BuildAPIError exit 3.
At the SDK boundary, WrapDoAPIError preserves any already-classified
error (legacy *output.ExitError or typed *errs.*) so output.ErrAuth
from missing credentials surfaces with the auth category and exit 3
intact instead of being downgraded to a network error. Policy responses
classified by BuildAPIError (codes 21000 / 21001) extract challenge_url
and the canonical hint from the response body, matching what the
auth transport already surfaces at the HTTP layer; non-https
challenge URLs are dropped.
First PR in the feat/error-contract-* series.
|